Key benefits
Zero Hallucinations
The HDC cosine-similarity gate operates at threshold 0.55. Queries that don’t match stored hypervectors are blocked and returned as honest refusals — never passed to the LLM for creative gap-filling.
Under 1.2 GB VRAM
The full three-tier pipeline fits in less than 1.2 GB of VRAM, comfortably within GTX 1070-class hardware. CPU-only mode requires no GPU at all — useful for Raspberry Pi and edge deployments.
100% Offline & Private
No API keys, no cloud calls, no telemetry. Hillock talks only to your local Ollama instance over
localhost. Your documents never leave your machine.Fast Tensor Ingestion
TALON (Tensor-Accelerated Local Ontology Network) uses tensor-based O(1) matrix classification to extract facts — no LLM pass during ingestion. A 30-sentence document ingests in roughly 5 seconds.
OpenAI-Compatible API
Hillock exposes an OpenAI-compatible REST server on
http://localhost:8000/v1. Any tool that speaks the OpenAI chat completions format — Open-WebUI, AnythingLLM, Obsidian — works out of the box.Three Answering Modes
Switch between STRICT (single-sentence facts only), BALANCED (fact + light context + source citation), and CONVERSATIONAL (warm, associative responses using Hebbian memory links).
How it works
Hillock’s pipeline has three stages that happen every time you ingest a document or ask a question:- Ingest: TALON parses your document with spaCy and GLiREL, resolves coreferences via FastCoref, and writes normalized SPO triples into the SQLite Knowledge Graph. No LLM is involved in this step.
- Store: The Knowledge Graph holds entity nodes and typed relation edges. The Hebbian Plasticity Engine strengthens association weights between co-activated entities over repeated queries, surfacing contextually relevant priming.
- Retrieve with HDC gate: At query time, the Hyperdimensional Reservoir encodes the query into a 10,000-dimensional binary hypervector and computes cosine similarity against stored fact hypervectors. Only facts that clear the 0.55 threshold are forwarded to the LLM renderer — everything else triggers a refusal.
Who is Hillock for?
Hillock is purpose-built for four groups of engineers and researchers: AI engineers building agents that need auditable, deterministic memory over private document sets — contracts, internal wikis, research papers — where hallucination is a business risk, not just an annoyance. Local LLM developers running Ollama, LM Studio, or llama.cpp on consumer GPUs who’ve outgrown vector databases and need something that fits inside their existing VRAM budget alongside the main language model. Privacy-focused researchers working with sensitive data — medical records, legal documents, proprietary datasets — where sending document content to any cloud endpoint is a non-starter. Edge hardware users targeting Raspberry Pi, Jetson Nano, or low-power x86 hardware, where Hillock’s CPU-only mode and minimal RAM footprint make deployment practical without a dedicated GPU.Get Started in 5 Minutes
Install Hillock, pull llama3.2 from Ollama, and run your first query from the CLI.
Why Not Traditional RAG?
See a side-by-side comparison of Hillock vs. Chroma, Pinecone, and standard RAG pipelines.