Skip to main content
Traditional Retrieval-Augmented Generation has a dirty secret: it doesn’t actually prevent hallucinations, it just changes their flavour. Vector databases retrieve the closest embedding chunks probabilistically, and when those chunks don’t contain a precise answer, the LLM fills the gap with plausible-sounding fabrication. Add in the VRAM cost of running a dense embedding model alongside an 8B+ generative model, the fragility of chunk-boundary drift, and a hard dependency on cloud infrastructure for most production pipelines, and you have a system that is simultaneously expensive, unreliable, and impossible to audit. Hillock was built as a direct answer to these failure modes.
Common failure modes in Chroma / Pinecone / LlamaIndex pipelines:
  • High VRAM requirements: Running a dense embedding model (e.g. text-embedding-ada-002 class) alongside a 7B–13B generative LLM easily exceeds 8–16 GB VRAM, requiring a dedicated GPU or cloud endpoint.
  • Chunked embedding drift: Documents are split into fixed-size chunks, breaking semantic context at arbitrary boundaries. Embedding distance between a query and a relevant chunk degrades with chunk size mismatches.
  • Hallucination from LLM fill-in: When retrieved chunks don’t contain the answer, the LLM synthesizes a plausible-sounding response from its parametric memory. There is no hard block — the model always produces output.
  • Cloud and GPU infrastructure dependency: Most production RAG setups rely on managed vector databases (Pinecone, Weaviate, Qdrant cloud) or require a GPU-equipped server, making fully offline or air-gapped deployments impractical.
  • Probabilistic retrieval with no audit trail: Nearest-neighbor vector search returns the most similar chunks, not necessarily the correct ones. There is no structured record of what facts were retrieved or why.

Architecture comparison

The hallucination problem in RAG is structural, not accidental. When a user asks a vector RAG system a question for which no retrieved chunk contains a direct answer, there is nothing in the architecture that prevents the LLM from answering. The model was trained to produce fluent, helpful-sounding text — it will do exactly that, even when it has no grounded information. In high-stakes use cases (legal, medical, financial, compliance), this is not an acceptable failure mode. Confidence scores on vector similarity do not solve this: a 0.87 cosine similarity between a query embedding and a tangentially related chunk does not mean the chunk answers the question.
How Hillock’s HDC gate works: When you query Hillock, the Hyperdimensional Reservoir encodes your query into a 10,000-dimensional binary hypervector using GloVe-seeded SimHash. It then computes the late-interaction MaxSim score between the query hypervector and the hypervectors of all candidate facts retrieved from the Knowledge Graph. Only facts that score ≥ 0.55 AND achieve a predicate alignment score ≥ 0.35 are forwarded to the LLM renderer. If no facts clear both thresholds, Hillock returns "I do not have verified information about that." — a deterministic, non-hallucinated response.

When to use Hillock

Hillock is the right choice when:
  • You’re working with private or sensitive documents — contracts, medical records, legal briefs, proprietary research — that cannot leave your machine or be processed by cloud APIs.
  • You’re deploying on edge hardware — Raspberry Pi, Jetson Nano, consumer laptops, air-gapped servers — where GPU VRAM is scarce or unavailable.
  • You need offline-capable agents that must function without internet connectivity, e.g. in field deployments, secure facilities, or bandwidth-constrained environments.
  • You require an audit trail of exactly what facts were retrieved, from which source document, and with what confidence — for compliance, explainability, or debugging purposes.
  • Your document set is relatively structured and factual — technical documentation, research papers, structured reports — where SPO triple extraction is high-fidelity.
Traditional RAG may still be the better fit when:
  • You have a massive unstructured corpus (millions of documents, diverse formats, heavily narrative text) where symbolic triple extraction would be lossy and the sheer scale benefits from distributed vector indexing.
  • You need fuzzy semantic search at scale — e.g. “find documents similar in topic to this one” — where approximate nearest-neighbor search is the actual product requirement, not precise factual recall.
  • Your use case is exploratory or creative rather than factual — brainstorming, summarization, or content generation tasks where some degree of LLM latitude is a feature, not a bug.