> ## Documentation Index
> Fetch the complete documentation index at: https://hillock.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Hillock vs. Traditional RAG: A Technical Comparison

> Compare Chroma and Pinecone-style vector RAG against Hillock on VRAM usage, hallucination risk, retrieval determinism, and fully offline deployment.

Traditional Retrieval-Augmented Generation has a dirty secret: it doesn't actually prevent hallucinations, it just changes their flavour. Vector databases retrieve the closest embedding chunks probabilistically, and when those chunks don't contain a precise answer, the LLM fills the gap with plausible-sounding fabrication. Add in the VRAM cost of running a dense embedding model alongside an 8B+ generative model, the fragility of chunk-boundary drift, and a hard dependency on cloud infrastructure for most production pipelines, and you have a system that is simultaneously expensive, unreliable, and impossible to audit. Hillock was built as a direct answer to these failure modes.

<Tabs>
  <Tab title="Traditional RAG">
    **Common failure modes in Chroma / Pinecone / LlamaIndex pipelines:**

    * **High VRAM requirements:** Running a dense embedding model (e.g. `text-embedding-ada-002` class) alongside a 7B–13B generative LLM easily exceeds 8–16 GB VRAM, requiring a dedicated GPU or cloud endpoint.
    * **Chunked embedding drift:** Documents are split into fixed-size chunks, breaking semantic context at arbitrary boundaries. Embedding distance between a query and a relevant chunk degrades with chunk size mismatches.
    * **Hallucination from LLM fill-in:** When retrieved chunks don't contain the answer, the LLM synthesizes a plausible-sounding response from its parametric memory. There is no hard block — the model always produces output.
    * **Cloud and GPU infrastructure dependency:** Most production RAG setups rely on managed vector databases (Pinecone, Weaviate, Qdrant cloud) or require a GPU-equipped server, making fully offline or air-gapped deployments impractical.
    * **Probabilistic retrieval with no audit trail:** Nearest-neighbor vector search returns the *most similar* chunks, not necessarily the *correct* ones. There is no structured record of what facts were retrieved or why.
  </Tab>

  <Tab title="Hillock">
    **How Hillock's three-tier architecture resolves each failure mode:**

    * **Under 1.2 GB VRAM:** The TALON extractor uses a lightweight MiniLM model for predicate routing (exportable to ONNX FP16), the HDC reservoir uses binary hypervectors, and the Knowledge Graph is plain SQLite. The entire pipeline runs alongside a local Ollama model on a GTX 1070 or equivalent.
    * **Symbolic SPO triples, not chunks:** Documents are decomposed into typed Subject-Predicate-Object relations (e.g. `alan_turing → cracked → enigma_cipher`). There is no chunking — every fact is a discrete, addressable unit.
    * **Mathematical HDC gate blocks hallucinations:** Queries are encoded as 10,000-dimensional binary hypervectors. Only queries whose cosine similarity to stored fact hypervectors clears the 0.55 threshold reach the LLM. Everything else returns an honest refusal.
    * **CPU and integrated GPU compatible:** CPU-only mode is explicitly supported. The ONNX export path (`export_to_onnx.py`) produces an FP16 MiniLM model that runs on the CPU execution provider with no GPU required.
    * **Deterministic retrieval with full audit trail:** Every retrieved fact includes its source document, confidence score, and predicate type. You can inspect exactly what the engine knows with `/inspect <entity>` in the CLI or by querying the SQLite database directly.
  </Tab>
</Tabs>

## Architecture comparison

| Aspect | Traditional RAG | Hillock |
| - | - | - |
| **Memory model** | Dense vector embeddings (floating-point) | Symbolic SPO triples + binary HDC hypervectors |
| **VRAM requirement** | 8–16 GB (embedding model + LLM) | Under 1.2 GB (CPU-only mode also available) |
| **Hallucination risk** | High — LLM fills gaps from parametric memory | Near-zero — HDC gate blocks unverifiable queries |
| **Retrieval method** | Probabilistic k-NN vector similarity | Deterministic HDC cosine gate at threshold 0.55 |
| **Ingestion method** | LLM or embedding model chunking pass | TALON tensor classification, \~5 sec / 30 sentences |
| **Offline support** | Limited — most vector DBs require cloud or GPU server | Full — SQLite + local Ollama, zero external calls |

<Warning>
  **The hallucination problem in RAG is structural, not accidental.** When a user asks a vector RAG system a question for which no retrieved chunk contains a direct answer, there is nothing in the architecture that prevents the LLM from answering. The model was trained to produce fluent, helpful-sounding text — it will do exactly that, even when it has no grounded information. In high-stakes use cases (legal, medical, financial, compliance), this is not an acceptable failure mode. Confidence scores on vector similarity do not solve this: a 0.87 cosine similarity between a query embedding and a tangentially related chunk does not mean the chunk answers the question.
</Warning>

<Note>
  **How Hillock's HDC gate works:** When you query Hillock, the Hyperdimensional Reservoir encodes your query into a 10,000-dimensional binary hypervector using GloVe-seeded SimHash. It then computes the late-interaction MaxSim score between the query hypervector and the hypervectors of all candidate facts retrieved from the Knowledge Graph. Only facts that score ≥ 0.55 AND achieve a predicate alignment score ≥ 0.35 are forwarded to the LLM renderer. If no facts clear both thresholds, Hillock returns `"I do not have verified information about that."` — a deterministic, non-hallucinated response.
</Note>

## When to use Hillock

**Hillock is the right choice when:**

* You're working with **private or sensitive documents** — contracts, medical records, legal briefs, proprietary research — that cannot leave your machine or be processed by cloud APIs.
* You're deploying on **edge hardware** — Raspberry Pi, Jetson Nano, consumer laptops, air-gapped servers — where GPU VRAM is scarce or unavailable.
* You need **offline-capable agents** that must function without internet connectivity, e.g. in field deployments, secure facilities, or bandwidth-constrained environments.
* You require an **audit trail** of exactly what facts were retrieved, from which source document, and with what confidence — for compliance, explainability, or debugging purposes.
* Your document set is **relatively structured and factual** — technical documentation, research papers, structured reports — where SPO triple extraction is high-fidelity.

**Traditional RAG may still be the better fit when:**

* You have a **massive unstructured corpus** (millions of documents, diverse formats, heavily narrative text) where symbolic triple extraction would be lossy and the sheer scale benefits from distributed vector indexing.
* You need **fuzzy semantic search at scale** — e.g. "find documents similar in topic to this one" — where approximate nearest-neighbor search is the actual product requirement, not precise factual recall.
* Your use case is **exploratory or creative** rather than factual — brainstorming, summarization, or content generation tasks where some degree of LLM latitude is a feature, not a bug.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.