- Traditional RAG
- Hillock
Common failure modes in Chroma / Pinecone / LlamaIndex pipelines:
- High VRAM requirements: Running a dense embedding model (e.g.
text-embedding-ada-002class) alongside a 7B–13B generative LLM easily exceeds 8–16 GB VRAM, requiring a dedicated GPU or cloud endpoint. - Chunked embedding drift: Documents are split into fixed-size chunks, breaking semantic context at arbitrary boundaries. Embedding distance between a query and a relevant chunk degrades with chunk size mismatches.
- Hallucination from LLM fill-in: When retrieved chunks don’t contain the answer, the LLM synthesizes a plausible-sounding response from its parametric memory. There is no hard block — the model always produces output.
- Cloud and GPU infrastructure dependency: Most production RAG setups rely on managed vector databases (Pinecone, Weaviate, Qdrant cloud) or require a GPU-equipped server, making fully offline or air-gapped deployments impractical.
- Probabilistic retrieval with no audit trail: Nearest-neighbor vector search returns the most similar chunks, not necessarily the correct ones. There is no structured record of what facts were retrieved or why.
Architecture comparison
How Hillock’s HDC gate works: When you query Hillock, the Hyperdimensional Reservoir encodes your query into a 10,000-dimensional binary hypervector using GloVe-seeded SimHash. It then computes the late-interaction MaxSim score between the query hypervector and the hypervectors of all candidate facts retrieved from the Knowledge Graph. Only facts that score ≥ 0.55 AND achieve a predicate alignment score ≥ 0.35 are forwarded to the LLM renderer. If no facts clear both thresholds, Hillock returns
"I do not have verified information about that." — a deterministic, non-hallucinated response.When to use Hillock
Hillock is the right choice when:- You’re working with private or sensitive documents — contracts, medical records, legal briefs, proprietary research — that cannot leave your machine or be processed by cloud APIs.
- You’re deploying on edge hardware — Raspberry Pi, Jetson Nano, consumer laptops, air-gapped servers — where GPU VRAM is scarce or unavailable.
- You need offline-capable agents that must function without internet connectivity, e.g. in field deployments, secure facilities, or bandwidth-constrained environments.
- You require an audit trail of exactly what facts were retrieved, from which source document, and with what confidence — for compliance, explainability, or debugging purposes.
- Your document set is relatively structured and factual — technical documentation, research papers, structured reports — where SPO triple extraction is high-fidelity.
- You have a massive unstructured corpus (millions of documents, diverse formats, heavily narrative text) where symbolic triple extraction would be lossy and the sheer scale benefits from distributed vector indexing.
- You need fuzzy semantic search at scale — e.g. “find documents similar in topic to this one” — where approximate nearest-neighbor search is the actual product requirement, not precise factual recall.
- Your use case is exploratory or creative rather than factual — brainstorming, summarization, or content generation tasks where some degree of LLM latitude is a feature, not a bug.