Skip to main content
Standard LLMs have no mechanism to refuse a question they cannot answer from verified knowledge — they generate plausible-sounding text regardless of whether that text is grounded in fact. Hillock solves this with a deterministic gate that sits between user queries and the LLM renderer. Before any generative call is made, every query is encoded as a 10,000-dimensional hypervector and scored against the knowledge graph using cosine similarity. If the score falls below a configurable threshold, the query is blocked and a refusal is returned — no LLM call, no hallucination possible. The gate is mathematically deterministic: the same query always produces the same gate decision.

The HYDRA MaxSim Gate

Every user query passes through the HYDRA scoring pipeline before any retrieval or generation happens.

Step-by-Step Gate Flow

Configuring the Threshold

The threshold is set as HDC_THRESHOLD = 0.55 in config.py. Adjust it based on your precision/recall tradeoff:
Lower values let more queries through (higher recall, more risk of weak matches). Higher values enforce strict semantic proximity (higher precision, may reject valid paraphrases).

Three Answering Modes

When a query passes the gate, the matched facts are handed to the LLM for rendering. The rendering behavior is controlled by the verbosity_mode setting, which governs the system prompt and what additional context (priming, HDC traces) is included.
Only the verified facts from the knowledge graph, translated into one sentence. No inference, no added context, no elaboration. The LLM acts as a fact formatter, not a reasoner.System prompt excerpt:
“You are a professional fact renderer. Translate ONLY the provided fact into one sentence. Do not add any extra context, historical assumptions, or details.”
Best for: Production agents where hallucination is unacceptable. Regulatory, medical, or legal contexts. Automated pipelines where responses are parsed programmatically.

Mathematical Guarantee

The gate is a cosine similarity comparison in a 10,000-dimensional Euclidean space. There are no random tie-breaking operations, no sampling, and no temperature parameters involved. The same query string, encoded with the same GloVe vocabulary and the same fixed random projection matrix R (seeded at seed=42), always produces the same hypervector and thus the same gate decision.

The Cosine Similarity Formula

Where:
  • q — the query hypervector (mean of per-token hypervectors)
  • d — the fact hypervector (predicate + entity components)
  • · — dot product
  • |·| — L2 norm
In bipolar ±1 space, this simplifies to:
Where D = 10000. The cosine similarity score is bounded in [-1.0, 1.0]. Random orthogonal hypervectors score near 0.0; near-identical encodings score near 1.0. The threshold of 0.55 sits well above the noise floor of random pairs at this dimensionality.

Why High Dimensionality Matters

At 10,000 dimensions, the expected cosine similarity between two randomly generated bipolar hypervectors is 0.0 with a standard deviation of 1/√D ≈ 0.01. This means scores above 0.55 are more than 55 standard deviations from the noise floor — statistically impossible for unrelated concepts to produce false positives. This is the concentration of measure property that makes HDC gates reliable.

What a Blocked Query Looks Like

When a query fails the gate, Hillock returns a structured refusal. The exact format depends on verbosity_mode:
No LLM call is made. The response is a hardcoded string, returned in microseconds.

Inspecting Gate Decisions

Enable debug logging to see HYDRA scores for every evaluated fact:

Seed knowledge and benchmarks: 4 of Hillock’s 7 seed triples overlap common knowledge-graph evaluation target sets (e.g., Marie_Curie discovered Radioactivity, Alan_Turing cracked Enigma). If you’re running your own accuracy benchmarks, either reset the database with kg.clear_and_reinitialize() before seeding with your own data, or account for the 4 overlapping triples in your evaluation methodology. Failure to do so will artificially inflate recall metrics on standard KG benchmarks.
Use STRICT mode for production agents. CONVERSATIONAL mode is designed for exploratory, human-facing interactions. In any automated pipeline — tool-use agents, RAG pipelines, structured data extraction — set hillock.verbosity_mode = "STRICT" to eliminate any possibility of the LLM renderer adding unverified elaborations. STRICT mode also skips the LLM call entirely on refusals, making blocked queries essentially free in terms of latency and token cost.