Tier 1: SQLite Knowledge Graph
The base tier is a relational fact store backed by SQLite. Every piece of knowledge Hillock holds is represented as a Subject–Predicate–Object (SPO) triple persisted in a structured schema. Retrieval is deterministic: given an entity ID and a predicate, the answer is either in the table or it isn’t — there is no probability distribution, no approximate match, no hallucination pathway.Schema
The knowledge graph uses four tables:
The
relations table schema:
The Predicate Taxonomy
All extracted relations are constrained to a fixed taxonomy of 53 Wikidata-style predicates. This prevents predicate explosion and keeps the graph queryable. The taxonomy covers:- Origin / location:
born_in,died_in,place_of_birth,place_of_death,country_of_citizenship,resided_in,migrated_to,moved_to - Social / professional:
collaborated_with,worked_with,partnered_with,spouse_of,child_of,parent_of,member_of,affiliated_with - Creation / discovery:
discovered,invented,co_invented,developed,designed,founded,created,authored,wrote,published,formulated,proposed,cracked,patented - Institutional:
educated_at,studied_at,employed_by,worked_at,field_of_work - Awards:
award_received,won,nominated_for - Geo / hierarchy:
capital_of,located_in,headquartered_in,subclass_of,part_of,instance_of,has_part,contains - Influence / succession:
influenced_by,student_of,teacher_of,successor_to,predecessor_to - Industry:
manufactured,operated_by
Sample Triple
(source_id, predicate, target_id), duplicate extraction runs are idempotent — re-ingesting the same document will never create duplicate rows.
Tier 2: Hebbian Synaptic Memory
The second tier implements gradient-free co-activation learning. Whenever two or more entities appear together in an answered query, their association weight increases. Weights that are not recently reinforced gradually decay. This mirrors the neuroscientific Hebbian principle: neurons that fire together, wire together.The Learning Rule
For every pair(entity_a, entity_b) that co-activates during a successful retrieval:
η = 0.15 (HEBBIAN_ETA in config.py). This is a saturating update — weights approach 1.0 asymptotically and never exceed it.
For all pairs not active in the current turn, weights decay by:
γ = 0.01 (HEBBIAN_DECAY). Decay is applied globally every turn, so stale associations fade over time.
Retrieval Threshold
When assembling context for a query, only entity pairs withweight > 0.05 are surfaced as priming context. The get_associated_priming_context() method returns associated entities ordered by descending weight:
Role in Query Answering
Priming context from Tier 2 is passed directly to the LLM prompt in BALANCED and CONVERSATIONAL modes. This enables the system to surface related topics the user didn’t explicitly ask about — without inventing facts. In CONVERSATIONAL mode, the LLM is instructed to ask the user if they want to hear about the associated entities, creating a natural memory-browsing experience.Tier 3: Hyperdimensional Computing (HDC / VSA)
The third tier is a Vector Symbolic Architecture operating in 10,000-dimensional bipolar hypervector space (±1 per dimension). This tier serves two distinct functions: semantic similarity gating (blocking queries that don’t match any known facts) and fading-memory context tracking (maintaining a rolling state of what the conversation has been about).
Hypervector Dimensions
All hypervectors areD = 10000 dimensions, configured as HDC_DIMENSION in config.py. At this dimensionality, the probability of two random hypervectors having cosine similarity above 0.1 is negligibly small — this is the mathematical foundation of the hallucination gate.
Dual-Path Encoding
Every token and entity name is encoded via two complementary paths, then fused: Path A — SubwordHDCEncoder (morphological): Character n-grams of lengths 3, 4, and 5 are extracted from the token. Each n-gram is hashed via MD5 to a deterministic random seed, which generates a±1 bipolar hypervector. All n-gram hypervectors are superposed (summed and sign-binarized) into a single morphological representation. Typos and inflected forms that share most n-grams will land in nearby hypervector regions.
Path B — SignRandomProjectionSimHash (semantic):
If the token exists in the GloVe 50d vocabulary (top 50,000 words, ~10MB RAM), its dense 50-dimensional vector is projected into 10,000 dimensions via a fixed random matrix R (shape 10000 × 50). The sign of each projected component produces a bipolar SimHash hypervector. Semantically similar words — whose GloVe vectors are close — produce SimHash vectors with high cosine similarity.
Fusion: Both paths are superposed and binarized:
Fading Memory Reservoir
The reservoir maintains a continuous state vector that summarizes recent conversational context. Each time a query is processed, every token in the query steps the reservoir:decay = 0.95 (HDC_DECAY). The roll(state) * token_hv term is the binding operation — it creates a context-sensitive representation that encodes not just what tokens appeared, but the sequential structure of the conversation. Older tokens decay away exponentially, giving the reservoir a natural short-term memory window of roughly 20–30 tokens before they fade below significance.
HYDRA MaxSim Two-Stage Gate
Query-to-fact scoring uses the HYDRA MaxSim algorithm — a two-stage pipeline that minimises wasted computation: Stage 1 — Early rejection (2,000 dimensions): The first 2,000 dimensions of all query and fact hypervectors are packed into compactuint64 words. Pairwise cosine similarity is computed via Hamming distance on the packed representation. If the best score across all fact hypervectors is below τ_early = 0.20, the fact is immediately rejected without computing the full-dimensional score.
Stage 2 — Full scoring (10,000 dimensions):
Queries that survive Stage 1 are re-scored against all 10,000 dimensions. The final score is the mean of MaxSim across all query hypervector components.
A fact passes the gate if its full-dimensional HYDRA score is ≥ 0.55 (HDC_THRESHOLD) and the predicate alignment score is ≥ 0.35.
Multi-Hop Path Binding
During document ingestion, Hillock constructs multi-hop relational paths through the knowledge graph and binds each path into a single macro-hypervector using cyclic permutation:Π^k denotes k cyclic left-rotations (np.roll(..., shift=k)), * is elementwise multiplication, and each E_i, R_i is a bipolar hypervector. The anchor entity E₀ and first predicate R₁ are bound without permutation; subsequent entities and predicates receive incrementally larger shifts. The permutation operator breaks commutativity — E₀ * E₁ ≠ E₁ * E₀ — so the path direction is preserved. Up to 3-hop paths are generated within a locality window of 6 triples, then superposed into a per-document accumulator and persisted to the hdc_reservoirs table as a compact int8 BLOB.
How the Tiers Interact
Seed knowledge: Hillock ships with 10 pre-seeded entities (
France, Paris, London, UK, Marie_Curie, Poland, Radioactivity, Nobel_Prize, Alan_Turing, Enigma) and 7 seed triples covering capital cities and basic biographical facts. These are inserted via seed_initial_knowledge() at startup and are safe to query immediately — no ingestion required. They are useful for smoke-testing your installation before loading your own documents. Note that 4 of these seed triples overlap common evaluation benchmarks, which can inflate accuracy metrics; see Deterministic Gating for benchmark guidance.