Skip to main content
Hillock is a memory engine for AI agents that solves a problem every local LLM developer eventually hits: retrieval-augmented generation that still hallucinates, burns through VRAM, and falls apart the moment your document topics get specific. Instead of dense vector embeddings and probabilistic nearest-neighbor lookups, Hillock extracts structured Subject-Predicate-Object triples from your documents, stores them in a relational SQLite Knowledge Graph, and applies a hard mathematical gate using Hyperdimensional Computing before any query reaches the language model. If the answer isn’t in the graph, Hillock says so — it never fills the gap with fabrication.

Key benefits

Zero Hallucinations

The HDC cosine-similarity gate operates at threshold 0.55. Queries that don’t match stored hypervectors are blocked and returned as honest refusals — never passed to the LLM for creative gap-filling.

Under 1.2 GB VRAM

The full three-tier pipeline fits in less than 1.2 GB of VRAM, comfortably within GTX 1070-class hardware. CPU-only mode requires no GPU at all — useful for Raspberry Pi and edge deployments.

100% Offline & Private

No API keys, no cloud calls, no telemetry. Hillock talks only to your local Ollama instance over localhost. Your documents never leave your machine.

Fast Tensor Ingestion

TALON (Tensor-Accelerated Local Ontology Network) uses tensor-based O(1) matrix classification to extract facts — no LLM pass during ingestion. A 30-sentence document ingests in roughly 5 seconds.

OpenAI-Compatible API

Hillock exposes an OpenAI-compatible REST server on http://localhost:8000/v1. Any tool that speaks the OpenAI chat completions format — Open-WebUI, AnythingLLM, Obsidian — works out of the box.

Three Answering Modes

Switch between STRICT (single-sentence facts only), BALANCED (fact + light context + source citation), and CONVERSATIONAL (warm, associative responses using Hebbian memory links).

How it works

Hillock’s pipeline has three stages that happen every time you ingest a document or ask a question:
  • Ingest: TALON parses your document with spaCy and GLiREL, resolves coreferences via FastCoref, and writes normalized SPO triples into the SQLite Knowledge Graph. No LLM is involved in this step.
  • Store: The Knowledge Graph holds entity nodes and typed relation edges. The Hebbian Plasticity Engine strengthens association weights between co-activated entities over repeated queries, surfacing contextually relevant priming.
  • Retrieve with HDC gate: At query time, the Hyperdimensional Reservoir encodes the query into a 10,000-dimensional binary hypervector and computes cosine similarity against stored fact hypervectors. Only facts that clear the 0.55 threshold are forwarded to the LLM renderer — everything else triggers a refusal.

Who is Hillock for?

Hillock is purpose-built for four groups of engineers and researchers: AI engineers building agents that need auditable, deterministic memory over private document sets — contracts, internal wikis, research papers — where hallucination is a business risk, not just an annoyance. Local LLM developers running Ollama, LM Studio, or llama.cpp on consumer GPUs who’ve outgrown vector databases and need something that fits inside their existing VRAM budget alongside the main language model. Privacy-focused researchers working with sensitive data — medical records, legal documents, proprietary datasets — where sending document content to any cloud endpoint is a non-starter. Edge hardware users targeting Raspberry Pi, Jetson Nano, or low-power x86 hardware, where Hillock’s CPU-only mode and minimal RAM footprint make deployment practical without a dedicated GPU.

Get Started in 5 Minutes

Install Hillock, pull llama3.2 from Ollama, and run your first query from the CLI.

Why Not Traditional RAG?

See a side-by-side comparison of Hillock vs. Chroma, Pinecone, and standard RAG pipelines.