Skip to main content
By the end of this guide you’ll have Hillock running locally, a document ingested into the Knowledge Graph, and the interactive CLI answering questions from your document’s extracted facts — with the HDC hallucination gate active from the first query. The entire setup runs on your machine with no cloud dependencies.
Prerequisites before you begin:
  • Python 3.10 or later — check with python --version
  • Ollama installed and running — download from ollama.com
  • A local Ollama model pulled — Hillock defaults to llama3.2. Run ollama pull llama3.2 if you haven’t already. Any Ollama-compatible model works; smaller models like phi3:mini also run well.
1

Install Hillock

The fastest way to install is via pip:
Alternatively, use the one-click launcher to clone the repo, create a virtual environment, install all dependencies, and start the console automatically:
If you cloned manually and want to install in editable mode:
2

Start Ollama

Pull the default model and make sure the Ollama server is running:
Ollama listens on http://localhost:11434 by default. Hillock connects to this endpoint automatically — no additional configuration needed unless you change the port.
You can swap to any model at any time from inside the Hillock CLI using /model phi3:mini. Smaller models run faster and use less VRAM alongside Hillock’s pipeline.
3

Launch the CLI

If you installed via pip:
If you’re running from source:
You’ll see Hillock’s banner, a list of available local Ollama models, and the interactive prompt. The Knowledge Graph is created automatically at hillock_kg.db in your working directory on first launch.
4

Ingest a document

Feed a document into Hillock’s memory using the /ingest command. Both plain text (.txt) and PDF (.pdf) files are supported:
TALON processes the document in parallel blocks (5 sentences per block, 2 sentences of overlap), extracts Subject-Predicate-Object triples using spaCy and GLiREL, resolves coreferences, and writes the resulting facts into the SQLite Knowledge Graph. For a typical 30-sentence document, this takes roughly 5 seconds.After ingestion, Hillock may launch the Active Disambiguation Quiz — see the tip below.
Active Disambiguation Quiz: After ingesting, Hillock scans the Knowledge Graph for unresolved pronoun references (e.g., facts where the subject is stored as “he” or “she”). It presents these ambiguous facts one-by-one and asks you to clarify the referent. Resolving these improves retrieval precision significantly — take a few seconds to answer each prompt.
5

Ask a question

Type any question about the content you just ingested:
If you ask something Hillock doesn’t have verified facts for, the HDC gate blocks the query and returns an honest refusal instead of a fabricated answer:
You can also use these CLI commands to explore and tune Hillock’s behaviour:

Configuration

Hillock’s key settings live in config.py. The defaults work for most setups, but you can edit them directly for your hardware and workflow:
Raising HDC_THRESHOLD above 0.55 makes the gate stricter — fewer facts pass, but false positives decrease. Lowering it below 0.55 lets more candidate facts through but increases the risk of the LLM receiving weakly matched context. The default of 0.55 was calibrated to eliminate hallucination leaks on the test suite.

Next steps

Architecture Deep Dive

Understand how the SQLite Knowledge Graph, Hebbian plasticity engine, and HDC reservoir interact at query time.

Python Library

Embed IntegratedHillock directly in your Python application for programmatic ingestion and querying.