Skip to main content
Hillock uses Ollama’s OpenAI-compatible API for all language model rendering. The separation is intentional: Hillock owns memory, retrieval, and HDC gating — Ollama owns text generation. Only facts that pass the HDC similarity gate ever reach the model prompt, which eliminates the hallucination surface by construction. This page explains how to configure the integration, choose the right model for your hardware, and diagnose connection issues.
Ollama must be installed and running at http://localhost:11434 before you call execute_chat_turn(). Download Ollama from ollama.com and start it with ollama serve if it is not already running as a background service.

Supported Models

Hillock works with any model available in the Ollama library — it calls the standard /v1/chat/completions endpoint and does not depend on model-specific features. The default model is llama3.2.

High Quality

Best answers on modern hardware (8GB+ VRAM or 16GB+ RAM):
  • llama3.2
  • mistral
  • gemma2

Low VRAM

Runs on 4GB VRAM or CPU-only machines:
  • phi3-mini
  • qwen2.5:3b
  • tinyllama

Code-Focused

Technical document ingestion and code Q&A:
  • deepseek-coder
  • codellama
  • qwen2.5-coder
Pull any model before using it:

Setting Your Model

You can switch models at three levels: CLI runtime, config file, or Python code.

Via the CLI (runtime switch)

In an active Hillock console session, use the /model command to switch models immediately without restarting:
Run /model with no argument to list every locally available model and see which one is active:

Via config.py (persistent default)

Edit OLLAMA_MODEL in config.py to change the default for all future sessions and API server instances:

Via Python (per-instance override)

Override the model for a specific IntegratedHillock instance without touching config.py:
You can also patch config before importing IntegratedHillock to change the default at module load time:

How the Integration Works

Understanding the request flow helps you tune performance and debug issues.

Request path

  1. execute_chat_turn(query) links query tokens to KG entities and retrieves candidate facts.
  2. Each candidate fact is scored against the query hypervector. Facts with HDC cosine similarity below HDC_THRESHOLD (default 0.55) are dropped.
  3. Passing facts and their Hebbian-primed associations are assembled into a system prompt.
  4. Hillock POSTs to LLM_BASE_URL using the OpenAI chat completions format with "stream": true.
  5. Tokens arrive as SSE chunks (data: {...}) and are written to sys.stdout as they stream.

Endpoint configuration

The full endpoint URL is controlled by LLM_BASE_URL in config.py:
If you run Ollama on a different host or port (e.g., a remote machine on your LAN, or a Docker container), update this value:

What the LLM prompt looks like

The system prompt and user prompt are constructed by _get_mode_prompts() and vary by verbosity mode. In BALANCED mode (the default), the prompt that reaches the LLM is similar to:
The LLM never sees unverified facts or raw document text — only the HDC-gated, plasticity-primed subset.

Verbosity modes

Switch the response style without changing the model: In STRICT mode, if the HDC gate fails Hillock returns a deterministic refusal without calling Ollama at all.

Troubleshooting

Hillock cannot reach http://localhost:11434. Start Ollama:
On macOS, Ollama typically runs as a menu-bar application and starts automatically after install. On Linux, enable the systemd service:
If you changed LLM_BASE_URL in config.py, verify the host and port match where Ollama is actually listening.
Hillock sends the model name from OLLAMA_MODEL (or hillock.ollama_model) to the Ollama API. If that model has not been pulled, the request fails silently and execute_chat_turn() returns a fallback response.Fix: pull the model first.
Run /model in the Hillock CLI to confirm the model appears in your local list before using it.
The first request after Ollama starts (or after a long idle period) triggers model loading from disk into VRAM/RAM. This cold-start delay is normal and typically takes 3–15 seconds depending on model size and storage speed. Subsequent calls within the same session are served from memory and respond in under a second on modern hardware.To pre-warm the model before serving requests, send a dummy request:
Hillock’s streaming parser expects the standard OpenAI SSE format: lines prefixed with data: containing JSON with choices[0].delta.content. Most Ollama versions after 0.1.20 emit this format correctly.Update Ollama to the latest version if you see garbled output:
You can verify Ollama’s streaming format directly:
This indicates the HDC gate is passing facts it should not, or the verbosity mode is too permissive. Try the following:
  1. Raise HDC_THRESHOLD in config.py (e.g., from 0.55 to 0.65).
  2. Switch to STRICT mode: /mode strict — the LLM receives only the single matched fact, no context.
  3. Run /debug low and re-ask the question. The [DEBUG HDC HYDRA] output shows MaxSim and PredAlign scores for each candidate fact, helping you identify which facts are incorrectly passing the gate.