Ollama must be installed and running at
http://localhost:11434 before you call execute_chat_turn(). Download Ollama from ollama.com and start it with ollama serve if it is not already running as a background service.Supported Models
Hillock works with any model available in the Ollama library — it calls the standard/v1/chat/completions endpoint and does not depend on model-specific features. The default model is llama3.2.
High Quality
Best answers on modern hardware (8GB+ VRAM or 16GB+ RAM):
llama3.2mistralgemma2
Low VRAM
Runs on 4GB VRAM or CPU-only machines:
phi3-miniqwen2.5:3btinyllama
Code-Focused
Technical document ingestion and code Q&A:
deepseek-codercodellamaqwen2.5-coder
Setting Your Model
You can switch models at three levels: CLI runtime, config file, or Python code.Via the CLI (runtime switch)
In an active Hillock console session, use the/model command to switch models immediately without restarting:
/model with no argument to list every locally available model and see which one is active:
Via config.py (persistent default)
Edit OLLAMA_MODEL in config.py to change the default for all future sessions and API server instances:
Via Python (per-instance override)
Override the model for a specificIntegratedHillock instance without touching config.py:
config before importing IntegratedHillock to change the default at module load time:
How the Integration Works
Understanding the request flow helps you tune performance and debug issues.Request path
execute_chat_turn(query)links query tokens to KG entities and retrieves candidate facts.- Each candidate fact is scored against the query hypervector. Facts with HDC cosine similarity below
HDC_THRESHOLD(default0.55) are dropped. - Passing facts and their Hebbian-primed associations are assembled into a system prompt.
- Hillock POSTs to
LLM_BASE_URLusing the OpenAI chat completions format with"stream": true. - Tokens arrive as SSE chunks (
data: {...}) and are written tosys.stdoutas they stream.
Endpoint configuration
The full endpoint URL is controlled byLLM_BASE_URL in config.py:
What the LLM prompt looks like
The system prompt and user prompt are constructed by_get_mode_prompts() and vary by verbosity mode. In BALANCED mode (the default), the prompt that reaches the LLM is similar to:
Verbosity modes
Switch the response style without changing the model:
In
STRICT mode, if the HDC gate fails Hillock returns a deterministic refusal without calling Ollama at all.
Troubleshooting
Ollama not found / connection refused
Ollama not found / connection refused
Hillock cannot reach On macOS, Ollama typically runs as a menu-bar application and starts automatically after install. On Linux, enable the systemd service:If you changed
http://localhost:11434. Start Ollama:LLM_BASE_URL in config.py, verify the host and port match where Ollama is actually listening.Model not found
Model not found
Hillock sends the model name from Run
OLLAMA_MODEL (or hillock.ollama_model) to the Ollama API. If that model has not been pulled, the request fails silently and execute_chat_turn() returns a fallback response.Fix: pull the model first./model in the Hillock CLI to confirm the model appears in your local list before using it.Slow first response
Slow first response
The first request after Ollama starts (or after a long idle period) triggers model loading from disk into VRAM/RAM. This cold-start delay is normal and typically takes 3–15 seconds depending on model size and storage speed. Subsequent calls within the same session are served from memory and respond in under a second on modern hardware.To pre-warm the model before serving requests, send a dummy request:
Garbled or incomplete streaming output
Garbled or incomplete streaming output
Hillock’s streaming parser expects the standard OpenAI SSE format: lines prefixed with You can verify Ollama’s streaming format directly:
data: containing JSON with choices[0].delta.content. Most Ollama versions after 0.1.20 emit this format correctly.Update Ollama to the latest version if you see garbled output:LLM invents facts not in my documents
LLM invents facts not in my documents
This indicates the HDC gate is passing facts it should not, or the verbosity mode is too permissive. Try the following:
- Raise
HDC_THRESHOLDinconfig.py(e.g., from0.55to0.65). - Switch to
STRICTmode:/mode strict— the LLM receives only the single matched fact, no context. - Run
/debug lowand re-ask the question. The[DEBUG HDC HYDRA]output shows MaxSim and PredAlign scores for each candidate fact, helping you identify which facts are incorrectly passing the gate.