> ## Documentation Index
> Fetch the complete documentation index at: https://hillock.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Configure Hillock to Run with Any Local Ollama Model

> Configure Hillock to use any Ollama model for LLM rendering — swap models at runtime, tune the endpoint, and keep every token 100% local.

Hillock uses Ollama's OpenAI-compatible API for all language model rendering. The separation is intentional: Hillock owns memory, retrieval, and HDC gating — Ollama owns text generation. Only facts that pass the HDC similarity gate ever reach the model prompt, which eliminates the hallucination surface by construction. This page explains how to configure the integration, choose the right model for your hardware, and diagnose connection issues.

<Note>
  Ollama must be installed and running at `http://localhost:11434` before you call `execute_chat_turn()`. Download Ollama from [ollama.com](https://ollama.com) and start it with `ollama serve` if it is not already running as a background service.
</Note>

## Supported Models

Hillock works with any model available in the Ollama library — it calls the standard `/v1/chat/completions` endpoint and does not depend on model-specific features. The default model is `llama3.2`.

<CardGroup cols={3}>
  <Card title="High Quality" icon="star">
    Best answers on modern hardware (8GB+ VRAM or 16GB+ RAM):

    * `llama3.2`
    * `mistral`
    * `gemma2`
  </Card>

  <Card title="Low VRAM" icon="microchip">
    Runs on 4GB VRAM or CPU-only machines:

    * `phi3-mini`
    * `qwen2.5:3b`
    * `tinyllama`
  </Card>

  <Card title="Code-Focused" icon="code">
    Technical document ingestion and code Q\&A:

    * `deepseek-coder`
    * `codellama`
    * `qwen2.5-coder`
  </Card>
</CardGroup>

Pull any model before using it:

```bash theme={null}
ollama pull llama3.2
ollama pull phi3-mini
ollama pull mistral
```

## Setting Your Model

You can switch models at three levels: CLI runtime, config file, or Python code.

### Via the CLI (runtime switch)

In an active Hillock console session, use the `/model` command to switch models immediately without restarting:

```
User > /model llama3.2
Hillock [SYSTEM]: Active Ollama model switched to [llama3.2].
```

Run `/model` with no argument to list every locally available model and see which one is active:

```
User > /model

Hillock [SYSTEM]: Active Model: [llama3.2]
 [Available Local Ollama Models on your PC]:
  * llama3.2 (Active)
  * mistral
  * phi3-mini
```

### Via `config.py` (persistent default)

Edit `OLLAMA_MODEL` in `config.py` to change the default for all future sessions and API server instances:

```python theme={null}
# config.py
OLLAMA_MODEL = "mistral"   # was "llama3.2"
```

### Via Python (per-instance override)

Override the model for a specific `IntegratedHillock` instance without touching `config.py`:

```python theme={null}
from engine import IntegratedHillock

# Pass model name at construction
hillock = IntegratedHillock(ollama_model="gemma2")

# Or switch mid-session
hillock.ollama_model = "phi3-mini"

response, _, _, _ = hillock.execute_chat_turn("What is radioactivity?")
```

You can also patch `config` before importing `IntegratedHillock` to change the default at module load time:

```python theme={null}
import config
config.OLLAMA_MODEL = "deepseek-coder"

from engine import IntegratedHillock
hillock = IntegratedHillock()  # picks up deepseek-coder
```

## How the Integration Works

Understanding the request flow helps you tune performance and debug issues.

### Request path

1. `execute_chat_turn(query)` links query tokens to KG entities and retrieves candidate facts.
2. Each candidate fact is scored against the query hypervector. Facts with HDC cosine similarity below `HDC_THRESHOLD` (default `0.55`) are dropped.
3. Passing facts and their Hebbian-primed associations are assembled into a system prompt.
4. Hillock POSTs to `LLM_BASE_URL` using the OpenAI chat completions format with `"stream": true`.
5. Tokens arrive as SSE chunks (`data: {...}`) and are written to `sys.stdout` as they stream.

### Endpoint configuration

The full endpoint URL is controlled by `LLM_BASE_URL` in `config.py`:

```python theme={null}
# config.py
LLM_BASE_URL = "http://localhost:11434/v1/chat/completions"
```

If you run Ollama on a different host or port (e.g., a remote machine on your LAN, or a Docker container), update this value:

```python theme={null}
# config.py — remote Ollama instance
LLM_BASE_URL = "http://192.168.1.50:11434/v1/chat/completions"
```

### What the LLM prompt looks like

The system prompt and user prompt are constructed by `_get_mode_prompts()` and vary by verbosity mode. In `BALANCED` mode (the default), the prompt that reaches the LLM is similar to:

```
System: You are a knowledgeable assistant. Answer the question using the
        verified facts provided. Always briefly mention the source document.

User:   Verified fact: [Marie_Curie discovered radioactivity] (Source: curie_bio.pdf)
        Related context from memory: pierre_curie (strength 0.87), polonium (strength 0.73)
        Question: What did Marie Curie discover?
```

The LLM never sees unverified facts or raw document text — only the HDC-gated, plasticity-primed subset.

### Verbosity modes

Switch the response style without changing the model:

| Mode | CLI command | Behaviour |
| - | - | - |
| `STRICT` | `/mode strict` | LLM renders a single-sentence fact. No elaboration. |
| `BALANCED` | `/mode balanced` | LLM adds one sentence of context and cites the source document. |
| `CONVERSATIONAL` | `/mode conversational` | LLM wraps facts in warm, engaging dialogue and surfaces Hebbian associations. |

In `STRICT` mode, if the HDC gate fails Hillock returns a deterministic refusal without calling Ollama at all.

## Troubleshooting

<Accordion title="Ollama not found / connection refused">
  Hillock cannot reach `http://localhost:11434`. Start Ollama:

  ```bash theme={null}
  ollama serve
  ```

  On macOS, Ollama typically runs as a menu-bar application and starts automatically after install. On Linux, enable the systemd service:

  ```bash theme={null}
  sudo systemctl enable --now ollama
  ```

  If you changed `LLM_BASE_URL` in `config.py`, verify the host and port match where Ollama is actually listening.
</Accordion>

<Accordion title="Model not found">
  Hillock sends the model name from `OLLAMA_MODEL` (or `hillock.ollama_model`) to the Ollama API. If that model has not been pulled, the request fails silently and `execute_chat_turn()` returns a fallback response.

  Fix: pull the model first.

  ```bash theme={null}
  ollama pull llama3.2
  ```

  Run `/model` in the Hillock CLI to confirm the model appears in your local list before using it.
</Accordion>

<Accordion title="Slow first response">
  The first request after Ollama starts (or after a long idle period) triggers model loading from disk into VRAM/RAM. This cold-start delay is normal and typically takes 3–15 seconds depending on model size and storage speed. Subsequent calls within the same session are served from memory and respond in under a second on modern hardware.

  To pre-warm the model before serving requests, send a dummy request:

  ```bash theme={null}
  curl http://localhost:11434/api/generate \
    -d '{"model": "llama3.2", "prompt": "hi", "stream": false}'
  ```
</Accordion>

<Accordion title="Garbled or incomplete streaming output">
  Hillock's streaming parser expects the standard OpenAI SSE format: lines prefixed with `data: ` containing JSON with `choices[0].delta.content`. Most Ollama versions after 0.1.20 emit this format correctly.

  Update Ollama to the latest version if you see garbled output:

  ```bash theme={null}
  # macOS / Linux
  curl -fsSL https://ollama.com/install.sh | sh
  ```

  You can verify Ollama's streaming format directly:

  ```bash theme={null}
  curl http://localhost:11434/v1/chat/completions \
    -H "Content-Type: application/json" \
    -d '{"model": "llama3.2", "messages": [{"role": "user", "content": "ping"}], "stream": true}'
  ```
</Accordion>

<Accordion title="LLM invents facts not in my documents">
  This indicates the HDC gate is passing facts it should not, or the verbosity mode is too permissive. Try the following:

  1. Raise `HDC_THRESHOLD` in `config.py` (e.g., from `0.55` to `0.65`).
  2. Switch to `STRICT` mode: `/mode strict` — the LLM receives only the single matched fact, no context.
  3. Run `/debug low` and re-ask the question. The `[DEBUG HDC HYDRA]` output shows MaxSim and PredAlign scores for each candidate fact, helping you identify which facts are incorrectly passing the gate.
</Accordion>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.