> ## Documentation Index
> Fetch the complete documentation index at: https://hillock.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Hillock API Reference: HTTP Endpoints and Python SDK

> Full reference for Hillock's FastAPI HTTP endpoints, IntegratedHillock Python class methods, and the ingest_document_parallel ingestion function.

Hillock exposes two interfaces: an OpenAI-compatible HTTP API served by FastAPI, and a Python library you can import directly into any agent framework or script. Both interfaces share the same underlying `IntegratedHillock` engine instance — the FastAPI server is a thin HTTP wrapper around the same reasoning pipeline you call directly in Python. This page covers both in full.

***

## FastAPI Server

The HTTP server (`api.py`) starts a FastAPI application on port `8000` and exposes a chat completions endpoint that any OpenAI-compatible client can point to without modification.

**Base URL:** `http://localhost:8000`

**Start the server:**

```bash theme={null}
python api.py
```

```
🚀 Starting Hillock OpenAI-Compatible API Server on http://0.0.0.0:8000
```

<Note>
  The server binds to `0.0.0.0:8000` — it accepts connections on all interfaces. On a shared machine, firewall port 8000 if you don't want it accessible on your local network.
</Note>

***

### `GET /`

Health check. Returns a status object confirming the server is online.

**Response**

```json theme={null}
{
  "status": "online",
  "message": "Hillock API is running! Point your AI client to /v1/chat/completions",
  "version": "0.6.1"
}
```

<ResponseField name="status" type="string">
  Always `"online"` when the server is running.
</ResponseField>

<ResponseField name="message" type="string">
  A human-readable prompt pointing you to the chat completions endpoint.
</ResponseField>

<ResponseField name="version" type="string">
  The running Hillock API version string.
</ResponseField>

***

### `POST /v1/chat/completions`

OpenAI-compatible chat completions. Extracts the last user message from the `messages` array, routes it through `IntegratedHillock.execute_chat_turn()`, and returns an OpenAI-format response. Drop-in compatible with any client that supports the OpenAI Chat Completions API.

**Request body**

<ParamField body="model" type="string" required>
  Model identifier string. Hillock ignores this value and uses whichever Ollama model is currently configured on the server. Pass any string — `"hillock"` works fine.
</ParamField>

<ParamField body="messages" type="array" required>
  Array of message objects, each with `role` (`"user"` or `"assistant"`) and `content` (string). Only the last message with `role: "user"` is processed; earlier messages are not used for context retrieval.
</ParamField>

<ParamField body="stream" type="boolean">
  If `true`, returns a Server-Sent Events (SSE) streaming response in OpenAI delta format. Defaults to `false`.
</ParamField>

<ParamField body="temperature" type="number">
  Accepted in the request body but not forwarded to the LLM — Hillock always runs at `temperature: 0.0` for deterministic fact rendering. Defaults to `0.0`.
</ParamField>

**Request example**

```json theme={null}
{
  "model": "hillock",
  "messages": [
    {"role": "user", "content": "What did Marie Curie discover?"}
  ],
  "stream": false
}
```

**Response (non-streaming)**

<ResponseField name="id" type="string">
  Unique completion ID in the format `chatcmpl-{uuid}`.
</ResponseField>

<ResponseField name="object" type="string">
  Always `"chat.completion"`.
</ResponseField>

<ResponseField name="created" type="integer">
  Unix timestamp of when the response was generated.
</ResponseField>

<ResponseField name="model" type="string">
  Echoes back the `model` string from the request.
</ResponseField>

<ResponseField name="choices" type="array">
  Array containing a single completion choice object.
</ResponseField>

<ResponseField name="choices[0].index" type="integer">
  Always `0`.
</ResponseField>

<ResponseField name="choices[0].message.role" type="string">
  Always `"assistant"`.
</ResponseField>

<ResponseField name="choices[0].message.content" type="string">
  The rendered response from Hillock. CLI prefixes (e.g. `Hillock (Renderer) >`) are automatically stripped before this field is populated.
</ResponseField>

<ResponseField name="choices[0].finish_reason" type="string">
  Always `"stop"`.
</ResponseField>

<ResponseField name="usage" type="object">
  Token usage object. All fields (`prompt_tokens`, `completion_tokens`, `total_tokens`) return `0` — Hillock does not track token counts.
</ResponseField>

```json theme={null}
{
  "id": "chatcmpl-3f7a2b1e9d4c6a0f8b5e2d1c7a4f9b3e",
  "object": "chat.completion",
  "created": 1720000000,
  "model": "hillock",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "Marie Curie discovered Radioactivity."
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 0,
    "completion_tokens": 0,
    "total_tokens": 0
  }
}
```

**Streaming response**

When `stream: true`, the server returns `Content-Type: text/event-stream`. The full response text is sent as a single content delta, followed by a stop chunk and `[DONE]`.

```
data: {"id":"chatcmpl-...","object":"chat.completion.chunk","created":1720000000,"model":"hillock","choices":[{"index":0,"delta":{"content":"Marie Curie discovered Radioactivity."},"finish_reason":null}]}

data: {"id":"chatcmpl-...","object":"chat.completion.chunk","created":1720000000,"model":"hillock","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}

data: [DONE]
```

<Note>
  Hillock's streaming implementation yields the complete reply text in a single delta chunk rather than token-by-token. Clients that expect incremental token streaming will receive the full response at once before the stop chunk arrives.
</Note>

**Error responses**

| Status | Cause |
| - | - |
| `422 Unprocessable Entity` | Request body failed Pydantic validation (missing required fields, wrong types). |
| `500 Internal Server Error` | Ollama is unreachable, the configured model is not pulled, or an unhandled exception occurred in the engine. |

**Example: OpenAI Python client pointed at Hillock**

```python theme={null}
from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:8000/v1",
    api_key="not-needed",  # Hillock doesn't require authentication
)

response = client.chat.completions.create(
    model="hillock",
    messages=[{"role": "user", "content": "What did Marie Curie discover?"}],
)

print(response.choices[0].message.content)
# Marie Curie discovered Radioactivity.
```

***

## Python Library

Import `IntegratedHillock` directly to embed Hillock into any Python application, agent loop, or evaluation harness.

```python theme={null}
from engine import IntegratedHillock
from ingestor import ingest_document_parallel
from config import DB_FILE

hillock = IntegratedHillock(DB_FILE)
```

***

### `IntegratedHillock(db_path, ollama_model)`

Constructor. Initializes all Hillock subsystems in sequence: `SQLiteKnowledgeGraph`, `HebbianPlasticityEngine`, `HyperdimensionalReservoir`, and GloVe encoder. Seeds the HDC codebook with every entity currently in the knowledge graph.

<ParamField path="db_path" type="string">
  Path to the SQLite database file. Created automatically if it doesn't exist. Defaults to `DB_FILE` from `config.py` (`"hillock_kg.db"`).
</ParamField>

<ParamField path="ollama_model" type="string">
  Name of the Ollama model to use for LLM rendering. Defaults to `"llama3.2"`.
</ParamField>

```python theme={null}
from engine import IntegratedHillock

# Default configuration
hillock = IntegratedHillock("hillock_kg.db")

# Custom model and database path
hillock = IntegratedHillock(
    db_path="/data/my_project.db",
    ollama_model="mistral",
)
```

**Instance attributes you can read and set at runtime:**

| Attribute | Type | Default | Description |
| - | - | - | - |
| `ollama_model` | `str` | `"llama3.2"` | Active Ollama model. Assignable at any time. |
| `verbosity_mode` | `str` | `"BALANCED"` | Response personality: `"STRICT"`, `"BALANCED"`, or `"CONVERSATIONAL"`. |
| `debug_level` | `str` | `"OFF"` | Trace verbosity: `"OFF"`, `"LOW"`, or `"FULL"`. |
| `kg` | `SQLiteKnowledgeGraph` | — | Direct handle to the knowledge graph. |
| `plasticity` | `HebbianPlasticityEngine` | — | Direct handle to the Hebbian plasticity engine. |
| `hdc` | `HyperdimensionalReservoir` | — | Direct handle to the HDC reservoir. |

***

### `execute_chat_turn(query)`

Main reasoning pipeline. Runs the full neuro-symbolic inference chain: entity linking → HDC reservoir step → knowledge graph fact retrieval → HDC gate scoring → LLM rendering. This is the method the FastAPI server calls on every request.

```python theme={null}
reply, primed, fingerprint, mode = hillock.execute_chat_turn(
    "What did Marie Curie discover?"
)
print(reply)
# Hillock (Renderer) > Marie Curie discovered Radioactivity.
```

<ParamField path="query" type="string" required>
  The user's natural-language question or statement.
</ParamField>

**Returns:** `Tuple[str, List[Tuple[str, float]], List[Tuple[str, float]], str]`

<ResponseField name="reply" type="string">
  The full response string from Hillock, prefixed with `"Hillock (Renderer) > "` on LLM-rendered responses or `"Hillock > "` on deterministic fallbacks.
</ResponseField>

<ResponseField name="primed" type="list[tuple[str, float]]">
  Top Hebbian priming activations from the plasticity engine: list of `(entity_id, synaptic_strength)` tuples, ordered by strength descending. Empty list if no facts matched.
</ResponseField>

<ResponseField name="fingerprint" type="list[tuple[str, float]]">
  Top HDC context fingerprint entries: list of `(entity_id, cosine_similarity)` tuples from the reservoir's fading memory state. Empty list if no HDC context is active.
</ResponseField>

<ResponseField name="mode" type="string">
  Execution mode tag indicating how the response was generated. One of: `"RENDER_SUCCESS"`, `"RENDER_FALLBACK"`, `"GREETING"`, `"DETERMINISTIC_GATED_FALLBACK"`, `"CONVERSATIONAL_REFUSAL"`.
</ResponseField>

<Note>
  When `verbosity_mode` is `"STRICT"` and no matching facts are found, the method returns immediately with `"I do not have verified information about that."` — no LLM call is made, making it the fastest and most deterministic mode.
</Note>

***

### `get_ambiguous_facts()`

Scans the knowledge graph for stored SPO triples where the subject or object is a bare pronoun (`he`, `she`, `it`, `they`, `this`, `that`, `who`, `whom`, `which`, `his`, `her`). Call this after ingestion to identify facts that need human disambiguation.

```python theme={null}
ambiguous = hillock.get_ambiguous_facts()
for subject, predicate, obj, source_doc in ambiguous:
    print(f"[{subject}] -[{predicate}]-> [{obj}]  (source: {source_doc})")
# [she] -[discovered]-> [Radioactivity]  (source: curie_biography.pdf)
```

**Returns:** `List[Tuple[str, str, str, str]]`

Each tuple contains `(subject, predicate, object, source_doc)`. The ambiguous term is whichever of `subject` or `object` is a pronoun.

***

### `resolve_ambiguous_fact(old_s, p, old_o, new_s, new_o)`

Resolves a pronoun ambiguity by deleting the ambiguous triple and replacing it with a human-confirmed triple at `confidence = 1.0`. Also allocates HDC hypervectors for any new entities introduced by the resolution.

```python theme={null}
# Resolve [she] -[discovered]-> [Radioactivity]
hillock.resolve_ambiguous_fact(
    old_s="she",
    p="discovered",
    old_o="Radioactivity",
    new_s="Marie_Curie",
    new_o="Radioactivity",
)
```

<ParamField path="old_s" type="string" required>
  Subject of the ambiguous fact to delete (as returned by `get_ambiguous_facts()`).
</ParamField>

<ParamField path="p" type="string" required>
  Predicate of the ambiguous fact to delete.
</ParamField>

<ParamField path="old_o" type="string" required>
  Object of the ambiguous fact to delete.
</ParamField>

<ParamField path="new_s" type="string" required>
  Resolved subject for the replacement fact. Will be normalized via `resolve_entity_identity()` before insertion.
</ParamField>

<ParamField path="new_o" type="string" required>
  Resolved object for the replacement fact. Will be normalized via `resolve_entity_identity()` before insertion.
</ParamField>

**Returns:** `None`

<Note>
  Resolved facts are stored with `source_doc = "human_disambiguation"` and `confidence = 1.0`. They are never overwritten by subsequent ingestion of the same document.
</Note>

***

### `link_entities(query)`

Performs token-level entity linking against the knowledge graph. Tokenizes the query, then checks each token against every entity ID in the graph, matching on individual word parts (split by `_`). Tokens shorter than three characters are skipped.

```python theme={null}
entities = hillock.link_entities("What did Marie Curie discover?")
print(entities)
# {'Marie_Curie'}
```

<ParamField path="query" type="string" required>
  Natural-language query string to link.
</ParamField>

**Returns:** `Set[str]`

Set of matched entity IDs from the knowledge graph. Empty set if no tokens match any registered entity.

***

### `resolve_entity_identity(entity_str)`

Normalizes an entity name for consistent knowledge graph lookup. Strips trailing possessives (`'s`, `_s`), converts spaces to underscores, lowercases, then attempts exact match → word-part match → fallback to the normalized input string.

```python theme={null}
print(hillock.resolve_entity_identity("Marie Curie's"))
# Marie_Curie

print(hillock.resolve_entity_identity("curie"))
# Marie_Curie  (matched via word-part "curie")

print(hillock.resolve_entity_identity("unknown_entity"))
# unknown_entity  (fallback: returned as-is)
```

<ParamField path="entity_str" type="string" required>
  Raw entity string from extraction output or user input.
</ParamField>

**Returns:** `str`

The canonical entity ID from the knowledge graph if a match is found, otherwise the normalized version of the input string.

***

## `ingest_document_parallel(file_path, hillock)`

Top-level ingestion function. Reads a `.txt` or `.pdf` document, runs the full TALON extraction pipeline (coreference resolution → predicate normalization → relation extraction), commits all extracted triples to the knowledge graph, updates Hebbian weights for co-occurring entities, generates and binds multi-hop hypergraph paths into the HDC reservoir, and persists a compact reservoir BLOB to SQLite.

```python theme={null}
from ingestor import ingest_document_parallel
from engine import IntegratedHillock
from config import DB_FILE

hillock = IntegratedHillock(DB_FILE)

summary, stats = ingest_document_parallel(
    "./research_papers/curie_biography.pdf",
    hillock,
)

print(summary)
print(f"Extracted {stats['extracted_triples']} triples in {stats['total_time']:.2f}s")
```

<ParamField path="file_path" type="string" required>
  Absolute or relative path to a `.txt` or `.pdf` file. Relative paths are resolved from the working directory.
</ParamField>

<ParamField path="hillock" type="IntegratedHillock" required>
  An initialized `IntegratedHillock` instance. The function writes directly to `hillock.kg`, `hillock.plasticity`, and `hillock.hdc`.
</ParamField>

**Returns:** `Tuple[str, Dict[str, float]]`

<ResponseField name="summary" type="string">
  Human-readable ingestion report (multi-line string) including file name, sentence count, extracted triple count, multi-hop path count, timing breakdown, and CPU/RAM utilization.
</ResponseField>

<ResponseField name="timing_stats" type="dict">
  Machine-readable timing and count data. Keys:

  | Key | Type | Description |
  | - | - | - |
  | `load_and_first_extraction_time` | `float` | Seconds from function entry to first extracted triple (model cold-start cost). |
  | `pure_extraction_time` | `float` | Seconds from first to last extracted triple (pure TALON throughput window). |
  | `pure_rate` | `float` | Sentences per second during the pure extraction window. |
  | `total_time` | `float` | Total wall-clock seconds for the entire ingestion call. |
  | `total_sentences` | `int` | Number of sentences detected in the source document. |
  | `extracted_triples` | `int` | Number of 1-hop SPO triples extracted and stored. |
</ResponseField>

<Warning>
  `ingest_document_parallel` requires the full TALON dependency stack: `glirel`, `fastcoref`, `spacy` (`en_core_web_sm`), and `torch`. If any are missing, the function prints a warning and returns an empty stats dict without modifying the knowledge graph. Run `pip install -r requirements.txt && python -m spacy download en_core_web_sm` to install everything.
</Warning>

<Note>
  The TALON engine instance is lazy-loaded and cached as a module-level singleton on first call. Subsequent calls within the same process skip the cold-start cost. The models are unloaded and the singleton is cleared after each ingestion call to free GPU VRAM.
</Note>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.