IntegratedHillock engine instance — the FastAPI server is a thin HTTP wrapper around the same reasoning pipeline you call directly in Python. This page covers both in full.
FastAPI Server
The HTTP server (api.py) starts a FastAPI application on port 8000 and exposes a chat completions endpoint that any OpenAI-compatible client can point to without modification.
Base URL: http://localhost:8000
Start the server:
The server binds to
0.0.0.0:8000 — it accepts connections on all interfaces. On a shared machine, firewall port 8000 if you don’t want it accessible on your local network.GET /
Health check. Returns a status object confirming the server is online.
Response
string
Always
"online" when the server is running.string
A human-readable prompt pointing you to the chat completions endpoint.
string
The running Hillock API version string.
POST /v1/chat/completions
OpenAI-compatible chat completions. Extracts the last user message from the messages array, routes it through IntegratedHillock.execute_chat_turn(), and returns an OpenAI-format response. Drop-in compatible with any client that supports the OpenAI Chat Completions API.
Request body
string
required
Model identifier string. Hillock ignores this value and uses whichever Ollama model is currently configured on the server. Pass any string —
"hillock" works fine.array
required
Array of message objects, each with
role ("user" or "assistant") and content (string). Only the last message with role: "user" is processed; earlier messages are not used for context retrieval.boolean
If
true, returns a Server-Sent Events (SSE) streaming response in OpenAI delta format. Defaults to false.number
Accepted in the request body but not forwarded to the LLM — Hillock always runs at
temperature: 0.0 for deterministic fact rendering. Defaults to 0.0.string
Unique completion ID in the format
chatcmpl-{uuid}.string
Always
"chat.completion".integer
Unix timestamp of when the response was generated.
string
Echoes back the
model string from the request.array
Array containing a single completion choice object.
integer
Always
0.string
Always
"assistant".string
The rendered response from Hillock. CLI prefixes (e.g.
Hillock (Renderer) >) are automatically stripped before this field is populated.string
Always
"stop".object
Token usage object. All fields (
prompt_tokens, completion_tokens, total_tokens) return 0 — Hillock does not track token counts.stream: true, the server returns Content-Type: text/event-stream. The full response text is sent as a single content delta, followed by a stop chunk and [DONE].
Hillock’s streaming implementation yields the complete reply text in a single delta chunk rather than token-by-token. Clients that expect incremental token streaming will receive the full response at once before the stop chunk arrives.
Example: OpenAI Python client pointed at Hillock
Python Library
ImportIntegratedHillock directly to embed Hillock into any Python application, agent loop, or evaluation harness.
IntegratedHillock(db_path, ollama_model)
Constructor. Initializes all Hillock subsystems in sequence: SQLiteKnowledgeGraph, HebbianPlasticityEngine, HyperdimensionalReservoir, and GloVe encoder. Seeds the HDC codebook with every entity currently in the knowledge graph.
string
Path to the SQLite database file. Created automatically if it doesn’t exist. Defaults to
DB_FILE from config.py ("hillock_kg.db").string
Name of the Ollama model to use for LLM rendering. Defaults to
"llama3.2".execute_chat_turn(query)
Main reasoning pipeline. Runs the full neuro-symbolic inference chain: entity linking → HDC reservoir step → knowledge graph fact retrieval → HDC gate scoring → LLM rendering. This is the method the FastAPI server calls on every request.
string
required
The user’s natural-language question or statement.
Tuple[str, List[Tuple[str, float]], List[Tuple[str, float]], str]
string
The full response string from Hillock, prefixed with
"Hillock (Renderer) > " on LLM-rendered responses or "Hillock > " on deterministic fallbacks.list[tuple[str, float]]
Top Hebbian priming activations from the plasticity engine: list of
(entity_id, synaptic_strength) tuples, ordered by strength descending. Empty list if no facts matched.list[tuple[str, float]]
Top HDC context fingerprint entries: list of
(entity_id, cosine_similarity) tuples from the reservoir’s fading memory state. Empty list if no HDC context is active.string
Execution mode tag indicating how the response was generated. One of:
"RENDER_SUCCESS", "RENDER_FALLBACK", "GREETING", "DETERMINISTIC_GATED_FALLBACK", "CONVERSATIONAL_REFUSAL".When
verbosity_mode is "STRICT" and no matching facts are found, the method returns immediately with "I do not have verified information about that." — no LLM call is made, making it the fastest and most deterministic mode.get_ambiguous_facts()
Scans the knowledge graph for stored SPO triples where the subject or object is a bare pronoun (he, she, it, they, this, that, who, whom, which, his, her). Call this after ingestion to identify facts that need human disambiguation.
List[Tuple[str, str, str, str]]
Each tuple contains (subject, predicate, object, source_doc). The ambiguous term is whichever of subject or object is a pronoun.
resolve_ambiguous_fact(old_s, p, old_o, new_s, new_o)
Resolves a pronoun ambiguity by deleting the ambiguous triple and replacing it with a human-confirmed triple at confidence = 1.0. Also allocates HDC hypervectors for any new entities introduced by the resolution.
string
required
Subject of the ambiguous fact to delete (as returned by
get_ambiguous_facts()).string
required
Predicate of the ambiguous fact to delete.
string
required
Object of the ambiguous fact to delete.
string
required
Resolved subject for the replacement fact. Will be normalized via
resolve_entity_identity() before insertion.string
required
Resolved object for the replacement fact. Will be normalized via
resolve_entity_identity() before insertion.None
Resolved facts are stored with
source_doc = "human_disambiguation" and confidence = 1.0. They are never overwritten by subsequent ingestion of the same document.link_entities(query)
Performs token-level entity linking against the knowledge graph. Tokenizes the query, then checks each token against every entity ID in the graph, matching on individual word parts (split by _). Tokens shorter than three characters are skipped.
string
required
Natural-language query string to link.
Set[str]
Set of matched entity IDs from the knowledge graph. Empty set if no tokens match any registered entity.
resolve_entity_identity(entity_str)
Normalizes an entity name for consistent knowledge graph lookup. Strips trailing possessives ('s, _s), converts spaces to underscores, lowercases, then attempts exact match → word-part match → fallback to the normalized input string.
string
required
Raw entity string from extraction output or user input.
str
The canonical entity ID from the knowledge graph if a match is found, otherwise the normalized version of the input string.
ingest_document_parallel(file_path, hillock)
Top-level ingestion function. Reads a .txt or .pdf document, runs the full TALON extraction pipeline (coreference resolution → predicate normalization → relation extraction), commits all extracted triples to the knowledge graph, updates Hebbian weights for co-occurring entities, generates and binds multi-hop hypergraph paths into the HDC reservoir, and persists a compact reservoir BLOB to SQLite.
string
required
Absolute or relative path to a
.txt or .pdf file. Relative paths are resolved from the working directory.IntegratedHillock
required
An initialized
IntegratedHillock instance. The function writes directly to hillock.kg, hillock.plasticity, and hillock.hdc.Tuple[str, Dict[str, float]]
string
Human-readable ingestion report (multi-line string) including file name, sentence count, extracted triple count, multi-hop path count, timing breakdown, and CPU/RAM utilization.
dict
Machine-readable timing and count data. Keys:
The TALON engine instance is lazy-loaded and cached as a module-level singleton on first call. Subsequent calls within the same process skip the cold-start cost. The models are unloaded and the singleton is cleared after each ingestion call to free GPU VRAM.