> ## Documentation Index
> Fetch the complete documentation index at: https://hillock.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Use Hillock as an Embedded Python Library in Your App

> Import IntegratedHillock directly into your Python app to ingest documents, query the knowledge graph, and get grounded responses — no CLI required.

Hillock is designed to run embedded inside your own Python applications, not just as a standalone CLI tool. You can import `IntegratedHillock` directly, feed it documents, resolve ambiguous facts programmatically, and drive the full neuro-symbolic reasoning pipeline from your own code. This page walks through installation, the core ingestion and query loop, streaming output, and the complete public API surface.

## Installation

```bash theme={null}
pip install hillock
```

After installation, download the spaCy language model required by the TALON extraction pipeline:

```bash theme={null}
python -m spacy download en_core_web_sm
```

<Note>
  `IntegratedHillock` requires Ollama running locally at `http://localhost:11434` to execute `execute_chat_turn()`. Document ingestion and HDC gating work fully offline without Ollama — you only need it for the final LLM rendering step.
</Note>

## Basic Usage

<Steps>
  <Step title="Import and initialize">
    Instantiate `IntegratedHillock` to boot all subsystems: the SQLite knowledge graph, the Hebbian plasticity engine, and the hyperdimensional computing reservoir. The constructor seeds the KG with built-in knowledge and allocates HDC hypervectors for every registered entity.

    ```python theme={null}
    from engine import IntegratedHillock

    hillock = IntegratedHillock()
    ```

    By default, Hillock uses the database file path and Ollama model defined in `config.py`. You can override both at construction time:

    ```python theme={null}
    hillock = IntegratedHillock(db_path="my_project.db", ollama_model="mistral")
    ```
  </Step>

  <Step title="Ingest a document">
    Pass any local `.txt` or `.pdf` file through the TALON extraction pipeline. TALON extracts subject-predicate-object triples, resolves coreferences, binds multi-hop relational paths into the HDC reservoir, and persists everything to SQLite. The function returns a human-readable summary string and a `dict` of timing statistics.

    ```python theme={null}
    from ingestor import ingest_document_parallel

    summary, timing = ingest_document_parallel("path/to/document.pdf", hillock)
    print(summary)
    # ========================================================
    #         TALON ENGINE BULK INGESTION SUMMARY REPORT
    # ========================================================
    #   * File Processed             : document.pdf
    #   * Total Sentences            : 142
    #   * Extracted 1-Hop Triples    : 87
    #   * Multi-Hop Paths Bound      : 214
    #   * Pure Extraction Rate       : 12.4 sentences/sec
    #   * Total Processing Time      : 11.46 seconds
    # ========================================================

    print(timing)
    # {'load_and_first_extraction_time': 3.21, 'pure_extraction_time': 8.25,
    #  'pure_rate': 17.2, 'total_time': 11.46, 'total_sentences': 142,
    #  'extracted_triples': 87}
    ```

    The `timing` dict keys are:

    | Key | Description |
    | - | - |
    | `load_and_first_extraction_time` | Seconds to cold-start the TALON model and emit the first triple |
    | `pure_extraction_time` | Net extraction time after model warm-up |
    | `pure_rate` | Sentences processed per second during extraction |
    | `total_time` | Wall-clock time for the full ingestion call |
    | `total_sentences` | Total sentences detected in the document |
    | `extracted_triples` | Number of 1-hop SPO triples committed to the KG |
  </Step>

  <Step title="Handle disambiguation (recommended)">
    TALON's coreference resolver occasionally extracts facts where a subject or object is an unresolved pronoun (e.g., `he`, `she`, `it`). Call `get_ambiguous_facts()` after ingestion to surface these, then resolve them with `resolve_ambiguous_fact()`.

    ```python theme={null}
    ambiguous = hillock.get_ambiguous_facts()
    for s, p, o, source_doc in ambiguous:
        print(f"Ambiguous: [{s}] -[{p}]-> [{o}]  (from: {source_doc})")

    # Programmatic resolution — replace the pronoun side with the correct entity name
    # Signature: resolve_ambiguous_fact(old_s, p, old_o, new_s, new_o)
    hillock.resolve_ambiguous_fact("he", "discovered", "radioactivity", "Pierre_Curie", "radioactivity")
    ```

    Resolved facts are written back to SQLite with `confidence = 1.0` and tagged `source_doc = "human_disambiguation"`, so they are treated as high-trust anchors during HDC gating.
  </Step>

  <Step title="Ask a question">
    Call `execute_chat_turn()` with any natural-language question. Hillock links the query tokens to KG entities, gates candidate facts through the HDC similarity threshold, primes associated concepts via Hebbian weights, and streams the grounded answer through Ollama to `sys.stdout`. The full concatenated response is also returned as the first element of the tuple.

    ```python theme={null}
    response, primed_info, hdc_fingerprint, status = hillock.execute_chat_turn(
        "What did Marie Curie discover?"
    )
    # Tokens stream to stdout as they arrive from Ollama.
    # 'response' holds the complete answer once streaming finishes:
    print(f"Status: {status}")
    print(f"Full response: {response}")
    # Status: RENDER_SUCCESS
    # Full response: Hillock (Renderer) > Marie Curie discovered radioactivity...
    ```

    The return tuple carries:

    | Position | Type | Description |
    | - | - | - |
    | `response` | `str` | The rendered answer (or refusal), prefixed with `Hillock (Renderer) >` |
    | `primed_info` | `list[tuple[str, float]]` | Hebbian-associated concepts and their synaptic strengths |
    | `hdc_fingerprint` | `list[tuple[str, float]]` | Top-k HDC context traces with cosine similarity scores |
    | `status` | `str` | One of `RENDER_SUCCESS`, `RENDER_FALLBACK`, `DETERMINISTIC_GATED_FALLBACK`, `CONVERSATIONAL_REFUSAL`, `GREETING` |
  </Step>
</Steps>

## Streaming Responses

`execute_chat_turn()` drives `query_ollama_stream()` internally, which writes tokens directly to `sys.stdout` as they arrive from Ollama's SSE stream. The full concatenated response is returned as the first element of the tuple once streaming completes.

If you need to capture tokens as they stream (e.g., to pipe them to a websocket), patch `sys.stdout`:

```python theme={null}
import sys
import io
from engine import IntegratedHillock

class TokenCapture(io.TextIOBase):
    def __init__(self):
        self.tokens = []

    def write(self, s: str) -> int:
        self.tokens.append(s)
        return len(s)

    def flush(self):
        pass

hillock = IntegratedHillock()

capture = TokenCapture()
original_stdout = sys.stdout
sys.stdout = capture

response, _, _, status = hillock.execute_chat_turn("Who cracked the Enigma cipher?")

sys.stdout = original_stdout  # restore

streamed_text = "".join(capture.tokens)
print(f"Captured {len(capture.tokens)} token chunks")
print(f"Full response: {streamed_text.strip()}")
```

<Tip>
  When building a web server on top of Hillock, call `execute_chat_turn()` in a thread and stream `sys.stdout` output to the client using a queue. See `api.py` for how the built-in FastAPI server handles this pattern.
</Tip>

## API Reference

### `IntegratedHillock(db_path, ollama_model)`

Initializes all Hillock subsystems. Safe to call multiple times with different `db_path` values to maintain separate knowledge bases.

<ParamField path="db_path" type="str" default="hillock_kg.db">
  Path to the SQLite database file. Created automatically if it does not exist.
</ParamField>

<ParamField path="ollama_model" type="str" default="llama3.2">
  Name of the Ollama model to use for LLM rendering. Must be pulled locally via `ollama pull <model>`.
</ParamField>

***

### `execute_chat_turn(query)`

The main entry point for the full reasoning pipeline. Links query tokens to KG entities, gates facts through the HDC threshold, primes Hebbian associations, and streams a grounded response to `sys.stdout`.

<ParamField path="query" type="str" required>
  Natural-language question or statement. Greetings and single-word inputs are handled gracefully without triggering the KG lookup.
</ParamField>

**Returns:** `Tuple[str, List[Tuple[str, float]], List[Tuple[str, float]], str]`

***

### `get_ambiguous_facts()`

Scans the knowledge graph for stored triples where the subject or object is a pronoun (`he`, `she`, `it`, `they`, `this`, `that`, `who`, `whom`, `which`, `his`, `her`).

**Returns:** `List[Tuple[str, str, str, str]]` — list of `(subject, predicate, object, source_doc)` tuples.

***

### `resolve_ambiguous_fact(old_s, p, old_o, new_s, new_o)`

Deletes the ambiguous triple and replaces it with a human-clarified version tagged as `confidence = 1.0`.

<ParamField path="old_s" type="str" required>
  The original subject value (often a pronoun like `"he"`).
</ParamField>

<ParamField path="p" type="str" required>
  The predicate of the ambiguous triple, exactly as stored (e.g., `"discovered"`).
</ParamField>

<ParamField path="old_o" type="str" required>
  The original object value.
</ParamField>

<ParamField path="new_s" type="str" required>
  Replacement subject. Pass the same value as `old_s` if the subject was not ambiguous.
</ParamField>

<ParamField path="new_o" type="str" required>
  Replacement object. Pass the same value as `old_o` if the object was not ambiguous.
</ParamField>

***

### `ingest_document_parallel(file_path, hillock)`

Routes a `.txt` or `.pdf` document through the TALON extraction pipeline, commits triples to the knowledge graph, and binds multi-hop paths into the HDC reservoir.

<ParamField path="file_path" type="str" required>
  Absolute or relative path to the document to ingest. Supports `.txt` and `.pdf` formats.
</ParamField>

<ParamField path="hillock" type="IntegratedHillock" required>
  An initialized `IntegratedHillock` instance. The function writes extracted entities and triples directly into this instance's knowledge graph and HDC reservoir.
</ParamField>

**Returns:** `Tuple[str, Dict[str, float]]` — a human-readable summary string and a timing statistics dictionary.

***

### `link_entities(query)`

Tokenizes `query` and matches tokens against all entity IDs in the knowledge graph. Used internally by `execute_chat_turn()` to build the candidate fact set.

**Returns:** `Set[str]` — set of matched entity IDs.

***

### `resolve_entity_identity(entity_str)`

Normalizes raw entity strings: strips possessive suffixes (`'s`, `_s`), lower-cases, and resolves partial matches against registered KG entity IDs. Use this before storing or looking up entities manually.

**Returns:** `str` — canonical entity ID (underscore-delimited, lowercase).

## Configuration

Hillock's runtime behavior is controlled by constants in `config.py`. Edit them before initializing `IntegratedHillock` to change defaults across your application.

```python theme={null}
# config.py — key constants

DB_FILE = "hillock_kg.db"           # SQLite database path
OLLAMA_MODEL = "llama3.2"           # Default LLM model
LLM_BASE_URL = "http://localhost:11434/v1/chat/completions"  # Ollama endpoint

HDC_DIMENSION = 10000               # Hypervector dimensionality
HDC_DECAY = 0.95                    # Fading memory decay rate
HDC_THRESHOLD = 0.55                # Minimum HDC similarity to pass the gate

HEBBIAN_ETA = 0.15                  # Synaptic learning rate
HEBBIAN_DECAY = 0.01                # Synaptic decay rate
```

The three constants you are most likely to tune:

| Constant | Effect |
| - | - |
| `OLLAMA_MODEL` | Controls which local model renders answers. Switch to `phi3-mini` for low-VRAM machines. |
| `HDC_THRESHOLD` | Raising this (e.g., `0.65`) makes the gate stricter — fewer facts pass, answers are more conservative. Lowering it (e.g., `0.45`) allows more facts through at the cost of occasional false positives. |
| `DB_FILE` | Set to an absolute path to persist memory across working directory changes. |

You can also override model and database at the instance level without touching `config.py`:

```python theme={null}
import config
config.OLLAMA_MODEL = "mistral"
config.HDC_THRESHOLD = 0.60

from engine import IntegratedHillock
hillock = IntegratedHillock()
```

## Full Working Example

A self-contained script that ingests a local PDF and queries it in under 25 lines:

```python theme={null}
from engine import IntegratedHillock
from ingestor import ingest_document_parallel

# 1. Boot the memory engine
hillock = IntegratedHillock(db_path="research.db", ollama_model="llama3.2")

# 2. Ingest your document
pdf_path = "papers/curie_biography.pdf"
summary, timing = ingest_document_parallel(pdf_path, hillock)
print(summary)
print(f"\nExtracted {timing['extracted_triples']} triples in {timing['total_time']:.1f}s")

# 3. Resolve any ambiguous pronouns
ambiguous = hillock.get_ambiguous_facts()
print(f"\nFound {len(ambiguous)} ambiguous facts — resolving automatically...")
for s, p, o, doc in ambiguous:
    # In production, prompt the user; here we skip for brevity
    print(f"  Skipping: [{s}] -[{p}]-> [{o}]")

# 4. Query the grounded knowledge graph
questions = [
    "What did Marie Curie discover?",
    "Where was Marie Curie born?",
    "Who did Marie Curie collaborate with?",
]

print("\n--- Query Results ---")
for q in questions:
    # Tokens stream to stdout as Ollama responds.
    # 'response' holds the full answer once the stream completes.
    response, _, _, status = hillock.execute_chat_turn(q)
    print(f"\nStatus: {status}")
```


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.