> ## Documentation Index
> Fetch the complete documentation index at: https://hillock.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# TALON Extraction Engine: From Raw Text to SPO Triples

> TALON converts raw documents into SPO triples via a 3-stage pipeline: FCoref coreference, MiniLM predicate routing, and GLiREL relation extraction.

TALON (Tensor-Accelerated Local Ontology Network) is Hillock's ingestion pipeline. It converts raw text documents — `.txt` or `.pdf` — into structured Subject–Predicate–Object triples and writes them into the SQLite knowledge graph, without making a single call to a generative language model. Instead, TALON chains three purpose-built NLP models: a coreference resolver that disambiguates pronouns, a bi-encoder sentence classifier that selects relevant predicates from a 53-predicate Wikidata taxonomy, and a zero-shot span relation extractor that identifies entity pairs. This page covers each stage's design, the ONNX CPU-only path, VRAM management, and what happens when required models are missing.

***

## Stage 1: Coreference Resolution

Before any relation extraction happens, TALON resolves pronouns throughout the full document. This prevents triples like `(she, discovered, Radioactivity)` — where the subject is an unresolved pronoun — from polluting the knowledge graph.

**Model:** `FCoref` from the `fastcoref` library.

`FCoref` processes the entire document at once and returns **mention clusters** — groups of character-offset spans that all refer to the same entity. The first mention in each cluster is treated as the canonical head. Every subsequent mention (pronoun, short reference, abbreviated form) is replaced in-place using reverse-ordered character offset substitution, ensuring downstream sentence splits see fully resolved entity names.

**Example:**

```
INPUT:
"Marie Curie was a physicist born in Warsaw.
 She discovered radioactivity and won the Nobel Prize."

AFTER Stage 1:
"Marie Curie was a physicist born in Warsaw.
 Marie Curie discovered radioactivity and won the Nobel Prize."
```

The model falls back to CPU automatically if CUDA is unavailable. If `fastcoref` itself is not installed, the resolver returns the original text unchanged and logs a warning — extraction continues with potentially unresolved pronouns.

***

## Stage 2: Predicate Routing

After coreference resolution, the document is split into sentences. TALON does **not** attempt to match every possible Wikidata predicate against every sentence — that would be 53 sequential comparisons per sentence in a loop. Instead, it uses a bi-encoder to score all 53 predicates against all sentences simultaneously in a single batched GPU matrix multiply.

**Model:** `all-MiniLM-L6-v2` SentenceTransformer (GPU path) or `minilm.onnx` FP16 (CPU path).

At load time, TALON pre-computes and caches sentence embeddings for all 53 predicate labels (with underscores replaced by spaces for natural phrasing). For each incoming batch of sentences, it encodes the batch in a single forward pass and computes cosine similarity against the cached predicate matrix. The top-10 scoring predicates per sentence are returned as candidate labels for Stage 3.

**Complexity:** O(1) in terms of architecture — all comparisons are expressed as a single matrix multiply, regardless of taxonomy size.

### The Predicate Taxonomy

<Accordion title="View all 53 Wikidata predicate labels">
  | # | Predicate | # | Predicate |
  | - | - | - | - |
  | 1 | `born_in` | 28 | `formulated` |
  | 2 | `died_in` | 29 | `proposed` |
  | 3 | `place_of_birth` | 30 | `educated_at` |
  | 4 | `place_of_death` | 31 | `studied_at` |
  | 5 | `country_of_citizenship` | 32 | `employed_by` |
  | 6 | `collaborated_with` | 33 | `worked_at` |
  | 7 | `worked_with` | 34 | `member_of` |
  | 8 | `partnered_with` | 35 | `affiliated_with` |
  | 9 | `spouse_of` | 36 | `award_received` |
  | 10 | `child_of` | 37 | `won` |
  | 11 | `parent_of` | 38 | `nominated_for` |
  | 12 | `discovered` | 39 | `capital_of` |
  | 13 | `invented` | 40 | `located_in` |
  | 14 | `co_invented` | 41 | `headquartered_in` |
  | 15 | `developed` | 42 | `field_of_work` |
  | 16 | `designed` | 43 | `subclass_of` |
  | 17 | `founded` | 44 | `part_of` |
  | 18 | `created` | 45 | `instance_of` |
  | 19 | `cracked` | 46 | `has_part` |
  | 20 | `authored` | 47 | `contains` |
  | 21 | `wrote` | 48 | `influenced_by` |
  | 22 | `published` | 49 | `student_of` |
  | 23 | `patented` | 50 | `teacher_of` |
  | 24 | `manufactured` | 51 | `successor_to` |
  | 25 | `operated_by` | 52 | `predecessor_to` |
  | 26 | `migrated_to` | 53 | `moved_to` |
  | 27 | `resided_in` | — | — |
</Accordion>

### ONNX CPU Path

For deployments without a GPU, export the SentenceTransformer to FP16 ONNX once at setup time:

```bash theme={null}
python export_to_onnx.py
```

This produces `minilm.onnx` (\~10MB, FP16 quantized). When `TalonEngine` initializes, it checks for `minilm.onnx` in the working directory. If found, it instantiates `ONNXDynamicPredicateRouter` instead of the PyTorch router — no code change required:

```python theme={null}
# TalonEngine auto-selects at init:
if os.path.isfile("minilm.onnx"):
    self.router = ONNXDynamicPredicateRouter(model_path="minilm.onnx")
else:
    self.router = DynamicPredicateRouter(device=device)
```

The ONNX router uses `onnxruntime` with `CPUExecutionProvider` and produces identical predicate rankings to the PyTorch version — it is not an approximation.

***

## Stage 3: Zero-Shot Relation Extraction

With the top-10 candidate predicates in hand, TALON runs each sentence through a zero-shot span-level relation extractor.

**Model:** `jackboyla/glirel-large-v0` (GLiREL Large).

GLiREL receives the tokenized sentence, its NER spans (detected by `en_core_web_sm`), and the candidate predicate labels. It returns relation predictions above a confidence threshold of `0.30`. TALON then applies four post-processing steps to clean the output:

<Steps>
  ### Directionality Auto-Correction

  For **origin predicates** (`born_in`, `died_in`, `migrated_to`, etc.), if the head entity is a location and the tail is a person, the pair is swapped automatically. For **creation predicates** (`discovered`, `invented`, `designed`, etc.), if the head is an inanimate object and the tail is a person, the pair is swapped. This corrects common GLiREL span ordering errors without requiring a re-run.

  ### Creation Predicate Fix

  A hardcoded set of inanimate stopwords (`radioactivity`, `enigma machine`, `compiler`, `algebra`, etc.) prevents abstract concepts from appearing as the subject of action predicates. Triples like `(Radioactivity, discovered, Marie_Curie)` are corrected to `(Marie_Curie, discovered, Radioactivity)`.

  ### Inverted Pair Purging

  For non-symmetric predicates, if both `(A, P, B)` and `(B, P, A)` are extracted from the same document, the second occurrence is dropped. Only the first seen ordering is kept, preventing contradictory directional triples in the knowledge graph.

  ### Canonical Deduplication

  For **symmetric predicates** (`collaborated_with`, `worked_with`, `partnered_with`, `spouse_of`), order is normalized to `(min(A, B), P, max(A, B))` before inserting into the seen-keys set. This ensures `(Curie, collaborated_with, Einstein)` and `(Einstein, collaborated_with, Curie)` collapse into a single canonical triple.
</Steps>

After post-processing, surviving triples are passed to `hillock.kg.update_relations_batch()` for atomic insertion into SQLite.

***

## VRAM Management

TALON is designed for single-pass ingestion on consumer hardware. All three models are loaded into GPU VRAM at the start of an ingestion job, and **immediately unloaded afterward**:

```python theme={null}
# Called automatically at end of ingest_document_parallel():
talon.unload_models()
```

`unload_models()` sets all model references to `None`, calls `gc.collect()`, and runs `torch.cuda.empty_cache()`. The singleton `_talon_instance` is also set to `None`, so the next ingestion call triggers a fresh cold start. This prevents VRAM accumulation when running multiple ingestion jobs in sequence.

### Ingestion Timing Report

After each ingestion, Hillock prints a full timing breakdown:

```
========================================================
        TALON ENGINE BULK INGESTION SUMMARY REPORT
========================================================
  * File Processed             : my_document.txt
  * Total Sentences            : 84
  * Extracted 1-Hop Triples    : 31
  * Multi-Hop Paths Bound      : 47
  * Model Load & Cold-Start    : 12.43 seconds
  * Pure Extraction Duration   : 6.81 seconds
  * Pure Extraction Rate       : 12.3 sentences/sec
  * Total Processing Time      : 19.24 seconds
========================================================
```

**Cold-start time** includes model downloads (first run only), VRAM allocation, and taxonomy embedding pre-computation. **Pure extraction rate** measures only the sentence→triple throughput after models are warm.

***

<Warning>
  **TALON stack is required for ingestion.** If any of the required packages (`glirel`, `fastcoref`, `spacy`, `torch`) are missing or fail to initialize, ingestion halts immediately and prints a loud warning banner — it does **not** silently continue with empty output. Install all dependencies with:

  ```bash theme={null}
  pip install -r requirements.txt
  python -m spacy download en_core_web_sm
  ```

  The TALON stack is only required at ingestion time. The query/retrieval path (`IntegratedHillock`) has no NLP model dependencies and runs entirely on SQLite + NumPy.
</Warning>

<Tip>
  **Edge and CPU deployments:** Run `python export_to_onnx.py` once during setup to generate `minilm.onnx`. This eliminates the PyTorch GPU dependency for predicate routing entirely. The FCoref and GLiREL models still require PyTorch, but predicate routing is typically the highest-throughput bottleneck — moving it to ONNX can improve CPU extraction rates by 2–3× on multi-core machines.
</Tip>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.