.txt or .pdf — into structured Subject–Predicate–Object triples and writes them into the SQLite knowledge graph, without making a single call to a generative language model. Instead, TALON chains three purpose-built NLP models: a coreference resolver that disambiguates pronouns, a bi-encoder sentence classifier that selects relevant predicates from a 53-predicate Wikidata taxonomy, and a zero-shot span relation extractor that identifies entity pairs. This page covers each stage’s design, the ONNX CPU-only path, VRAM management, and what happens when required models are missing.
Stage 1: Coreference Resolution
Before any relation extraction happens, TALON resolves pronouns throughout the full document. This prevents triples like(she, discovered, Radioactivity) — where the subject is an unresolved pronoun — from polluting the knowledge graph.
Model: FCoref from the fastcoref library.
FCoref processes the entire document at once and returns mention clusters — groups of character-offset spans that all refer to the same entity. The first mention in each cluster is treated as the canonical head. Every subsequent mention (pronoun, short reference, abbreviated form) is replaced in-place using reverse-ordered character offset substitution, ensuring downstream sentence splits see fully resolved entity names.
Example:
fastcoref itself is not installed, the resolver returns the original text unchanged and logs a warning — extraction continues with potentially unresolved pronouns.
Stage 2: Predicate Routing
After coreference resolution, the document is split into sentences. TALON does not attempt to match every possible Wikidata predicate against every sentence — that would be 53 sequential comparisons per sentence in a loop. Instead, it uses a bi-encoder to score all 53 predicates against all sentences simultaneously in a single batched GPU matrix multiply. Model:all-MiniLM-L6-v2 SentenceTransformer (GPU path) or minilm.onnx FP16 (CPU path).
At load time, TALON pre-computes and caches sentence embeddings for all 53 predicate labels (with underscores replaced by spaces for natural phrasing). For each incoming batch of sentences, it encodes the batch in a single forward pass and computes cosine similarity against the cached predicate matrix. The top-10 scoring predicates per sentence are returned as candidate labels for Stage 3.
Complexity: O(1) in terms of architecture — all comparisons are expressed as a single matrix multiply, regardless of taxonomy size.
The Predicate Taxonomy
View all 53 Wikidata predicate labels
View all 53 Wikidata predicate labels
ONNX CPU Path
For deployments without a GPU, export the SentenceTransformer to FP16 ONNX once at setup time:minilm.onnx (~10MB, FP16 quantized). When TalonEngine initializes, it checks for minilm.onnx in the working directory. If found, it instantiates ONNXDynamicPredicateRouter instead of the PyTorch router — no code change required:
onnxruntime with CPUExecutionProvider and produces identical predicate rankings to the PyTorch version — it is not an approximation.
Stage 3: Zero-Shot Relation Extraction
With the top-10 candidate predicates in hand, TALON runs each sentence through a zero-shot span-level relation extractor. Model:jackboyla/glirel-large-v0 (GLiREL Large).
GLiREL receives the tokenized sentence, its NER spans (detected by en_core_web_sm), and the candidate predicate labels. It returns relation predictions above a confidence threshold of 0.30. TALON then applies four post-processing steps to clean the output:
After post-processing, surviving triples are passed to hillock.kg.update_relations_batch() for atomic insertion into SQLite.
VRAM Management
TALON is designed for single-pass ingestion on consumer hardware. All three models are loaded into GPU VRAM at the start of an ingestion job, and immediately unloaded afterward:unload_models() sets all model references to None, calls gc.collect(), and runs torch.cuda.empty_cache(). The singleton _talon_instance is also set to None, so the next ingestion call triggers a fresh cold start. This prevents VRAM accumulation when running multiple ingestion jobs in sequence.