Knowledge graphs and automated content pipelines
- Question
- How do you build a large body of published text in which every sentence can be traced to a source, using language models that are not trusted?
- Status
- Running pipelines for more than a hundred towns, at the scale of hundreds of thousands of sourced claims. A person approves everything that is published.
A layered knowledge graph for place-based knowledge. Every generated sentence points back through claims to a revision-pinned source sentence. Automated pipelines build the graph and turn it into narrated content.
- Source
- Sentence-level evidence
- Atomic claim
- Entity
- Dossier
- Hook and block
- Public text
What I built
- Stable, content-derived claim ids that merge identical facts across sources, and revision-pinned citations on every generated sentence.
- Explicit uncertainty: claims keep the source's modality (certain, attributed, legend, disputed) and date precision as intervals. Conflicts between sources are typed and retained.
- Constrained language-model steps: the model sees numbered sentences and candidate ids and answers through a JSON schema with ids and indices only. Deterministic code verifies every number, name, hedge and citation before anything is accepted.
- Lower-cost automated checks (entailment cascades, name-span extractors, rules) that replace model judges, run in shadow mode first, and switch themselves off when audit disagreement exceeds a bound.
- Resumable, content-addressed pipelines: immutable Parquet builds with recorded inputs, staleness derived from lineage instead of timestamps, and "could not check" as a first-class result that is never read as "none".
- Postgres with PostGIS and pgvector, DuckDB over the columnar builds, and binary-quantised embeddings with exact Hamming search, chosen after measuring recall against storage.
- Source licences and freshness treated as data: per-source terms, a gate in front of every fetch, and durability classes for facts that expire.
- A test that fails if any place, language or script literal appears in the code, so that the system generalises by rule.
What I found
- A language-model judge disagreed with itself on about a fifth of items and caught only about half of the unsupported sentences, which is why the judges were replaced by measured checks.
- Most gold sets were labelled by model readers and are marked as not yet checked by a person. I do not report accuracy on them.