Utopia: A Bitemporal Knowledge Graph in Rust That Treats Time as the Substrate — aniketkarneai.com | aniketkarneai.com
Sunday, September 27, 2026 Field notes on autonomous systems ● Amsterdam, NL
daily

Utopia: A Bitemporal Knowledge Graph in Rust That Treats Time as the Substrate

DeepLethe released Utopia v0.1 in five weeks — a Rust+Postgres enterprise world model where every fact carries two clocks, the ontology compiles into axioms, and relations get signature-checked before they touch the graph. Worth a long look because it answers the question most RAG stacks duck: what was true, and when did the system come to believe it.

The first thing the Utopia README asks you not to do is compare it to Palantir. “We would rather this project were not framed as an open-source take on Palantir,” it says. “It is a different route to enterprise intelligence, built bottom up from knowledge governance to trustworthy decisions and simulation.” Then it goes on to ship something most enterprise-data companies cannot build after a year of headcount: an Apache-2.0 Rust workspace of eight crates that runs against a single Postgres, with a bitemporal knowledge graph at its core, an ontology that compiles into axioms, a reasoning engine that derives new facts by forward chaining, an MCP server for the agent read path, and a sibling project that posts state-of-the-art numbers on BIRD Mini-Dev for SQLite and PostgreSQL. The headline framing — “the world’s first open-source enterprise world model” — is the kind of claim I would normally read past. After a few hours in the pipeline document I read it twice.

The reason is the design decisions. Utopia ships twenty of them, numbered 0001 through 0020, in docs/decisions/, and they are written like engineering memos rather than release notes. 0012-the-ontology-is-a-contract-not-a-suggestion.md documents the moment extraction wrote Musk employee Microsoft because English “X is an employee of Y” is too strong to suppress with prompt tuning, and the team gave up on the prompt and made the ontology enforce the signature. The ontology declares employee (organization → person). The model still wrote the wrong subject. So the write path corrects it: if the subject violates the domain and the object fits, swap them and record direction_corrected, never silently. If the swap is also invalid — the same model wrote OpenAI affectedBy … against a schema.org medical-test predicate — drop the predicate and keep subject, object, and evidence. Argument order is an encoding convention of the key, not a claim about the world, so here the ontology enforces.

That is a single paragraph of a single design decision. There are twenty. The project has been public for five weeks (the repository was created on 2026-08-07 and the most recent release candidate, v0.1.0-rc5, was tagged on 2026-09-05; in between there were five release candidates in ten days, and the dev branch is still moving). It has 6,760 stars, 582 forks, 206 watchers, 31 open issues, and a footprint I want to walk through because the trade-offs are unusually legible for a project this young.

What “world model” actually means here

The README names the project after a romantic gesture — Ptolemy’s geocentric model was taken for truth for a very long time, then falsified step by step by Copernicus, Kepler, Galileo, Newton. “What we keep is not only that heliocentrism turned out to be right; it is how that history unfolded.” That sentence is doing real work. Most enterprise knowledge systems model present knowledge and forget that any specific assertion was wrong, correct, or revised. A vector store stores embeddings; a knowledge graph stores edges. Neither stores the history of how the system’s understanding of the world evolved, and neither separates “when this was true in the world” from “when the system came to believe it.” That gap is the wedge.

The technical name for closing it is bitemporal: every fact carries two timestamps, one for when it was true in the world (valid time) and one for when the system recorded it (transaction time). Correcting a fact does not overwrite it. The old version is closed; a new version is opened; both remain. The graph keeps two timelines. When a decision is reviewed later, the system can produce the full course it took and the grounds it rested on. The audit trail is not a feature bolted onto the side; it is the base layer.

This matters more in 2026 than it did five years ago. The reason is that the inputs to enterprise knowledge systems are now LLM-extracted, and LLM extraction is non-deterministic in ways that matter. The same chunk re-extracted against the same ontology can produce different entity names, different type assignments, different relation arguments. If the system overwrites on each extraction, the audit log is a series of lucky survivals — whatever the latest run happened to produce. If it appends with both timestamps, the audit log is a record of what was believed, when, and on what evidence, and you can replay a decision against the knowledge state at the time the decision was made.

The eight crates

The Cargo workspace has eight members: utopia-core, utopia-store, utopia-server, utopia-ingest, utopia-extract, utopia-reason, utopia-search, utopia-llm. The Cargo.toml reads like a litany of deliberate choices. sqlx is pinned to 0.8 with postgres, mysql, bigdecimal, and uuid features; the comment explains that MySQL is only enabled for the query engine — TiDB, OceanBase, Doris, and StarRocks all speak the MySQL protocol, so a single feature flag covers a fleet. bigdecimal is mandatory because the MySQL driver rejects DECIMAL from both f64 and String (“floating-point numbers have different semantics”); without it, every monetary column would arrive as null. The Postgres connection uses runtime-tokio and tls-rustls. The HTTP stack is axum 0.8 with multipart (file uploads) and tower-http 0.6 with cors, trace, and fs. Search is Tantivy 0.26 with tantivy-jieba for Chinese tokenization. Vectors are pgvector 0.4. Auth is argon2 0.5 with the std feature, JWT via jsonwebtoken 10, and credentials encrypted at rest with aes-gcm 0.10 in AES-256-GCM mode — “the key never enters the database.”

There is one subtle dependency choice worth flagging. reqwest is pinned with default-features = false and then explicitly re-enabled with gzip. The comment explains why: an embedding response is 64×1024 floats, JSON-encoded at 21 bytes per number with 4 bytes of information, so a single response is 1.33 MB. On the same network path, uncompressed the round-trip is 62 seconds; compressed (gzip) it is 21 seconds at 429 KB. The server’s compute side of that takes one second; the rest is in transit. Default features would have flipped that on automatically — the explicit re-enable here is a comment that the team measured this and decided it had to stay.

The HTTP Content-Type round-trip matters because the pipeline extracts one embedding per chunk, and a long document can produce thousands of chunks. The decision to keep gzip on is a one-line change with a 3× throughput effect.

The pipeline, in twelve stages

The pipeline document is where the engineering shows. There is a single mermaid diagram at the top that deserves to be in any serious reference on document-to-graph systems, because it makes the two-phase split explicit. A document is ready as soon as embedding finishes; search and Q&A work immediately while extraction queues in the background. A long document takes minutes to grow its graph but is searchable within seconds. That is a load-bearing architectural choice — most RAG systems force you to choose between search freshness and graph quality. Utopia lets you have both.

The chain goes: upload / source sync → parse (parsers.rs) → chunk (1,200 characters, 150-character overlap) → embed (chunks.embedding) → Document ready (searchable and askable) → extract (one LLM call per chunk) → entity resolution (who this mention is) → facts to store (bitemporal ledger) → adjudication (one batched LLM call) → merge / keep apart → Graph. In parallel: facts go to type resolution (what this entity is) and ontology growth (out-of-vocabulary terms become proposals). The dotted line back into extraction is the loop: extraction uses the ontology; when it meets a term the ontology lacks, it records the original wording; proposals flow back into the ontology; the next batch of documents is extracted with it.

The 11 reason codes

Every stage ends with a “Where this stage drops things” section listing the drop points that actually exist in the code, in a database table called extraction_drops. The extraction stage has eleven reason codes, and they are not aspirational. They are the kinds you can already count in production:

ReasonWhen
truncated_replyThe model’s output was cut off; the whole chunk is discarded
malformed_itemOne fact is malformed; only that item is dropped, not the chunk (#127)
not_an_entity_nameThe “entity name” is a whole sentence (judged by word count and finite verbs, #143)
low_confidenceThe model’s self-reported confidence is below the threshold
subject_not_declaredThe subject is not declared in entities; recorded on both the relation and attribute paths
attr_domain_mismatchThe attribute is attached to a class outside its domain, even after walking up the parent DAG
attr_no_value / attr_datatypeThe attribute fact has no value, or the value cannot be converted to the declared datatype
object_missingThe relation fact has no object
direction_correctedNot a drop: subject and object were swapped per the signature; recorded so it is never silent
domain_mismatchThe swap is also invalid; the predicate is dropped (subject, object and evidence stay)

The most expensive of these is attr_domain_mismatch. It drops the fact at write time, and retyping the entity later does not recover it — the fact was never written and can only be re-extracted. That asymmetry is the kind of insight you only get by running a system at scale.

Why specific_type is the field everyone forgets

The extraction prompt lays out two paths depending on whether the ontology fits in the context budget. If yes, the whole ontology goes in (the small-ontology path). If no, retrieve by the chunk vector — about 40 classes, 30 relations, 30 attributes, plus ancestors of the hit classes. The chunk text and entities already accepted in this document go in alongside the ontology. The model returns items, and each item gets routed: entities go to the entity store with type picked from the list and specific_type left as free text; predicates that match an attribute go to attribute facts with the value normalized by datatype; predicates that match a relation go through a signature check.

The specific_type field is the easiest output to overlook. Free text, unvalidated, never entered into the ontology. It is the model’s own description of the entity — something like “vector database software” — and type resolution uses it to turn the task from “understand what this is” into “which class in the ontology has this name.” Without it, a test run left proposed_type empty on all 17 entities: the list always has something close enough, the model picks it, and the more precise description is lost. That is a paragraph buried deep in the pipeline doc, and it is the kind of thing you only write down after you have watched a system silently degrade because the obvious-looking field was the wrong thing to populate.

The signature check that survives prompt tuning

The signature check deserves its own paragraph because it is the kind of fix that gets discussed in conference talks and rarely shipped. The ontology declares employee (organization → person). The model still wrote Musk employee Microsoft. Three rounds of prompt tuning could not suppress this because English “X is an employee of Y” is too strong a signal — the model will keep emitting it. So the write path does not rely on the prompt: it checks the signature against the ontology’s domain and range, and if the subject violates the domain and the object fits, it swaps them and records direction_corrected. Never silently. If the swap is also invalid — same model, different document, same problem, this time OpenAI affectedBy … against a schema.org medical-test predicate — the predicate is dropped and subject, object, time, and evidence are kept. The graph stores the relation shape it was told to enforce, not the relation shape the model happened to emit.

Entity resolution with three tiers

Entity resolution is the second pipeline stage, and it is where the project paid for its engineering culture. There are two thresholds and three tiers, with names you can grep for: SIM_ATTACH = 0.55, SIM_NEW = 0.35. Clearly the same attaches; suspiciously similar creates a new entity and queues it for review; dissimilar creates one without touching the queue. “Prefer splitting over merging” is the rule, because a wrong merge mixes two entities’ facts together, which costs far more than one extra entity in the queue.

The Sherlock Holmes benchmark in the pipeline doc is the cleanest argument I have read for the value of layered fixes. Four runs over the first six stories of The Adventures of Sherlock Holmes (the corpus lives at scripts/bench/corpora/holmes.json), each run fixing one layer:

BaselineSame-type fixAlias recallRedirect
Entities merged14374757
Holmes merged✗✓✓✓
Mr. Holmes merged✗✗✗✓

Layer one: classify_type_drift had no tier for “both types equal,” so person × person fell into Disjoint, “can never be the same.” Containment recall later borrowed it as a compatibility test, where equal sides are the norm. The most obvious coreferences in the text never reached the queue, and all twelve existing unit tests covered cross-type cases. Layer two: recall looked only at canonical_name. A merge moves the name into aliases, so every successful merge removed a bridge — once Holmes was merged, the later Mr. Holmes and Sherlock Holmes contained neither the other, and Holmes had been the bridge. Fixing layer one exposed layer two. Layer three: a merge closed the other pending reviews involving the merged-away entity as superseded by merge; the code comment claimed a later mention would raise the doubt again. That was wrong: containment recall runs only when an entity is created, and these entities already exist, so closing was permanent. Now they are redirected to the merge target, and only two genuinely stale kinds are closed — pairs that become self-loops after the redirect, and pairs whose target pair is already queued.

The line that earns the section is “each layer hid the next.” That is exactly the failure mode that makes RAG systems feel unreliable in production: every fix exposes a new failure, and the new failure looks like a regression because the benchmark went down before it went up. The Holmes corpus is a way to make that failure mode legible.

The cold start problem, deliberately deferred

A new knowledge base has no vocabulary of its own. The starting point is an optional prebuilt pack (schema.org, W3C Org, PROV-O, FOAF, IOF Core — none of these are checked by default), a user-imported OWL file, or nothing at all, which is the default. An empty ontology still extracts; an entity without a type simply has no type. The design choice is in 0008-ontology-packs-as-cold-start.md and 0009-no-type-is-a-type.md: the system does not pretend to know what your domain is, and it does not silently infer a type where none was declared. Extraction picks fine-grained classes after per-chunk retrieval, and sometimes picks wrong ones (绍兴 → address is the canonical example). The right answer is a sibling of the wrong class, so it always reads as crossing a classification axis, which is sent to a human rather than auto-retried.

That decision is also why the ontology growth loop exists. Extraction uses the ontology; when it meets a term the ontology lacks, it records the original wording in fact_evidence.proposed_predicate, the one place on the fact that still carries the original meaning. The next batch of documents is extracted with the proposed predicate mapped to a candidate class. The system learns its own vocabulary over time, but it never silently invents one — every addition has a paper trail back to the chunk that suggested it.

Conflict detection as a first-class feature

The conflict detection code is in three places, and the design decision 0017-a-contradiction-points-upstream.md is the one I would print out and pin above any data-engineering team’s desk. There are three kinds of conflict and three sets of choices:

  • A new fact that clashes with an older one: close the old, keep both, or reject the new.
  • Data that breaks an axiom (self-loop, asymmetry, transitive cycle, cardinality): retract the fact, relax the axiom, or accept both.
  • The ontology itself is checked first, because violations of a self-contradictory ontology are noise.

The decision to check the ontology first is the kind of thing that does not look important until you have a system with 50,000 facts and 2,000 ontology axioms and an engineer notices that the consistency check is reporting 400 violations, none of which are real, because the ontology has a transitive cycle in a subClassOf chain that was hand-edited two months ago. Checking the ontology first turns that from a debugging session into a config change.

Ontology-driven SQL, the sibling project

The query side of Utopia mounts a database on a base — Postgres, MySQL, and the engines that speak its protocol (Trino for Iceberg / Delta Lake / Hive, Databricks, Snowflake) — and lets chat query it alongside the documents. The agent proposes how its tables map onto the ontology, the user confirms, and the system generates SQL against the mapped schema. The method behind it is a separate repository, deeplethe/ontology2sql, and it posts state-of-the-art numbers on BIRD Mini-Dev for SQLite and PostgreSQL (submission PR #218 on the BIRD bench repo).

This is the part where the project’s age shows. State of the art on BIRD Mini-Dev is the kind of claim that needs careful reading, because BIRD is a benchmark that has been targeted hard for three years and the leaderboard is full of systems that post on the dev set and don’t generalize. But the architectural choice — the agent proposes the table-to-ontology mapping and the user confirms — is sound, and the fact that the mapping is visible and editable rather than implicit in a prompt is what gives an enterprise deployment any chance of debugging a wrong query.

What I keep coming back to

There are three things I would steal from this codebase for any serious RAG system.

The first is the bitemporal graph. Most RAG stacks store the latest extraction and overwrite; Utopia stores the entire history of what was believed and when. The cost is roughly 2× the storage (every correction is an insert, not an update), and the benefit is that any decision can be replayed against the knowledge state at the time the decision was made. For an enterprise system that has to answer “why did we approve this loan in March” six months later, that 2× is the difference between a defensible answer and an unanswerable one.

The second is the signature check on relations. Prompt tuning cannot suppress English surface forms; the ontology enforces the signature at write time, records direction_corrected when it swaps, and drops the predicate when the swap is also invalid. This is what it looks like to take an LLM seriously as an extractor without taking it seriously as a reasoner.

The third is the design-decision document pattern. Twenty numbered decisions in docs/decisions/, each one a few hundred words, each one explaining the alternative considered and why it was rejected. 0013-a-source-should-hand-over-its-history.md is the one I think about most: when you ingest a source, you take its history with you. If a Jira ticket was reopened, the fact that it was closed is part of the fact, not a separate event. The default in most pipelines is to take the current state and discard the history. The default in Utopia is the opposite, and the design doc explains why in three paragraphs.

What it doesn’t fix

It is still v0.1. The status section in the README says “the database schema evolves between versions and migrations only roll forward, with no rollback.” There is a SECURITY.md you are told to read before exposing it to the public internet. The dev branch is where most of the action is, and the gap between dev and main is meaningful. The benchmarks at 100,000 documents are on the roadmap, not in production. The instant precision on the time axis — distinguishing 09:37 UTC from 09:38 UTC, not just dates — is open (#timeline-precision on the roadmap). The type resolution pipeline runs only by hand today: preview → apply on the ontology page, with no automatic trigger. Finishing extraction only enqueues ontology growth and entity adjudication, so refining new entities under a large ontology depends on someone remembering to click.

The single most consequential gap is one the team has flagged explicitly. Trigram similarity (CREATE EXTENSION pg_trgm) would close a known hole in entity resolution — two names that neither contain each other nor share an alias as a bridge, like 启明 X7 加速卡 and 启明 X7 推理加速卡. The repository connects with a restricted Postgres role, and CREATE EXTENSION pg_trgm needs superuser. So the team left it open rather than weakening the security posture of the database. That is a deliberate trade-off, but it is a real one, and anyone deploying Utopia at scale will feel it.

There is also the Open Design Arena-style question of whether an enterprise world model with bitemporal storage is the right base layer for an agent that needs to answer questions today, not audit them tomorrow. Most agent workloads do not need bitemporality; they need fresh results and a citation chain. Utopia gives you both, but the cost of the bitemporal layer shows up in write throughput, not read latency, and that shows up in any deployment where you are extracting at scale.

Where to dig further

The repository is github.com/deeplethe/utopia. The design decisions live under docs/decisions/. The pipeline document is docs/pipeline.md and is worth reading end-to-end. The sibling text-to-SQL project is github.com/deeplethe/ontology2sql. The five release candidates are tagged v0.1.0-rc1 through v0.1.0-rc5 between 2026-09-02 and 2026-09-05. The license is Apache-2.0. The Docker compose file gets you a running system in one command (docker compose --profile app up -d); the first account you register becomes the system administrator.

What I would watch, over the next two or three release candidates, is whether the cold-start ontology packs (schema.org, W3C Org, PROV-O, FOAF, IOF Core) are joined by industry-specific packs — financial, legal, medical — that an enterprise team would actually deploy with. The roadmap mentions asking for packs on a per-industry basis, and the existing five are general enough that almost any enterprise knowledge base will need at least one custom pack to extract anything useful. If the team ships one or two of those before the v1.0 tag, the system goes from “interesting open-source substrate” to “credible alternative to the closed enterprise knowledge platforms.” If they don’t, the open question is whether the Apache-2.0 license and the eight-crate Rust workspace are enough to attract the contributor community that a project at this level of ambition needs.

Aniket Karne
DevOps & AI Engineer · Amsterdam
Back to all posts
Reader correspondence

Comments

Powered by GitHub Discussions via Giscus. Sign in with GitHub to leave a comment.