{"path":"research/world-models-and-the-ontology-revival.md","content":"# World Models and the Ontology Revival: Where the Consistency Constraint Lives\n\n**Date**: 2026-08-16\n**Occasion**: The founder asked whether the 2026 world-models debate and the ontology revival bear on the structure-versus-scale question, and on his own development environment.\n**Status**: Substantial update to [structure-versus-scale.md](structure-versus-scale.md). It supplies a better articulation of the architecture bet than the one retracted on 2026-08-13, and it supplies the counter-evidence that framing has to survive.\n\n---\n\n## 1. The finding that matters most, and it is not about language at all\n\nVafa, Kleinberg, Mullainathan & Rambachan trained transformers on 4.7 billion tokens of turn-by-turn taxi trips in Manhattan. By surface metrics the models were excellent: next-turn predictions were valid close to 100% of the time, and internal state appeared to encode position.\n\nThen the researchers reconstructed the map the model had implicitly built. **It was an imagined New York** — streets that do not exist, flyovers crossing other streets, intersections joined at impossible angles. Close just 1% of the real streets and accuracy fell from near-perfect to 67%.\n\nThe model was a pile of locally correct heuristics that were never once required to add up to a single coherent Manhattan.\n\n**The cause is the objective, not the architecture.** Next-token prediction rewards being plausible at each step and never penalises the whole for failing to cohere. As one summary of the debate puts it, the fight is not language versus the physical world — LLMs demonstrably build world models, as Othello-GPT and the space-and-time probes showed. **The split is over where the consistency constraint lives: bought through architecture, or hoped for from scale.**\n\n## 2. Why this is the corpus's own diagnosis, arriving from outside for the third time\n\nLocally correct and globally incoherent is not a new shape here.\n\nThe **obfuscated-arguments threat** in the standing threat model is exactly it: locally valid, globally deceptive argument structures that spread error thin across many barely-checkable subclaims. **Dogfood run 6** produced a concrete instance: a locally correct judgment (these two sentences assert the same thing) that was globally wrong (their authors are in direct opposition). And the **sheaf framing** from the decomposition work names it formally — local sections that are individually consistent with no global section to glue them, which Abramsky showed is the same structure as contextuality and logical paradox.\n\nThree independent literatures, one project's own dogfood log, and now a large empirical result in machine learning, all describing one failure mode. That convergence is worth more than any single citation.\n\n## 3. The reformulation this licenses, and the test it has to pass\n\nOn 2026-08-13 a proposed restatement of the architecture bet was retracted the same day, because *reasoning must be persistent and contestable* could not lose — a forum with permalinks satisfies it, so it argued for a transcript rather than a claim graph. Any new framing has to clear that bar.\n\n**The candidate**: *coherence is not a property that emerges from accumulating locally plausible statements; it has to be imposed by something. Deliberus imposes it in the substrate.*\n\nDoes it clear the wiki test? **Yes, and this is the improvement.** A well-kept wiki has no mechanism that penalises global incoherence — it relies on an editor noticing. A claim graph with attack edges and propagated strength has one: an inconsistency between two claims is a computable relation, not a thing someone has to spot. The rival that defeated the previous framing does not satisfy this one.\n\nIs it falsifiable? **Yes, and it is currently failing in one place.** If the graph accumulates locally-correct-globally-incoherent structure at the same rate prose does, the constraint is not binding. Run 6 found exactly such an accumulation at the relationship layer, which is why the stance instrument exists and why it reports candidates rather than verdicts. **The honest statement is that the constraint is imposed at the claim layer and not yet at the relationship layer**, and that gap is measured rather than asserted.\n\nWhat this framing does *not* license: any claim that the graph is coherent. It is a claim about where the mechanism sits, not about whether it has succeeded.\n\n## 4. A second argument, about maintenance rather than capability\n\nA 2026 study of the Tower of Hanoi found something more useful than another capability ceiling. Frontier reasoning models **encode a faithful representation of the puzzle state at the end of the prompt** — the probe recovers it near-perfectly — and then **lose it during generation**. Restoring the prompt-time representation by injection at inference nearly doubled one model's optimal solve count. The authors' framing: the models build a world model and then forget it.\n\n**This is an argument for external persistent structure that has nothing to do with what a model can understand.** A representation held in a graph does not decay across a long trace. That is a mechanism-level point the corpus did not previously have, and it is narrower and more defensible than any claim about reasoning ability — the failure is maintenance, and externalised structure is a maintenance solution.\n\n## 5. The ontology revival, and the part of it that is uncomfortable\n\nOntologies came back in 2026 as guardrails for probabilistic agents. AWS shipped an open-source Context Ontology Accelerator in July with W3C standards and an MCP server; Forrester, BCG and Gartner are all pushing semantic layers; Neo4j's framing is business ontology plus technical ontology plus execution traces. The stated motive throughout is that agents without explicit context guess, and guessing agents act.\n\n**Three things transfer to this project, and one of them stings.**\n\n*The formalism matters less than the content, and the corpus already bet that way.* A widely-shared argument from the data-engineering side is that markdown in a filesystem works as well as OWL for most agentic uses, because the content of the ontology does the work and the syntax does not — the consumer is now a language model, and a language model reads language. Deliberus's agent-readable surface is markdown-primary and deliberately not AIF-conformed, which is the same call made independently. **The sting is in the general direction of the argument**: it is an argument against formalism for its own sake, and this project has a great deal of formalism. The defence has to be that each piece of typing earns its keep in something computable, which is exactly the open question in `decomposition-axes.md`.\n\n*The decision layer has to move into the ontology.* The sharper version of the revival's argument: ontologies were built as **comprehension** layers, because the only general-purpose decider was a human who would read the model and act. The decider got cheap. So decision logic that used to live in human heads now has to be written into the ontology, or agents reason over terrain with no rules. For Deliberus this reframes the critical-question layer — those questions *are* decision logic made explicit, and that is a stronger reason for them than pedagogy.\n\n*The maintenance problem is the known killer, and the founder's instinct matches the field's current answer.* The Semantic Web died on ontology maintenance. The 2026 answer being floated is agent-maintained ontologies — the agent updates definitions as it meets edge cases, changing the character of the problem. That is precisely what the founder described wanting: scheduled agentic workflows that continually structure, clean and understand the graph and its ontology. **His instinct is the field's current best answer to its own historical failure**, and the honest addendum is that everyone proposing it also calls it unsolved.\n\n## 6. Counter-evidence, carried rather than buried\n\n**Structured environments are where LLM world models work.** An ACL 2026 study across five text environments found that world-modelling benefits hinge on behavioural coverage and environment complexity: structured, rule-driven environments saturate quickly and cheaply, while open-ended ones scale badly and one showed no saturation at all. **Argumentation is open-ended.** So the encouraging results come from the regime least like this one.\n\n**The spectrum argument weakens the dichotomy the whole debate rests on.** A 2026 position paper argues LLMs are a degenerate special case of world models — the state space is token sequences and the only action is appending a token — with a continuous path from next-token prediction through multi-token and next-latent prediction to JEPA. If that is right, \"structure versus scale\" is a question of *how far along a spectrum* rather than a choice between paradigms, and the honest position is that the destination is agreed while the distance is not.\n\n**And the enterprise numbers do not transfer.** Figures circulating in the revival — accuracy rising from the teens on raw schemas to the seventies with an ontology-backed graph and a query check — come from database querying, where there is a ground truth to be right about. Discourse has no such target. Quoting them for Deliberus would be a category error, and they are noted here so that nobody later quotes them from this document.\n\n## 6b. The talk itself, and a correction to §7 below (2026-08-16)\n\nThe founder pointed at Frank Coyle's *Why Agentic Systems Need Ontologies* (AI Engineer World's Fair, July 2026) as the sense of \"ontology\" he meant. Reading the slides makes the architecture concrete and shows the sense is narrower than this document had been using.\n\n**Two gates around a tool call.** In a Claude agent loop, the model proposes a call and two natural gates surround it. **Gate 1, before the tool runs**, validates the *shape* of the call — types and parameters, Pydantic's job. **Gate 2, after the tool runs**, validates the *coherence* of the result against the domain model — the ontology's job. His summary: Pydantic guards what goes in, the ontology guards what comes out.\n\n**The layer is not probabilistic, and that is the whole point.** RDFS *infers*: declare `teaches` with domain Teacher and range Student, assert that Bob teaches Scooter, and a reasoner derives unasked that Bob is a Teacher and Scooter a Student. OWL *constrains*: transitive properties close chains, and a functional property means at most one — so if Stewie has father Peter and also father Peter_G, the reasoner concludes those are the same individual. These derivations sit beside the graph rather than in it, and become guardrails the loop must obey. The errors he cites are the giveaway: a second refund on the same order, a payout routed to the support desk instead of the buyer, an order status of *\"probably shipped.\"* **Catches that are painful to write in English become a few lines of logic.**\n\n**The correction I owe.** An earlier version of §7 said the founder's development environment \"already is the pattern the field converged on,\" on the grounds that his working ontology is markdown maintained by agents. Against Coyle's sense that is **wrong**, and the talk's central claim is precisely why: a prose instruction shapes probability, while an ontology constraint either passes or does not, and no amount of prompt engineering closes that gap. `CLAUDE.md` is documentation-as-context. It is not a validator.\n\nWhat is true, and more useful: **his hooks are Gate 1 and he has no Gate 2.** The PreToolUse denies validate the shape of a call before it runs, which is exactly the first gate. Nothing anywhere validates the *coherence of a result* against a domain model afterwards.\n\n**And Deliberus has been hand-rolling Gate 2, one constraint at a time, each after a bug.** The clearest instance was created the day before this was written: the requirement that the `REPORTS` relation carry no strength is enforced by **a test that greps the propagation modules for a string**. In OWL that is a property characteristic, declared once and machine-checked. Three more of the same shape: the `VALID_RELATIONSHIP_TYPES` whitelist is a range restriction written by hand after a Cypher-injection audit; `count_safe_summary`'s `unclassified` bucket is a range check on `claim_kind` written by hand because a silently absorbed new kind would erode a count guarantee; and the Literal-versus-Enum discipline is a value constraint whose failure cost **fifteen days of silent zero-claim extractions** in April.\n\nMeasured against the live graph on 2026-08-16, three such constraints currently hold with **zero violations** — no claim carries two terminus classifications, no `claim_kind` sits outside the known set, and no `EXTRACTED_FROM` edge points at a non-Source. That is clean by the current write paths rather than by construction, and nothing would catch it if a path changed.\n\n**The reflexive point is the one worth keeping.** Deliberus already *is* a neurosymbolic system for discourse: an LLM extracts, and a typed graph with computed strength constrains what the extraction can mean. What it does not yet have is that same discipline applied **to its own schema**, one level up. The platform imposes on the world a rigour it has not imposed on itself, which is the operator-reflexivity gap the July red team named, appearing in a new place.\n\n## 7. For the development environment, briefly\n\nThe founder asked. His setup is already an instance of the pattern the field is converging on: a machine-readable ontology of his own working practice, written in markdown rather than a formalism, consumed by agents, and maintained by those agents in the course of their work. The directive-scan hook and the conditional event hooks are the maintenance mechanism the Semantic Web never had — the thing that keeps the ontology from rotting is that it is exercised on every turn rather than curated on a schedule.\n\nThe transferable observation runs the other way, from the field back to him: **execution traces are the third ontology** in the Neo4j framing, and his are largely unmined. Session transcripts, hook firings and the record of which directives fired and which did not are the runtime signal about how the working ontology actually behaves, as opposed to how it is written down.\n\n---\n\n## Sources\n\n- Vafa, Kleinberg, Mullainathan & Rambachan, *Evaluating the World Model Implicit in a Generative Model* (NeurIPS 2024) — the Manhattan taxi study.\n- Li et al., *Emergent World Representations* (Othello-GPT); Gurnee & Tegmark, *Language Models Represent Space and Time*.\n- *Transformers Struggle to Use Their Emergent World Models: Revisiting the Tower of Hanoi*, arXiv:2608.07077 — build-then-lose, and recovery by injection.\n- Dubois, *From Tokens to States: LLMs as a Special Case of World Models*, arXiv:2606.28127 — the spectrum argument.\n- *From Word to World: Can Large Language Models be Implicit Text-based World Models?*, ACL 2026 — structured versus open-ended scaling.\n- Jon Moshier, *World Models vs. LLMs* (2026-06-06) — where the consistency constraint lives.\n- Richard MacManus, *Ontologies Are So Back*, Latent Space (2026-07-30) — the revival, neurosymbolic guardrails, Neo4j's three ontologies.\n- dltHub, *Ontology engineering: what it is, why it's back, and why agents need it* — markdown over OWL, and the decision layer.\n- AWS Context Ontology Accelerator (2026-07-01); Forrester and BCG semantic-layer reports (2026).\n\n**Read depth**: abstracts and results sections for the arXiv and ACL papers; the Latent Space and dltHub pieces in full; the enterprise reports at summary depth, and their numbers deliberately not imported.\n\n## See also\n\n- [structure-versus-scale.md](structure-versus-scale.md) — the bet this updates, and the retraction any new framing has to survive\n- [dogfood-run-6-israel-palestine-cross-domain.md](dogfood-run-6-israel-palestine-cross-domain.md) — the corpus's own locally-correct-globally-wrong instance\n- [decomposition-axes.md](decomposition-axes.md) — the sheaf framing, and whether each piece of typing earns its keep\n- [scaffolding-versus-difficulty.md](scaffolding-versus-difficulty.md) — the opposing pressure, that formal structure can substitute for the competence it means to support\n"}