{"path":"research/session19-decomposition-corrections-and-the-ontology-turn.md","content":"# Session 19: Four Corrections, a Decomposition Arc, and the Ontology Turn\n\n**Date**: 2026-08-14 to 08-16\n**Type**: Correction, research synthesis, and three builds. The organising pattern is that **every substantive advance this session began as something the founder challenged that turned out to be wrong in the docs rather than in the code.**\n\n---\n\n## 1. The pattern worth naming first\n\nFour times the founder questioned a claim, and four times the record disagreed with the recollection — mine, and once his own.\n\n**The run-6 stance claim was overstated and two days from shipping into grant applications.** Asked *\"when exactly did that happen and why was an inconsistency not found?\"*, the primary artifact showed two different claim pairs collapsed into a single before-and-after, and the live graph held no such edge at all. What survived was narrower and stronger: the classifier *returns* agreement at maximum confidence for a same-fact-opposite-use pair, so the risk is real and **unrealised**.\n\n**The DeepMind citation was half wrong.** Asked whether it had been analysed before, it had — in six places, at abstract depth, with an open reading task. Reading the paper closed the task and found their consensus measure sat near a ceiling and measures numeric allocations, so half of what the corpus cited from it does not hold.\n\n**His own memory beat the corpus twice.** The structuring gradient and the mother-claim pattern were both already his, documented, and inert.\n\n**And my dev-environment claim was wrong** once the talk's actual sense was visible.\n\nThe generalisable lesson, and it applied to me at least as often as to the docs: **the record beats the recollection, and the check is nearly always cheap.**\n\n---\n\n## 2. The decomposition arc\n\n### The founder's reasoning (verbatim)\n\nOn why critical questions cannot be the whole unbundling template, and what the standard should be:\n\n> \"Ok but then let's simply do it incredibly robustly and well thought through.\"\n\nOn the fair-share discount worked through as if a publication record were one unit of evidence:\n\n> \"This is not nearly complex enough. It needs to consider not only the topic of each paper but really it needs to decompose all of them: needs to exhaustively consider and evaluate the full contents of all 47 papers to know how well-backed they are. Then all relevant papers and relevant subclaims/subtrees of indirectly related papers also — EVERYTHING RELEVANT — needs to be taken into account in the graph.\"\n\nThat correction resized the discount rather than removing it. Once evidence decomposes, most *apparent* sharing dissolves — different children draw on different subsets — so the discount was partly correcting an artifact of coarse evidence, and its remaining job is genuine sharing. The implementation is granularity-indifferent, so what is missing is an evidence-decomposition pass distinct from routing.\n\n### The registered prediction (verbatim)\n\n> \"This is a living knowledge substrate which our interactions with will over time increase in what we get from them in terms of instructive/constructive value while asymptotically lowering the investment in terms of defining the right ontology and filling up the graph with extractions and human- or auto-populated implicit premises, evidence and shared fundamental knowledge/assumptions about how the world works. The costs here will likely go down over time while the usefulness will go up. This is my prediction.\"\n\nRecorded as a prediction rather than an assumption because it is testable and currently **unmeasured** — cost tracking was never implemented, so it is unmeasurable rather than unproven. The first sentence carries the content and had been dropped from both earlier records: what falls is **the investment required to keep the substrate populated and coherent**, the variable is corpus maturity rather than time, and the shape is asymptotic. Its April antecedent names the mechanism, and the split reading of the August measurement against it, in [self-similar-decomposition-and-claim-ontology.md](self-similar-decomposition-and-claim-ontology.md) § Registered prediction.\n\n### The decision\n\n**Always mint the reported content**, chosen against the analysis's own lean. Rationale: a position under attack is in play, and a map omitting it is not a map of that argument. Full reasoning, the four failure modes it creates, and the straw-man detector it unlocks: [self-similar-decomposition-and-claim-ontology.md](self-similar-decomposition-and-claim-ontology.md).\n\n---\n\n## 3. The civic framing, in the founder's words\n\nBrought with a climate thread, and it is the clearest statement this corpus has of what the platform is *for* at the scale of ordinary life:\n\n> \"There's some deep lessons here about the necessity of creating Deliberus to call (in the poker sense, or as x-ray vision) the various ideological water we swim in as the fish of society. Society is built up of relationships and trust but also mental models of how the world works and those can be mistaken.\"\n\nAnd the boundary he draws around it, which matters as much as the aim:\n\n> \"Companies and countries and families are bound together and I don't necessarily wanna split them up and have everyone join the hive mind of Deliberus hehe… but we can at least facilitate reflection about what ideological house of cards we prop up each day by our actions and words and when voting with our wallets and actions and non actions.\"\n\n**Two things to carry forward.** *Calling* is a better word than *challenging*: you are not asserting the other side is wrong, you are declining to take an implied strength on trust and paying to see it. And the hesitation is not sentimentality: it points at the documented gap where some agreements survive only while unspoken, and the ontology has no concept of load-bearing illegibility. The settled position is **make the descent available, never compulsory**.\n\n---\n\n## 4. The ontology turn, and what it asks of this project\n\nThe founder pointed at Coyle's *Why Agentic Systems Need Ontologies* as the sense he meant, which is narrower than documentation-as-context: **two gates around a tool call**, the first validating the shape of the call and the second validating the coherence of the result against a domain model, with the second layer deliberately non-probabilistic.\n\nThe finding that lands hardest is reflexive. **Deliberus has been hand-rolling that second gate one constraint at a time, each after a bug** — most vividly a rule enforced by a test that greps source files for a string, created the day before. Full analysis, including the three constraints measured clean against the live graph and the counter-evidence not to transfer: [world-models-and-the-ontology-revival.md](world-models-and-the-ontology-revival.md).\n\nThe same research supplied the framing that the 2026-08-13 retraction left missing — *coherence must be imposed, and the question is where* — which survives the wiki test the previous reformulation failed.\n\n---\n\n## 5. The experiment that came out of it, and whose it is\n\nAsked what decision the world-models literature forces, the honest answer was **none** — the spectrum argument says explicitly there is no fork to pick. What it bears on is a question already open: not whether to hold reasoning in a graph, but how much.\n\n**And the literature cannot answer that. The corpus can.** Take the cases where the structure fired — a stance-conflict candidate, a completeness gap, a hinge score — and give the same raw material to a model with no graph at all, asking the same question. If it catches everything, the structure is decoration at this size. If it misses, that is the first real number on what typed structure buys, on this project's own material rather than on someone else's benchmark.\n\nThe founder's response, and it is the right ownership: *\"It's an interesting idea. I'll mull over how to run that test in the best way.\"* The design is his; it is filed rather than pre-empted.\n\n---\n\n## 6. Late additions (2026-08-16)\n\n**The reported-speech pass was switched on**, and its first real output found a gap in its own validation. Three of 82 candidates ran before Gemini quota bit: two refused correctly (a plain assertion, and one at medium confidence), one minted — **and the minted one was a fragment**, cut off at *\"…is based on the claim that\"*, with content beginning mid-sentence in lower case. Every other pass here enforces the stranger test and this one did not, so validation now refuses halves that do not begin as sentences or end mid-clause, pinned with the actual failing strings. **A fragment in the graph is worse than no claim, because it looks like a claim and cannot be argued with.** The safety property held under a real write: asserted stayed at 1,678, the minted claim counts separately, and the only edge type touching it is `REPORTS`.\n\n**The uplift evaluation frame was rejected, and it contradicted the corpus's own economics.** Told that \"whether people reason better\" was the single most useful missing measurement, the founder answered that this is an individualist productivity measure and the project is a commons:\n\n> \"This is a bit like asking whether a single reader 'reads better' from a book from a public library. The benefits of having public libraries far outweigh that which is individually beneficial or productivity-increasing… I'm much more collectivistic and commons-oriented.\"\n\nThe sharper finding is that [lowering-the-cost.md](lowering-the-cost.md) §5 had already argued this formally — Romer nonrivalry gives the warrant that **value scales with corpus coverage rather than user count**. So the framing had drifted from the corpus rather than differing from it in taste, and it is the same legibility-shrinkage the fourth-identity work already named, this time wearing an evaluation frame instead of an identity. Two things kept alongside: the fit measures are commons-shaped and **each can still fail** (claim-level reuse already reads badly), so this is not an escape from falsifiability; and a library rests on a claim nobody disputes while this project argues a *new kind of artifact* is worth building, so the collective frame settles the unit of value without exempting the artifact from being good.\n\n**The registered prediction was audited against the code and the graph, and it has never been given its own conditions.** The April mechanism is checkable — a saturating layer of shared premises, so new descents terminate in already-known subclaims — and the graph holds **16 implicit-premise claims in total**, 18 of 25 sources have none, one of the sixteen is linked to anything, and `decompose_claim` performs no lookup of existing claims at all. Four independent build gaps, none of them a refutation: the interactive extraction path never runs the pass, the prompt is instructed toward scarcity, the object produced is a within-argument bridge rather than shared background, and reuse-on-descent does not exist. **The most consequential of the four is the third**, because bridge premises and background premises turn out to be different things and only the first was ever built. Full audit: [self-similar-decomposition-and-claim-ontology.md](self-similar-decomposition-and-claim-ontology.md) § The mechanism audited.\n\n**And the tooling question got a measured answer.** Asked why the semantic search over the life corpus was too slow to reach for, timing the stages showed startup at 0.10 s and a single short query embed at 2.67, 3.40 then 5.77 seconds across identical consecutive probes, against a documented idle figure near 80 ms. The rising curve is a growing queue rather than a hardware floor, and the shared embedder is single-slot with no parallel arguments — **so every consumer serializes, including this project's own auto-connect and similarity linking.** Full write-up in dotfiles: `docs/embedder_saturation_2026_08_16.md`.\n\n---\n\n## 7. Late additions (2026-08-16, second half)\n\n**The four build gaps were ratified.** Founder, on the implicit-premise audit: *\"All of this should be built.\"* So the reuse lookup in `decompose_claim`, wiring Pass 2c into the stream path, the background-premise decision, and raising bridge-premise recall are approved work rather than candidates. The ordering recorded in `TODO.md` puts the reuse lookup first, because it is the only one that makes the April mechanism *possible* rather than merely likelier, and it carries the April caution: propose, never auto-merge.\n\n**A cross-project bug arrived and turned out to be ours too.** The founder's semantic-search tool was found to have tokenised on `[a-z0-9]+`, silently shredding every Swedish word containing å, ä or ö — half its corpus, blind since the tool was built. Checking whether Deliberus carried the same failure class found **three instances**, and the measurement is on real Swedish words rather than invented shapes:\n\n- **`_sense_key`** (concept near-duplicate detection) used `[^a-z0-9]+`, so `rätt` and `rött` — right and red — produced an identical key. **The failure direction is manufactured sameness in the concept layer**, which is the one layer whose entire purpose is preserving distinction, and which `lowering-the-cost.md` §6.3 had already named as the cost curve's safety condition. Fixed and pinned.\n- **`reported_speech._normalize`** deleted accented characters outright, so `hål` and `häl` both became `hl`. Direction is a false refusal to mint rather than a bad write. Fixed and pinned.\n- **Source slug generation** shredded the same way, visible in the live graph as `…-egoistiskt-sk-l-…` where `skäl` lost its ä. Raised as a founder call because slugs are public URLs; answered the same day — *\"I don't care about broken old slug links\"* — and fixed. The two call sites now share `deliberus/slug.py`, which transliterates (`skäl` → `skal`, `Søren` → `soren`) and truncates before stripping hyphens, a second defect found alongside the first. **Slugs stay ASCII on purpose**: `_SOURCE_ID_RE` is a path-traversal guard on public GET endpoints, so admitting Unicode would loosen a security boundary to fix a cosmetic problem, and a test asserts every generated slug still passes that guard.\n\n**The slug work also sharpened the rule.** A test written to assert that `hål` and `häl` keep distinct slugs failed, correctly: ASCII folding *should* collide there. **The same fold is right in an identifier and catastrophic in a comparison key** — a slug names a URL, a key judges sameness — so the pinned assertion now states both halves together, with `_sense_key` and `_normalize` required to distinguish exactly the pair the slug is allowed to merge.\n\n**The lesson is the one the fixture directive names, one level up.** No English fixture can fail on any of these. The directive's own words for it: *not \"I invented the sample text\" but \"I invented the language\"* — and the exemption for pure logic reads safest exactly where it is most dangerous, because the more language-neutral a function looks, the likelier its correctness rests on the alphabet of the data it will actually meet.\n\n**And the always-mint decision's two hard requirements were found unwired**, both satisfied by accident: the badge endpoint matches edges *untyped* and filters on a `scheme` property, so `REPORTS` is non-propagating only because REPORTS edges happen to carry no scheme; and `count_safe_summary` has no caller, so `/graph/stats` still publishes undifferentiated totals. Detail and the fix shape: [self-similar-decomposition-and-claim-ontology.md](self-similar-decomposition-and-claim-ontology.md) § Both hard requirements. **This is also the approved ontology layer's first concrete constraint**, arriving from the graph rather than from the literature.\n\n*One correction to this session's own record: the search-latency diagnosis given here earlier attributed the cost to query length. It follows the matched term's **frequency** instead, because ranking scores every match to return twelve. Both that and the tokeniser are fixed upstream; the practical consequence is that long queries are fine again, and any pre-2026-08-16 \"found nothing\" on a Swedish query is worth re-running.*\n\n## Cross-references\n\nWritten or substantially updated this session: [self-similar-decomposition-and-claim-ontology.md](self-similar-decomposition-and-claim-ontology.md) · [decomposition-axes.md](decomposition-axes.md) · [world-models-and-the-ontology-revival.md](world-models-and-the-ontology-revival.md) · [scaffolding-versus-difficulty.md](scaffolding-versus-difficulty.md) · [conviction-and-critique.md](conviction-and-critique.md) · [structuring-gradient-lineage.md](structuring-gradient-lineage.md) · [dogfood-run-6-israel-palestine-cross-domain.md](dogfood-run-6-israel-palestine-cross-domain.md) · [nativism-thread-and-false-dissolution.md](nativism-thread-and-false-dissolution.md) · [synthesis-build-plan.md](synthesis-build-plan.md) · [structure-versus-scale.md](structure-versus-scale.md) · [curiosity-as-growth-fuel.md](curiosity-as-growth-fuel.md)\n\nShipped: `deliberus/stance.py` (deployed) · `deliberus/attribution.py` · `deliberus/evidence_routing.py` · `deliberus/reported_speech.py` · the Swedish introduction's corrected confession chapter (deployed).\n"}