{"path":"research/dogfood-run-7-swedish-election-structured-vs-unstructured.md","content":"# Dogfood Run 7: The Swedish Election, With Structure and Without\n\n**Date**: 2026-08-17\n**Status**: PREREGISTRATION — written and committed **before any source was read**. Results sections are appended below in order.\n**Prompted by**: the founder, who wanted the structure-earns-its-cost experiment run on material he does not already know. Verbatim: *\"It is election year in Sweden. I know pretty much nothing about contemporary Swedish politics... analyze the debate with structure and without.\"*\n\n---\n\n## 0. Why this design, and what it is a test of\n\nThe corpus has wanted this experiment since 2026-08-16, and it has a specific prior warning attached: **the control arm must be prose-plus-model, not \"no structure\"** ([structure-versus-scale.md](structure-versus-scale.md)), because against no structure at all the graph wins trivially and proves nothing. The primary rival identified there is *a well-kept wiki with a language model over it*.\n\nThis design satisfies that by construction. The unstructured arm is a frontier model reading primary sources and explaining a debate as well as it can — which **is** prose-plus-model. That is the rival, doing its best.\n\n**Order is fixed and it is the conservative direction.** Unstructured runs first, then structured on the same sources. Anything the structured pass contributes must therefore be *additional to what the reader already knows*, which is the harder test for structure. The reverse order would have flattered it.\n\n### What this can and cannot measure\n\nThe extraction pipeline requires an authenticated submission and the Gemini key is on free-tier remnants, so the structured arm is executed the way runs 3F, 5 and 6 were: **by hand, at pipeline fidelity**. That has consequences which must be stated before results, not after:\n\n- **This compares protocol, not systems.** The same reader performs both arms, so what is isolated is the contribution of the *ontology and the instruments*, not of the model. A finding that structure helped is a finding about the protocol.\n- **Two real components are genuinely exercised, not simulated.** The shipped instrument functions (`compute_claim_badge`, `hinge_scores`, `assess_completeness`, `assess_disagreement_preservation`, `find_stance_conflicts`, `detect_weighing`, `summarize_residue_map`) are pure and are run for real on the hand-extracted structure. And the **auto-connect candidate stage runs against the live 4,769-claim production corpus** through the public `/claims/search` endpoint.\n- **One component is unavailable and its absence favours the hand extraction:** the constrained scheme enum. Run 6 J14 measured that a careful human reader names schemes plausibly and *wrongly* — thirteen of fourteen remapped — where the pipeline's closed enum cannot make that error. Any scheme naming below is therefore weaker than production, not stronger.\n- **The corpus is English and the sources are Swedish.** This is a documented and unmeasured axis, not a nuisance: `TODO.md` carries *\"sweep the remaining lexical instruments for the same blindness against non-English sources\"* as unbuilt work.\n\n---\n\n## 1. Registered predictions\n\nWritten before reading. Each states what would count as wrong.\n\n**P1 — Division of labour.** The unstructured arm produces the better *narrative* (coalition dynamics, strategic context, who gains from which framing); the structured arm produces the better *map of disagreement* and locates at least one crux the narrative left implicit.\n*Wrong if*: the structured pass surfaces nothing the prose read did not already contain.\n\n**P2 — The crux is implicit on at least one side.** The most stable cross-run finding in this corpus (3F both batches; run 6 on both sides at once). Replication on a new register — partisan electoral rhetoric — is a real test rather than a rerun.\n*Wrong if*: both sides state their decisive premise explicitly.\n\n**P3 — Cross-domain auto-connect finds essentially nothing.** Run 6 measured 4 candidate pairs against the entire rest of the corpus, peak 0.685, zero at the 0.80 auto-link threshold. Predict Swedish election claims score **below 0.80** against the existing 25 sources, and that any near-misses are vocabulary-driven (welfare, labour, economics) rather than structural.\n*Wrong if*: anything crosses 0.80, or if the top matches are structurally kindred while lexically distant.\n\n**P4 — The English lexicons fire zero times on Swedish text.** `detect_weighing`, `detect_sacredness` and `stance.py`'s reported-speech patterns are English regexes. Predict zero firings on Swedish-language claims *even where a weighing is plainly present*, and non-zero once the same claims are rendered in English — which would locate the instrument's dependency on a translation pass that is itself a flattening step.\n*Wrong if*: they fire on Swedish, or if they stay silent on English renderings of the same content.\n\n**P5 — Structure costs coverage.** The structured arm will cover materially fewer topics per unit of effort. Registered because \"earns its cost\" has a cost side and this is where it shows up.\n\n**P6 — The risky one, registered because it can embarrass the project.** Structure may add *less* here than on the philosophical and legal debates the corpus has tested. The founding conviction says lowering the cost dissolves **confusion**-conflict and only *clarifies* **interest**-conflict. An election is closer to the second: parties differ on distribution and on identity, and much of the disagreement may be genuine opposed interest rather than recoverable misunderstanding. If so, decomposition should shrink these disputes less than it shrank the white-lie descent.\n*Wrong if*: the deep topics decompose to small shared-premise residues the way the everyday value clash did.\n\n### What a loss looks like\n\n**Structure loses this experiment if** the structured pass surfaces nothing the prose read missed, at a visibly higher cost. That is a publishable negative and it will be reported as one. The corpus's own rule applies: an instrument that can only confirm is decoration.\n\n---\n\n## 2. Method\n\n1. **Landscape survey** — what is actually contested in the 2026 Swedish campaign, from current sources rather than from model memory. Breadth, cheap.\n2. **Selection** — two topics carried forward, chosen for a genuine adversarial pair of primary sources. Selection rule fixed before reading the pairs: the topic must have (a) two identifiable sides publishing substantive argument, (b) a mix of empirical and normative content, (c) no dependence on facts after the model's May 2026 cutoff for the *reasoning* to be legible.\n3. **Arm A, unstructured** — read the sources, produce the best available explanation of the disagreement as prose.\n4. **Arm B, structured** — hand-extract the same sources at run-3F fidelity, then run the real instruments and the live-corpus probe.\n5. **Comparison** — what each surfaced that the other did not, with the confounds named.\n\n---\n\n*Results are appended below this line as each phase completes. Nothing above this line is edited after the fact; corrections are appended and dated, per the corpus convention.*\n\n---\n\n## 3. The landscape survey\n\nSweden votes on **13 September 2026**. Measured against the preregistered rule, the contested field as reported by current sources:\n\n| Issue | Status in the campaign |\n|---|---|\n| **Law and order / gang crime** | Consistently top-ranked by voters across polls |\n| **Immigration and integration** | Top-ranked; SD pushing a hijab ban, C attacking the 33,000 kr work-permit salary threshold |\n| **Healthcare** | Waiting lists, staffing, regional inequality; nationalisation is the sharpest *ideological* split |\n| **Economy, tax, cost of living** | \"It's never been more expensive to be Swedish\" (S) against warnings of tax rises and a bank-levy passthrough (right bloc) |\n| **Schools** | Recurrently high, lower salience this cycle |\n| **Climate and energy** | Risen to third by voter salience — but see below |\n| **Government formation** | Arguably the meta-issue: which bloc can form a stable government |\n\n**Blocs.** Tidö: M (Kristersson, 66), **SD (Åkesson, 70)**, KD (Busch, 19), L (Mohamsson, 16). Red-green plus C: S (Andersson, 106), V (Dadgostar, 21), MP (Lind/Helldén, 18), C (Thand Ringqvist, 24). SD is now the largest party in its own bloc, and the March 2026 \"Sweden Promise\" between L and SD opens cabinet seats to SD given a majority. First election since NATO accession.\n\n**One topic was surveyed and dropped, which is itself informative.** Energy looked like the obvious sharp divide and is not one any more: M, KD, L, SD **and S** all now support new nuclear, C accepts it on market terms, and only V holds a clear line against while MP has stopped pushing rapid phase-out. The live disagreement has narrowed to financing terms and pace. It fails selection criterion (a) — there is no longer a clean adversarial pair — and a reader working from priors rather than sources would very likely have picked it.\n\n**Carried forward:** gang crime (deterrence vs prevention) and healthcare (nationalisation and the queue fight).\n\n### Sources (fetched 2026-08-17; this list is exactly what a pipeline run would need)\n\n| # | Side | Source |\n|---|---|---|\n| S1 | SD | Sveriges Radio, Roger Hedlund — *\"Avskräckning förebygger gängkriminalitet\"* |\n| S2 | S | socialdemokraterna.se — *Gängkriminalitet* |\n| S3 | L (Tidö) | liberalerna.se — *Gängkriminalitet* |\n| S4 | KD | Läkartidningen — *\"KD kräver beslut om statlig vård för regeringsstöd\"* |\n| S5 | mixed | Fokus — *\"Kö-striden som kan avgöra valet\"* |\n\n---\n\n## 4. ARM A — the unstructured analysis (prose + model, no graph)\n\n*Written from the five sources above with no extraction, no ontology, no instruments. This is the control arm and it is the project's actual rival, so it is written to be as good as I can make it.*\n\n### 4.1 Gang crime: an argument that is narrower than it sounds\n\nThe public framing is punishment against prevention. The sources do not support that framing.\n\n**Every party in these sources supports harsher punishment.** The Social Democrats — the opposition — propose a Swedish \"mafia law\" for collective prosecution of networks, **double sentences for gang criminals**, and a ten-year maximum for economic crime. The Liberals have already supported sentence increases and reduced release discounts. The Sweden Democrats want higher penalties and immediate confiscation of status symbols. On the punitive axis the parties are close to indistinguishable, and the government has already legislated a doubling provision for network-linked crime.\n\nSo the disagreement is not *whether* to punish harder. It is about what else to do, and about the *marginal* question: given finite money and attention, where does the next unit go.\n\n**And the sides do not actually hold symmetric positions on that.** Reading closely, there are three distinct stances, not two:\n\n- **SD (Hedlund) argues substitution.** Prevention is \"less effective than deterrent methods\" and is dismissed as \"too difficult and costly\". Note the structure: this is not the claim that prevention *fails*. It is a claim about tractability and cost-effectiveness — prevention might work and still not be worth doing. Those are different propositions and the interview runs them together.\n- **The Liberals explicitly reject the dichotomy** — \"both a steel glove and a soft mitten\", and gang-fighting \"is not just about harsher penalties\". They locate the causal work in school completion, naming gymnasium graduation as \"one of the most important protective factors\".\n- **The Social Democrats agree on punishment and relocate the emphasis** — \"all focus must be on stopping gangs' recruitment of new members\" — while proposing measures (mentors, ankle monitors, risk-family programmes, U-turn programmes) that are neither purely punitive nor purely preventive.\n\nThe interesting split is therefore **inside the governing bloc**, between SD's substitution framing and L's complementarity framing. That is invisible if you read the campaign as bloc-against-bloc.\n\n**What nobody in these sources does is defend the empirical crux.** The whole dispute turns on how much *marginal* deterrence a *marginal* increase in sentence severity buys, given that certainty of apprehension is generally the stronger lever in the criminological literature. No source cites a figure, a study, or an elasticity. The SD piece is a two-minute radio item with no evidence at all; the party pages are policy prose. The disagreement is conducted entirely at the level of stance.\n\n**What would resolve it:** an estimate of the marginal effect of severity versus certainty versus prevention spending, per krona, on recruitment and on serious violence. That estimate exists in the literature well enough to be argued over, and none of these sources reaches for it.\n\n### 4.2 Healthcare: an ideological inversion and two sides talking past each other\n\nThe striking feature is that **the usual left-right polarity is reversed**. KD — a party of the right — wants to abolish the 21 regions' responsibility for healthcare and run it nationally. The Social Democrats oppose nationalisation, not wanting power handed to \"political bureaucrats in Stockholm\". A reader importing standard left-right priors gets this exactly backwards.\n\nBusch's argument, as reported, has two goals: **equity** (uniform access across the country) and **efficiency** (shorter queues), with an implementation sketch of five or six care areas run by medical professionals under national requirements, inside four years. Her framing claim is that the system cannot be patched — \"det går inte längre att lappa och laga ett trasigt system\".\n\n**The queue dispute is the sharpest empirical exchange, and the two sides are not contradicting each other.** Busch says queues are down 30 percent against 2022, after roughly 25 billion kronor of performance-based compensation targeted at specific procedures. Fredrik Lundh Sammeli (S) replies that the government maintains \"a lower level in fixed prices than 2022\". Read carefully: **one is a claim about outcomes, the other a claim about inputs in real terms.** Both can be true simultaneously. They are not competing answers to one question; they are answers to two different questions presented as a disagreement.\n\n**The most substantive objection comes from neither party.** The Doctors' Association argues that prioritising selected diagnoses may violate healthcare law, because targeting queue metrics can push the *sickest* patients behind healthier ones — the severely ill needing spinal surgery wait while healthier patients are routed to private clinics, since \"the sicker must operate where intensive care exists\". SKR warns against a \"quantitative focus\". That is a Goodhart objection: it attacks the *measure* both parties are arguing over, and it is orthogonal to the nationalisation question.\n\n**Two weaknesses on the KD side worth naming.** First, the two goals can come apart — a national system could raise equity (levelling regional variation) while lowering efficiency (longer chains of command), and nothing in the argument shows they move together. Second, **Kristersson opposes major reorganisation**, so the right bloc is split on its own sharpest healthcare proposal, and M's stated reason is that there is no clear evidence nationalisation would solve the problems.\n\n### 4.3 What the two topics share\n\nBoth disputes are **narrower than their rhetoric** and **split within blocs rather than between them**. Both have a live empirical question that neither side defends with evidence. And in both, the most interesting objection comes from a professional body rather than a party — doctors on prioritisation, and (implicitly) criminology on deterrence elasticity.\n\n### 4.4 Honest limits of this arm\n\nIt took roughly a dozen fetches and searches. It is readable, it captures strategy and coalition dynamics, and I believe it is broadly accurate. What I cannot tell you is **what I left out** — there is no list of the claims I dropped, no record of which of my assertions rest on one source versus several, and nothing here can be pointed at and disputed except by disputing whole paragraphs. If a Swedish reader thinks paragraph 4.2 is wrong, the unit of argument available to them is the paragraph.\n\n---\n\n## 5. ARM B — the structured arm\n\nThe pipeline needs an authenticated `POST`, and the Gemini key has been on free-tier remnants since the August billing block, so the extraction was executed by hand at run-3F fidelity. Artifact: [`dogfood-run-7/extraction.json`](dogfood-run-7/extraction.json). Every instrument reading below comes from [`dogfood-run-7/run_instruments.py`](dogfood-run-7/run_instruments.py), which imports the same functions the live endpoints call — `assess_completeness`, `compute_claim_badge`, `hinge_scores`, `summarize_residue_map`, `detect_weighing`, `detect_sacredness`, `stance._reported_asymmetry`. Nothing here is a description of what an instrument would say.\n\n**Extraction shape**: 5 sources · **39 claims** (32 extracted, 7 implicit) · **31 typed edges** (22 SUPPORTS, 9 ATTACKS) · 2 decomposition trees · claim types 29 empirical / 7 normative / 3 value_premise. `assess_disagreement_preservation` was not runnable — it needs the LAN-only embedder.\n\n### 5.1 The implicit premises are the arm's substantive product\n\nSeven premises no source states, of which two are load-bearing:\n\n- **`ip3` — \"A marginal increase in sentence severity changes offending behaviour more than a marginal increase in the certainty of apprehension.\"** This is what the entire crime dispute turns on, and **no source on either side asserts it.** SD's substitution argument requires it; the Liberals' complementarity argument requires its denial; neither writes it down.\n- **`ip5` — \"Uniform national access to care and shorter waiting queues move together under national governance rather than trading off against each other.\"** Busch's proposal has two goals and needs them to be compatible. The claim that they are is never made.\n\nThe rest are the ordinary enthymematic layer: `ip1` (difficult-and-costly implies less effective in practice — the step that lets SD's tractability claim pass as an effectiveness claim), `ip2` (prospective members weigh expected costs against benefits — the rational-actor premise under all deterrence talk), `ip4` (recruitment-reducing measures beat incapacitation on future crime), `ip6` (waiting-list length is a valid measure of system performance — the premise the Doctors' Association actually attacks), `ip7` (regional responsibility *causes* unequal care rather than correlating with it).\n\n**P2 confirmed on a fourth register.** The crux is stated by nobody and reconstructable by anyone who looks — political rhetoric behaves like the philosophical, evidentiary and legal registers before it.\n\n### 5.2 The completeness oracle found one copout axiom, and flattered twice\n\n| root | nodes | exposure | verdict | unsupported value premises |\n|---|---|---|---|---|\n| `h1` SD substitution thesis | 5 | 0.80 | partially_exposed | 0 |\n| `s4` S recruitment-first | 4 | 1.00 | **fully_exposed** | 0 |\n| `l1` L complementarity | 4 | 0.25 | partially_exposed | **1** |\n| `k2` KD nationalisation | 7 | 0.71 | partially_exposed | 0 |\n| `f10` M no-clear-evidence | 1 | 1.00 | **fully_exposed** | 0 |\n\nThe one hit is real and useful: **`l7` \"Society must be hard on those who commit serious crimes\"** — the Liberals' value premise, asserted and unsupported, which is exactly the object the oracle exists to surface.\n\nThe two `fully_exposed` verdicts are not. `f10` is a **single undecomposed claim** and reads as fully exposed because it has no children to be missing; `s4` reaches the same verdict with four leaves nobody has opened. This is dogfood run 4's calibration finding recurring unchanged — *`fully_exposed` at depth 0 flatters by construction* — and run 7 adds that it also flatters at depth 1 when every leaf happens to read atomic-for-now. The verdict measures *whether the frontier has been pushed down*, and says nothing about how far down it is.\n\n### 5.3 The hinge whispers, and here it cannot do otherwise\n\n| root | baseline σ | top hinge |\n|---|---|---|\n| `h1` | 0.632 | 0.022 |\n| `k2` | 0.795 | 0.024 |\n\nUnder `k2` the four children tie at 0.024, 0.024, 0.024, 0.022. Under `h1` all three tie at exactly 0.022 — including `ip3`, the crux the whole debate rests on. That is not a calibration problem to tune away: the hinge is sensitivity through the decomposition channel, and with no evidence, no votes and no answered critical questions anywhere in the tree, **every child has an identical local strength, so the instrument has nothing to be sensitive to.** Dogfood run 2 recorded \"the hinge whispers at light engagement\"; run 7 shows the floor case, where engagement is zero and the whisper carries no information at all.\n\n### 5.4 Every badge is 0.500 — and this is true of the live corpus, not just this run\n\nIn the run-7 map, 12 of 39 claims carry incoming evidence. All 12 compute to **exactly 0.500**, while displaying different terms: `h6` reads WELL-EVIDENCED, `h1` reads SUPPORTED, `l1` reads GROUNDED. The terms differ because they name the scheme *cluster*; the strengths are identical.\n\nBecause that could have been an artifact of hand data, I measured production before the outage below. **250 random live claims probed: 116 carry evidence, and 115 compute to exactly 0.500** (the other to 0.501). Term distribution over those 116: VALID 26 · WELL-EVIDENCED 21 · STRONG ANALOGY 20 · SUBSTANTIATED 15 · BEST EXPLANATION 11 · GROUNDED 8 · CREDIBLE 6 · SUPPORTED 5 · SOUND 3 · INVALID 1. Nine distinct confident-sounding labels; one number under all of them.\n\nThe cause is not a defect and the August fix is not implicated. Contributions are offsets from the neutral point, an unanswered critical question contributes zero by design, and **no CQ polarity claim in the corpus has ever received evidence** — sampled polarity claims sit at confidence 0.5 with nothing attached. So the strength layer is *correct and inert*: post-fix it computes the right answer, and the right answer is currently \"no information\" everywhere.\n\nTwo consequences worth separating.\n\n1. **The number is not yet a product.** Anything that ranks by claim strength is ranking by a constant. (The bridging feed escapes this only because it does not read the badge at all — it averages the classifier's extraction-time `r.strength`, which is its own problem, recorded in `bridging.md`.)\n2. **0.500 falls on the positive side of the badge boundary.** `strength_to_badge` assigns the high term at `>= 0.5`, so a claim nobody has examined displays as WELL-EVIDENCED in blue rather than as unevaluated. By the confession principle that is a system reporting success in the absence of any signal — the `NO DATA` gray already exists for the no-edges case and there is no equivalent for the have-edges-but-no-answers case. Reported, not fixed: it re-labels a shipped surface, which is a founder call.\n\nA boundary artifact sits next to it and is minor: `claim_3f622da496a6` returns `INVALID`/amber with `strength: 0.5`, because the displayed number is rounded to 3dp while the term is chosen from the unrounded value. The term and the number disagree at the threshold.\n\n### 5.5 P4: the instruments are blind to Swedish, cleanly demonstrated\n\nThe lexicons fired **zero** times on the run-7 claims — but zero in *both* languages, which is not a language finding, because these claims may simply not be weighing-shaped in any language. To separate the two failures I built nine sentence pairs that are weighing / sacred / reported-shaped **by construction**, in the idiom these debates actually use, and ran each lexicon over the Swedish and the English rendering:\n\n**0 of 9 Swedish · 9 of 9 English.**\n\n*Trygghet väger tyngre än den personliga integriteten* is invisible; \"Security outweighs personal privacy\" fires. *Människolivet är okränkbart* is invisible; \"Human life is inviolable\" fires. *Enligt Socialdemokraterna…* is invisible; \"According to the Social Democrats…\" fires. Weighing detection, the sacredness brake and `stance.py`'s reported-speech suppressor are all English-only, so on Swedish input the brake cannot brake and the suppressor cannot suppress. The instruments' Swedish coverage is not weak — it is nil, and it depends on a translation pass that is itself one of the four known flattening mechanisms.\n\n**Underneath the language gap sits a register gap, and the run separates them.** Ten run-7 claims are *type*-eligible (`normative` or `value_premise`) and the weighing lexicon misses all ten in English too. Three of the ten are genuine weighings in dialects the lexicon has never held:\n\n| claim | English text | dialect |\n|---|---|---|\n| `s4` | \"Stopping gangs' recruitment should be **the primary focus** of gang crime policy\" | priority-ordering |\n| `l1` | \"Combating gang criminality **requires both** punishment … **and** prevention\" | balance |\n| `l6` | \"The fight against gangs is **not solely** a matter of harsher penalties\" | negated exclusivity |\n\nRun 3F added *rank-ordering* and *superlative* to the dialect list. Swedish party rhetoric adds these three, and notably **never once reaches for \"outweighs\"** — the idiom the lexicon is built around. This is the documented register ceiling, met in a register the corpus had not yet tested.\n\n### 5.6 Two instruments had nothing to report, for opposite reasons\n\n`_reported_asymmetry` fired on **0 of 7 cross-source edges** — correct, since no claim on either side is reported speech; these are party pages and interviews speaking in their own voice. `summarize_residue_map` returns `residue_fraction: null` — correct, since typing a terminus requires a human verdict and this run has none. Both are honest silences rather than failures, and both are worth recording precisely because an instrument that returns nothing is the easiest kind to mistake for one that is working.\n\n### 5.7 P3 is unscored, the run cost the site ~100 minutes, and I got the cause wrong twice\n\n**Operational specifics are held privately until the fix is deployed.** Scoring P3 required a cross-domain similarity probe, and the only route from a cloud session is a public read endpoint. The probing was heavy enough to make that endpoint fail, and the host it runs on then wedged for roughly 100 minutes. A performance property of a live public surface, described in enough detail to reproduce while the mitigation is committed but not yet deployed, is the project's own weakness stated *as a vulnerability to manage* — which its docs-visibility boundary sends to `.private/` without a judgment call. The full account, measurements, host state and mechanism live in `.private/incidents/2026-08-17-darwin-wedge.md`, and **promotion back into this section once the fix is live is a logged founder decision**, because by then it is answered self-critique rather than a live exposure, and this project publishes those on purpose.\n\nWhat belongs here is what the run learned, which does not depend on the specifics.\n\n**P3 is NOT SCORED.** No result was obtained. The empty results from the first batch cannot be read as evidence of no cross-domain similarity: every call in it either timed out or returned an error, and the endpoint filters below a threshold anyway. Treating those as measurements would be exactly the confession-channel failure this project names — **an error path reported as a result**.\n\n**I stated the cause three times, and the first two were wrong.** First that I had taken the site down, from timing plus one plausible mechanism. Then that it was not attributable to my probing, after finding two sibling sites equally dark — a retraction resting on the claim that an application-level overload cannot darken sites I never touched, which is **false**: it holds only under memory isolation, and on a shared host a memory storm takes everything with it, which is precisely the pattern I read as exculpating. Only the third, written from host state read by a session with LAN access, is supportable: **confirmed on mechanism, unconfirmed on this incident**, with the confounder named rather than argued away.\n\nThat is the same shape as the run-6 stance overclaim, in the same week: a real but weaker measurement written up as a stronger causal story, corrected on challenge, with the accurate version available the whole time.\n\n**Two rules earned, both cheap enough that there is no excuse.**\n\n- **A dogfood probe against production is itself an intervention, and the first call is the experiment.** One probe, timed, before forty. This survives the attribution correction intact — it was bad practice whether or not it caused the outage.\n- **Before attributing an outage to your own action, test a service on the same host that you never touched.** Two requests, and it inverted the conclusion.\n\n**And a third, about this document.** The first version of this section was written while the mitigation was undeployed, on a page served publicly, and it was thorough. Thoroughness in an incident write-up is a virtue exactly until the incident is still exploitable — at which point the same paragraph is a recipe. The boundary rule existed and named this case; it was applied late rather than at authoring time.\n\n## 6. The comparison\n\n### 6.1 Scoring the registered predictions\n\n| | prediction | verdict |\n|---|---|---|\n| **P1** | unstructured wins narrative, structured wins the disagreement map and finds a crux the narrative left implicit | **half confirmed, half refuted** — see below |\n| **P2** | the crux is implicit on at least one side | **confirmed** (`ip3`, `ip5`) |\n| **P3** | cross-domain auto-connect finds essentially nothing | **not scored** (§5.7) |\n| **P4** | English lexicons fire zero on Swedish, non-zero on English | **confirmed**, 0/9 vs 9/9 |\n| **P5** | structure costs coverage | **confirmed, narrowly** |\n| **P6** | structure adds less here than on philosophical/legal debates | **partly confirmed, wrong mechanism** |\n\n**P1 is the interesting failure.** The first half held: Arm A carries coalition dynamics, strategic incentives and the inversion of left-right polarity on healthcare, none of which survive into the graph. The second half did not. Arm A **also** found the crux — §4.1 says in plain prose that nobody defends the marginal-severity-versus-certainty question, and §4.2 catches that Busch and Lundh Sammeli are answering different questions. The structured arm did not discover something the prose missed. What it did was give the crux **an address**: `ip3` is a node with an id that someone can attack, cite, or supply evidence for, whereas Arm A's version of the same insight is a clause inside a paragraph that can only be disputed wholesale.\n\nThat is a real difference and a smaller one than predicted, and it lands precisely on the load-bearing claim in `structure-versus-scale.md`: **addressability is the substrate.** Run 7 is mild confirmation of that layer and says nothing good about the layers above it.\n\n**P5, precisely.** Both arms read the same five sources, so per-source coverage was equal; the cost difference is in effort, not reach. Arm A took roughly a dozen fetches. Arm B took the same reading plus a full hand extraction, and it dropped the landscape survey's other live topics — migration, energy, EU policy — entirely, which Arm A at least situated. At pipeline speed this asymmetry mostly disappears; by hand it is large.\n\n**P6 was right about the outcome and wrong about the reason.** I predicted structure would add less because an election is closer to interest-conflict than to confusion-conflict. That is not what the decomposition found. The crime dispute bottomed out in a genuinely **empirical** crux — a criminological elasticity — which is confusion-conflict and is the wager-favourable case. The healthcare queue exchange decomposed into two claims that are **not even contradictory** (one about outcomes, one about inputs in fixed prices), which is dissolution. Both results run *against* my stated reason.\n\nStructure did nonetheless add less, for a different and more concrete reason, which is §6.2.\n\n### 6.2 The finding the run actually produced\n\n**Exactly one half of typed structure earned anything here, and it is the half that needs no human input.**\n\nThe map earned its cost: 39 addressable claims, 31 typed edges, 7 reconstructed premises, and two decomposition trees that state what each thesis *requires* rather than what it asserts. All of that is machine-produced and immediately useful.\n\nThe numeric half contributed nothing: every badge 0.500, every hinge 0.022–0.024 with ties at the crux, the residue map null. Not because the code is wrong — the strength layer was audited and fixed a day earlier — but because **all three instruments consume human judgement and this corpus has never received any.** The measurement in §5.4 shows this is corpus-wide and not a property of run 7: 115 of 116 evidenced live claims sit at the neutral point.\n\nSo the honest one-line answer to \"does typed structure earn its cost\" as of this run: **the typed map does; the typed numbers are a promissory note.** They are the part that pays off only if people show up and answer critical questions, and nobody has yet.\n\n### 6.3 Where each arm is actually better\n\n**Arm A is better at:** narrative and causal story; strategic and coalition context (that the interesting split is *inside* the governing bloc, and that the right bloc is split on its own healthcare proposal); the ideological inversion; register and tone; and — this is not a small thing — being read. It also delivered every substantive analytical insight in this document.\n\n**Arm B is better at:** addressability (a disputant can attack `ip3` without attacking a paragraph); making a thesis's *requirements* explicit, which is what the two decomposition trees are; flagging an unsupported value premise mechanically rather than by taste; and persistence — Arm A's insights die with this document, while Arm B's claims can be linked to, contradicted, and accumulated against.\n\n**Neither is better at:** knowing what was left out. Arm A has no omission ledger, and neither does a hand extraction. `POST /query` has one; this run never touched that path.\n\n### 6.4 What this changes\n\nNothing in the ontology. Three things for the build queue, in order of how much they cost:\n\n1. **Corpus-wide similarity scans are unbounded, on four call sites and two public unauthenticated routes** (§5.7). Measured: 160 MB and 2.15 s of GIL-held CPU per call, with no `LIMIT` in the query and `limit=` capping only the rows returned. Escalated, and the interim guard shipped the same day — bounded concurrency that sheds rather than queues. This is the only item that is urgent. *(Earlier wording here said \"public DoS surface, demonstrated\"; the cost is demonstrated, the site-wide outage is not attributable to it.)*\n2. **A claim with edges and no answered CQs should not display a confident badge.** The gray `NO DATA` state exists for zero edges; the have-edges-no-answers state has no visual equivalent and currently borrows the positive one. Founder call, because it re-labels a shipped surface.\n3. **The lexicons need a language dimension, and it is a different axis from the register ceiling.** Run 7 measures both independently: 0/9 across the language gap, 0/10 type-eligible across the register gap, with three new weighing dialects named. Adding Swedish patterns fixes neither alone.\n\nAnd one thing for how runs are conducted: **§5.7's rule** — a probe against production is an intervention, and the first call is the experiment.\n\n\n---\n\n## 8. Addendum, 2026-08-17: the language policy, decided after the measurement\n\n*Appended below the line, per the doc's own convention. An earlier version of this session inserted it into §1 — above the freeze line, and framed as a decision pending, when §5.5 had already measured the thing. Both errors are recorded rather than quietly fixed: editing a preregistration's prediction section after results are known is precisely the failure this project's conventions exist to prevent, and the correction is the more interesting artefact than the original.*\n\n**The founder's question came after the finding, which is the right order.** Asked whether the system should handle both languages or translate everything to English on ingest, and answering *\"do it that way this time and we can consider other options further down the line.\"*\n\nSo the policy is: **claims enter the graph in English.** What makes that decision better-grounded than it would have been yesterday is §5.5 — **0 of 9 Swedish, 9 of 9 English** on sentence pairs built to be weighing, sacred and reported-shaped by construction. The instruments' Swedish coverage is not weak, it is nil.\n\n**But the corpus already holds an argument against making the policy permanent.** `convergence-wager-red-team.md` names forced translation as the mechanism by which a pipeline loses *\"the very thing in dispute\"*, because thick evaluative concepts carry their justificatory force in their home vocabulary. That was written about rival moral frameworks; a language boundary is the same mechanism with a harder edge. The Lean analogy adds the matching warning: formal machinery cannot rule out errors in translation between the formal statement and the intended one, so **formal structure without semantic fidelity is hollow**.\n\n**The framing worth carrying: translation would be a fifth flattening mechanism, and the only one applied by policy.** The four already named — paraphrase at extraction, side-dropping at synthesis, evidential averaging in an aggregate source, stance loss at the relationship layer — happen incidentally. A translate-on-ingest rule happens deliberately, to every non-English claim, permanently. And unlike the others it is **invisible in the stored artefact**: a translated claim looks native.\n\n**This run supplies the live instances.** *Trygghet väger tyngre än den personliga integriteten* is not \"security outweighs privacy\" — *trygghet* is neither safety nor security, and the run's own §5.5 uses that pair as its first example. *Folkhemmet* and *arbetslinjen* are the same problem in the political register. These are contested-concept-layer terms, which is exactly the layer Deliberus exists to preserve.\n\n| Option | Cost | Buys |\n|---|---|---|\n| **Translate on ingest** (this run) | A fifth flattening mechanism, by policy, invisible afterwards | Every instrument works; one corpus; cross-source comparability |\n| **Store original, translate a display layer** | Per-language lexicons, or the semantic tier §5.5 already implies | Fidelity preserved; the original stays challengeable |\n| **Store both, link them** | Claim-sameness *across languages* | Makes the translation itself contestable — the project's own answer to everything else |\n\n**The third is the one consistent with the rest of the design, and it is unbuilt.** It also lands on the same blocker as the straw-man detector and the reuse lookup, which now makes three separate pieces of work waiting on claim-sameness — a concentration worth noticing on its own.\n\n**And §5.5's second half constrains the fix.** Adding Swedish patterns to the lexicons would close the language gap and leave the register gap open: the same lexicon misses **ten of ten** type-eligible claims in *English*, including three genuine weighings in dialects it has never held. Neither gap is fixed by the other, and a semantic tier addresses both where a second word-list addresses one.\n\n"}