{"path":"research/dogfood-run-6-israel-palestine-cross-domain.md","content":"# Dogfood Run 6 — Israel/Palestine, and the First Real Cross-Domain Auto-Connect Test\n\n*August 13, 2026. Founder brief, verbatim: \"Do a dogfooding experiment based on Israel vs Palestine. See if it auto connects to any existing claims etc.\" Findings are **J-series**, continuing F (run 1), G (run 2), H (run 3F), and the unlettered findings of runs 4 and 5.*\n\n*Method note, stated up front because it constrains everything below: the Gemini balance is again short. Measured rather than assumed — `gemini-3.5-flash` returns 429 RESOURCE_EXHAUSTED while `gemini-3-flash-preview` and `gemini-3.1-flash-lite` still answer, which is the signature of a key that has dropped to free tier (roughly twenty requests per day **per model**). That is far too little for eight passes over two texts, and enough for a handful of surgical calls. So the extraction is executed by hand at run-3F fidelity. **But the question the founder actually asked has a free and fully real answer**: auto-connect's candidate stage is pure cosine against locally stored embeddings, served by Darwin's Qwen3-Embedding-4B. No Gemini is involved in it at all. So the connection result in this document is not a simulation — it is the production mechanism, at the production threshold, against the production corpus.*\n\n## Why this topic is the right stress test\n\nEvery previous run used a debate where the parties were, in the end, addressable to each other: retribution, minimum wage, assisted dying, bullshit jobs, AI risk, psychedelics. Israel/Palestine is the case where the platform's thesis is least obviously true. If reasoning is legible all the way down, it should be legible here; if the convergence wager is going to lose anywhere, this is a good place to look for the loss.\n\nIt is also the first genuine **cross-domain** auto-connect test. Every prior cross-source result came from same-topic opposing pairs, where connection is unsurprising. The corpus currently holds 23 sources across death penalty, assisted dying, minimum wage, UBI, bullshit jobs, altruism/game theory, wisdom of crowds and a large AI/AGI cluster. Whether a new debate in an unrelated domain connects to any of it is the falsifiable core of the metacrisis structural-self-similarity claim, and it has never been measured. (This is the cousin of the three-scales experiment queued in TODO.md, run on one domain rather than three scales.)\n\n## Sources — a same-proposition adversarial pair\n\nThe two texts affirm and deny **the same proposition**, both in doctrinal legal register, both by international-law academics. This is the cleanest adversarial pairing the corpus has held.\n\n| Side | Text | Author | Words |\n|---|---|---|---|\n| **Affirms** | \"Israel's Right to Self-Defence against Hamas\", Articles of War (Lieber Institute, West Point), Israel–Hamas 2023 Symposium | Nicholas Tsagourias, Professor of International Law, Sheffield | ~1,250 (body) |\n| **Denies** | \"No, Israel Does Not Have the Right to Self-Defense In International Law Against Occupied Palestinian Territory\", Jadaliyya | Noura Erakat, Rutgers | ~2,560 |\n\nA third candidate was dropped and the reason is itself worth recording (**J0**): the JURIST piece titled \"Israel's Right to Self-Defense Under International Law\" reads as a pro-Israel title and argues the opposite case in its second paragraph. Side assignment by title is unreliable, which matters because \"make adversarial pairing a rule at ingest\" is a queued design item — any automated version of that rule must read the text, not the headline.\n\n## Registered predictions\n\nWritten and committed before a single query was run. Each is scoreable.\n\n**P1 — Cross-domain auto-connect fires substantially.** At least 15 candidate pairs at or above the production prefilter threshold of 0.65 (top-5 per new claim) against the existing corpus.\n\n**P2 — The death-penalty corpus is the strongest attractor.** The single existing source contributing the most candidate pairs will be one of the two death-penalty texts (ACLU or van den Haag), via proportionality, desert, collective punishment, state killing and the duty owed to a population in one's power.\n\n**P3 — Peak similarity lands in the 0.70–0.82 band.** High enough to be a real candidate, short of the paraphrase range.\n\n**P4 — SIMILAR_TO stays nearly empty.** Fewer than five pairs reach the 0.80 auto-link threshold. The two thresholds answer different questions, and the honest answer to \"does it auto-connect\" will differ depending on which mechanism is meant.\n\n**P5 — Zero lexical weighing detections, in a debate that is *about* weighing.** Legal-doctrinal register is the fourth register to map, after moral philosophy (\"outweigh\"), policy analysis (\"net benefits\") and rights advocacy (nothing). This case is sharper than run 2's G10: Erakat explicitly describes the law of armed conflict as \"a crude balance between humanitarian concerns on the one hand and military advantage and necessity on the other\" — an *explicitly named balancing test*, in words the v2 lexicon cannot see. If this holds it is the strongest evidence yet for the semantic detection tier.\n\n**P6 — The surviving residue is structural, not fittingness.** The two authors agree on nearly every value in play (civilians deserve protection, force must be limited, law should constrain power). They disagree about *which legal framework governs* — occupation law or jus ad bellum. Per H9, register drives residue type, and a doctrinal clash should bottom out in \"which standard applies\", the structural type.\n\n**P7 — The two texts do not join issue, and the graph should expose it.** Tsagourias rebuts a *statehood*-based argument (if Palestine is not a State, Article 2(4) and hence 51 do not apply) and answers it well. Erakat's argument is not statehood-based but *occupation*-based (an occupant cannot invoke self-defence against territory it controls, because the ICJ held the threat originates within rather than outside). Prediction: **zero directly stated claims in Tsagourias attack Erakat's central occupation premise**; the collision is reachable only through an implicit premise neither author states. This is H1 (\"the crux is stated on one side and implicit on the other\") in its strongest available form — not one implicit premise but a whole missing edge where the public debate assumes an argument is happening.\n\n**P8 — The AI-risk cluster does not connect.** Maximum cosine from any Israel/Palestine claim to any AI/AGI claim stays below 0.70. If this fails — if the AI cluster *does* connect — that is evidence for structural self-similarity across domains and a genuinely interesting result rather than a nuisance.\n\n**Honest note on which predictions were blind.** P1–P4 and P8 are about the corpus and were written without looking at it. P5–P7 were written after reading both texts but before any extraction or measurement. That is a weaker kind of prediction and is labelled as such rather than presented as equal.\n\n## What happened\n\n### Scoring\n\n| | Prediction | Result | |\n|---|---|---|---|\n| P1 | ≥15 cross-domain candidates | **4** | ✗ failed |\n| P2 | Death-penalty corpus is the strongest attractor | tied on count, second on peak | ~ partial |\n| P3 | Peak similarity 0.70–0.82 | **0.685** | ✗ failed, low |\n| P4 | Fewer than 5 pairs reach 0.80 | **0** | ✓ confirmed |\n| P5 | Zero lexical weighing detections | **0**, both texts | ✓ confirmed |\n| P6 | Residue is structural, not fittingness | structural | ✓ confirmed |\n| P7 | The texts do not join issue | decisively | ✓ confirmed |\n| P8 | AI cluster stays below 0.70 | **0.589** | ✓ confirmed |\n\nThe two failures are the valuable ones, and they failed in the same direction.\n\n### J3 — Cross-domain auto-connect barely happens, and the contrast is the finding\n\nBoth numbers come from the production mechanism: Darwin's embedder, the shipped cosine, the shipped prefilter of 0.65 with top-5 per claim, against all 1,492 embedded extracted claims in the live graph.\n\n| | Candidate pairs ≥0.65 | Peak cosine | Pairs ≥0.80 (`SIMILAR_TO`) |\n|---|---|---|---|\n| **Within the pair** (pro × con) | **402** | **0.891** | 9 |\n| **To the entire rest of the corpus** | **4** | 0.685 | **0** |\n\n140 hand-extracted claims produced four cross-domain candidates, from four distinct claims. Twenty-three existing sources, twenty-one of which produced nothing at all. Not one pair reached the auto-link threshold, so **the number of `SIMILAR_TO` edges this debate would add to the graph is zero.**\n\nSo the direct answer to the question: *yes, it auto-connects — to four claims, weakly, and to nothing at the threshold that actually creates an edge.* Meanwhile the two texts assembled into a dense debate with each other. Auto-connect is, empirically, a same-topic instrument.\n\nThe per-source ceiling makes the shape clear. Every existing source, ranked by its closest approach:\n\n```\n0.685  ordning-och-anarki (egoistic reasons for altruism)   <-- connects\n0.662  the-case-against-the-death-penalty (ACLU)            <-- connects\n0.603  ernest-van-den-haag (death penalty, pro)\n0.600  if-assisted-suicide-is-legal (Not Dead Yet)\n0.589  forging-a-new-agi-social-contract\n...\n0.448  all-ai-models-might-be-the-same\n```\n\nThe ordering is not random — the death-penalty and assisted-dying texts sit at the top of the non-connecting group, exactly where the hand analysis said the structural homologies were. The corpus's shape is visible in the numbers. It is just compressed into a band too narrow for a fixed threshold to act on.\n\n### J4 — Why: cosine cannot see structural homology, and this explains a measurement the corpus already had\n\nThe homologies are real. An occupier's duty toward a population in its power is structurally the argument Not Dead Yet makes about a vulnerable minority; proportionality in *jus ad bellum* is the same shape as desert in retribution. A reader sees these immediately. **The embedder scores them at 0.60.**\n\nTwo claims can share an argument structure while sharing almost no vocabulary, and claim-level cosine measures vocabulary. This is not a threshold that needs tuning — dropping to 0.60 would admit the homologies and also everything else, since the whole corpus sits in a 0.45–0.69 band.\n\nThis corroborates *and explains* the measurement in `lowering-the-cost.md` §6: reuse in this corpus runs through the **concept layer** (479 usages, 87% of concepts reused) and not the claim layer (48 `SIMILAR_TO` edges across 4,628 claims). That was recorded as a surprising fact. It now has a mechanism: concepts are shared vocabulary and claims are shared *structure*, and the only similarity function in the pipeline sees the first. The consequence stands as written there — concept governance ranks above claim-level dedup — and gains a second one: **if cross-domain structural connection is wanted, it will not come from embeddings.** Candidate routes are scheme-level matching (both debates run Argument-from-Weighing over an authority's duty to a population in its power) or concept-mediated bridging, and both are unbuilt.\n\nThe metacrisis structural-self-similarity claim survives this as a claim about *structure*. What died is the assumption that the shipped machinery can detect it.\n\n### J9 — The crux is unwritten on **both** sides, and this is now measured rather than argued\n\nRun 3F's H1 found that the crux of a real debate is stated on one side and implicit on the other. This pair is worse and cleaner.\n\n**The word \"occupation\", and every cognate, appears zero times in the pro text and 54 times in the con text.** Erakat's entire argument is that an occupant cannot invoke Article 51 against territory it controls. Tsagourias answers a *statehood* objection — if Palestine is not a State, Article 2(4) and therefore Article 51 do not apply — and answers it well. The two texts are not arguing with each other.\n\nRanking every pro-side claim against Erakat's pivot (the ICJ's holding that the threat \"originates within, and not outside\" occupied territory):\n\n| Rank | Cosine | Claim | Status |\n|---|---|---|---|\n| 1 | **0.776** | `ti3` — the Wall opinion's origin test does not bind | **implicit; never written** |\n| 2 | 0.705 | t18 | stated |\n| 3 | 0.689 | t36 | stated |\n\nThe best counterpart to the opposing side's decisive move is a premise the author never wrote, beating all 46 of his stated claims by a clear margin. The same holds in reverse: Erakat never states that a discrete later armed attack cannot restart the Article 51 trigger, which is precisely what her 2012 argument needs in order to reach October 7.\n\nThen the consequence, from the classifier spot-check:\n\n- Give it the implicit premise (`ti3`) against the pivot (`e71`) → **ATTACKS, strength 0.80**. The collision appears. (Stored in the live graph at 0.85; the spot-check and the stored edge differ slightly and the stored value is authoritative.)\n- Separately, the two claims that *state the same ICJ holding* — `t18` (Tsagourias) and `e69` (Erakat) — are classified **SUPPORTS at strength 1.00**, because both accurately report what the court held.\n\n> **Correction, 2026-08-14, on the founder's challenge.** The two bullets above were originally written as a single before-and-after on one pair, which said that leaving the premise unwritten made *the pivot* read as SUPPORTS 1.00. **That is not what was measured and the pairs are different.** The SUPPORTS 1.00 was measured on `t18 × e69`; `t18 × e71` was never submitted to the classifier. So \"omit the premise and the pivot enters as an agreement\" is an **inference from the ranking table, not a result** — the ranking shows `t18` as the pivot's nearest *stated* neighbour (0.705) behind the implicit `ti3` (0.776), which is suggestive and is not a classification.\n>\n> **Two further facts from the live graph, checked the same day.** No `SUPPORTS` edge between `t18` and `e69` exists; the only stored relation between them is `SIMILAR_TO`. And the pivot is correctly opposed in the graph today, since `ti3 → ATTACKS → e71` at 0.85 is stored. **Nothing inconsistent is currently in the corpus**, because the implicit-premise pass ran and did its job on this pair.\n>\n> What survives, and it is still the finding: **two claims that accurately report the same fact are classified as supporting each other at maximum confidence while their authors are in direct opposition.** That is measured, on `t18 × e69`, and it is J8 below. What does not survive is the claim that the debate's pivot was recorded as agreement. It was not.\n\nThe implicit-premise result stands on its own without the overstatement: the best counterpart to the opposing side's decisive move is a premise its author never wrote, outranking all 46 of his stated claims. That is the corpus's strongest case for implicit-premise recall being the pipeline's most important quality axis.\n\n### J8 — A third flattening mechanism: stance loss at the relationship layer\n\nThat `SUPPORTS 1.00` is not a bug in the classifier. Both claims *do* assert the same thing about what the court held — Tsagourias reporting it as \"the contrary position\" to his own argument, Erakat as the ground of hers. The propositions agree and the authors do not, and the relationship layer sees only propositions.\n\n**Scope this precisely** (tightened 2026-08-14): the finding is that the classifier *returns* `SUPPORTS 1.00` for such a pair, which is measured. It is **not** that such an edge is sitting in the graph — no `SUPPORTS` edge between `t18` and `e69` was written, and their only stored relation is `SIMILAR_TO`. The risk is real and currently unrealised, which is the right time to build the instrument rather than the wrong time to claim damage.\n\nThe missing instrument has an unusually cheap trigger available, and the corpus already holds everything it needs: these two sources are a documented adversarial pair with `ATTACKS` edges running between them. **A support edge between two claims whose sources are known to oppose each other is a checkable inconsistency**, and nothing in the system currently asks the question.\n\n### J8b — the instrument, built 2026-08-14\n\nShipped the same week the gap was named, on the founder's *\"why not build it while it's fresh\"*. `GET /graph/stance-conflicts`, module `deliberus/stance.py`, two graph queries and no model call.\n\n**What it does.** Establishes which source pairs are adversarial (at least two opposing edges between them, direction discarded since the authors are opposed either way), then pulls every *agreement* edge spanning such a pair, with both claim texts and the stored embedding cosine so a reader can judge on sight. Two signals are reported separately rather than combined, because one number would hide which term is doing the work: **near-identical** (the two authors said close to the same thing) and **reported asymmetry** (exactly one side is reporting someone else's position).\n\n**Measured against the live corpus the day it shipped.** 8 adversarial source pairs, all 8 carrying agreement edges, **84 candidates**, 1 near-identical, 4 reported-asymmetric. It is not inert, which was the thing to check first — the sacredness brake looked healthy for months while matching nothing.\n\n**It found the phenomenon in two debates besides this one, independently.** The sharpest is in the minimum-wage pair, where one source asserts in its own voice that a wage floor harms low-wage workers, and the opposing source states the same proposition as *\"Opponents of minimum wage legislation argue that…\"* on its way to rebutting it. Those two carry an agreement edge. A fourth instance sits in capital punishment, pairing \"Proponents of capital punishment commonly argue that the threat of execution deters…\" with van den Haag asserting the deterrence claim directly.\n\n**Two honest limitations, recorded rather than tuned away.**\n\n- **The near-identical band is set at 0.80 to match the auto-link threshold, and measurement says the phenomenon straddles it.** The run-6 pair sits at 0.891 and the strongest minimum-wage instance at 0.791. Retuning the threshold to capture what I already know I am looking for would be fitting the instrument to its own motivating example; the calibration question is logged for the founder instead, exactly as the hinge calibration was in run 2.\n- **The reported-speech signal misses this very case.** Tsagourias distances himself with \"took the contrary position\", which no current pattern covers. The generalizable axis is *authorial distance from a reported proposition* and it wants a semantic tier, not more strings — the same ceiling the weighing detector hit.\n\n**What it deliberately does not do**: emit verdicts, or touch an edge. Opposed authors share background facts constantly and a debate with no common ground would be the pathology, so distinguishing shared ground from same-fact-opposite-use is a judgment about *authors*, which is precisely what the graph cannot see and why this gap existed. The instrument says where to look and offers the next move; 80 of the 84 candidates are probably healthy common ground.\n\nThe corpus now has three distinct flattening mechanisms, each needing its own instrument:\n\n1. **Paraphrase flattening** at extraction — two opposing claims embedded as near-identical. Covered by `disagreement-preservation`.\n2. **Side-dropping** at synthesis — an endpoint of a conflict never appears in the answer. Covered by `conflict_coverage`.\n3. **Stance flattening** at the relationship layer — two claims that agree about a fact get a `SUPPORTS` edge although the authors deploy that fact in opposite directions. **No instrument covers this.**\n\nA fourth was named in `long-form-sources-and-meta-analysis-weighing.md` (an aggregate averaging its parts' evidential standing). The pattern is now familiar enough to state as a rule: *an instrument built for one flattening mechanism measures nothing about the others.*\n\n### J7 — What the classifier does with cross-domain candidates\n\nAll eight spot-check calls were served by `gemini-3.1-flash-lite`, the cheapest fallback, because the primary was shedding load and `gemini-3.5-flash` returned 429. So this measures the degraded path — which, under the current billing state, *is* the live path.\n\n| Pair | Cosine | Verdict | Edge? | Assessment |\n|---|---|---|---|---|\n| e69 × ACLU ICJ-2004 | 0.660 | UNRELATED 0.00 | no | **correct, and the hardest case** |\n| e22 × ACLU state killing | 0.662 | DECOMPOSES_INTO 0.70 | yes | related, but badly mistyped |\n| e23 × Swedish conscription | 0.685 | ATTACKS 0.60 | yes | false positive |\n| e27 × Swedish conscription | 0.671 | QUALIFIES 0.40 | yes | false positive |\n\nThree of four cross-domain candidates produced an edge and two are spurious. The good news is real: the corpus holds **two different ICJ rulings from 2004** — the Wall Advisory Opinion and the Avena consular-notification case — and the classifier caught that they \"concern entirely different legal cases\", refusing the single highest-risk false friend in the set. That is the LLM stage doing exactly the job the prefilter cannot.\n\nThe bad news has teeth. `DECOMPOSES_INTO` is not one relationship among five: it is the channel QEM propagates strength through and the channel the hinge score walks. A mistyped `decomposes_into` **injects a spurious parent-child relation into the strength computation across two unrelated domains**. The classifier is offered five relationship types and one refusal, with no signal that one of those types carries far more downstream weight than the others.\n\nTwo buildable consequences, neither requiring an ontology decision: raise the strength floor for `decomposes_into` specifically, and give the classifier prompt a sentence saying what that edge does to the graph.\n\n### J5 and J6 — What the deterministic instruments said\n\n**Zero weighing detections in both texts**, verified by running the shipped compiled regexes over the sources. This is the fourth register mapped, after moral philosophy (\"outweigh\"), policy analysis (\"net benefits\") and rights advocacy (nothing), and it is the sharpest case yet — because Erakat *names the balancing test in so many words*: the laws of armed conflict rest on \"a crude balance between humanitarian concerns on the one hand and military advantage and necessity on the other.\" A text that states its balancing test explicitly, in a debate whose central legal standard is proportionality, is invisible to a lexicon built to catch weighings. The semantic tier is no longer a nice-to-have.\n\n**Zero sacredness markers**, in the conflict the world treats as maximally sacred. The brake would not fire on either text. This bears directly on the open founder question about whether the sacredness brake is overrated: the evidence here is that argued legal prose about a sacred conflict contains no protected-value language at all, because the sacred content lives outside the argued register entirely. A brake keyed to lexical absolutes will fire on undergraduate moral philosophy and stay silent on Jerusalem. That is an argument for the re-scope, though not for removal — the brake protects *authored* input, and this run had none.\n\n### J10 — What the wager got, on the hardest topic available\n\nThe decomposition shrank the disagreement, and it shrank it a long way.\n\nErakat concedes, in her own text, that \"Israel has the right to protect itself and its citizens from attacks by Palestinians who reside in the occupied territories\" (`e13`), and separately that \"this is not to say that Israel cannot defend itself\" (`e40`). Tsagourias concedes that the excuse holds \"only if the self-defence action stays within the confines of self-defence\" (`t44`). Both hold that force in war is limited and that civilians are owed protection; the cross-source pass drew five `SUPPORTS` edges between them, including a clean one between \"defence is not punishment\" and Nuremberg's \"not for revenge or a lust to kill\".\n\nWhat survives is **structural**: whether occupation law operates as *lex specialis* displacing the Article 51 trigger, or whether *jus ad bellum* is a free-standing register that occupation does not touch. That is a question about which rule-system governs — the same residue type run 3F predicted for a worldview clash, arrived at from a doctrinal one.\n\n**The guard, stated explicitly because this is exactly where a reader would over-claim.** Two international lawyers converging on \"protect civilians, limit force, dispute the governing regime\" says nothing about whether the underlying conflict is a comprehension problem. It is not. The founding conviction is careful here: legibility *dissolves* confusion-conflict and only *clarifies* interest-conflict, and someone who understands you perfectly and wants your land was never a comprehension problem. What this run shows is narrower and still worth having: **the legal argument between these two texts is smaller than its reputation, and the graph can show exactly where it starts.**\n\n### J11 — The disagreement-preservation instrument gets a ceiling\n\nTwenty-two hand-drawn cross-source edges, profile **14 attacks / 3 qualifies / 5 supports** — attack-dominant with more bridging than run 3F's batch 1 managed. Seventeen conflict edges scored, **zero flattened, score 1.00**.\n\nThat number is not comparable to the pipeline scores of 0.818–0.942 and should never be quoted as if it were: a careful manual extractor does not flatten, so 1.00 is close to what the ceiling *should* look like. Its value is as a reference point the instrument did not previously have. If pipeline extractions land at 0.82–0.94 and the ceiling on a maximally charged debate is 1.00, then that 0.06–0.18 gap is attributable to the model rather than to the topic — which is exactly the decomposition the Flattening Eval in the funding materials proposes to make.\n\n### J2 — A third taxonomy with the same shape of hole\n\nRoughly a third of the claims in both texts are neither observations, nor causal hypotheses, nor predictions, nor counterfactuals, nor feasibility claims, nor normative assertions. They are **doctrinal-interpretive**: assertions that a legal proposition follows from a text. \"Article 2(4)'s wording does not say that force is prohibited except in self-defence\" is an observation about a document; \"therefore self-defence is not an exception to it\" is not, and `epistemic_status` has no slot for it. Both artifacts carry a `taxonomy_strain` flag on the affected claims.\n\nThis is the third taxonomy in three runs to show the same shape of hole — run 5 found a **claim-type** gap (the speaker-attitude hedge), the deep-descent experiment found a **residue-type** gap (an is-bedrock terminus the four ought-biased types cannot name), and this run finds an **epistemic-status** gap. Three different taxonomies, each induced from the register that produced it, each missing the slot a new register needs. The pattern is worth a founder decision of its own: whether these taxonomies should grow a slot each time, or whether the recurrence means they should carry an explicit *unclassified* value rather than forcing a nearest fit.\n\n### J0, J1, J12 — Three things about getting text into the system\n\n**J0**: the JURIST piece titled \"Israel's Right to Self-Defense Under International Law\" argues the opposite of what its title suggests, from its second paragraph. Any automated version of the queued \"adversarial pairing at ingest\" rule must read the text.\n\n**J1**: trafilatura returned 1,935 words for the Lieber page of which roughly 685 are the RELATED POSTS navigation tail — post titles, author names and dates that a scout pass would read as argument. About a third of that extraction is junk that would have gone to the LLM.\n\n**J12**: the fetcher was **causing** its own blocks. Measured A/B, one variable at a time:\n\n| Site | httpx + honest UA | httpx + spoofed Chrome | curl + honest UA |\n|---|---|---|---|\n| Wikipedia | **200** | 403 | 200 |\n| HRW | 403 | 403 | **200** |\n| B'Tselem | 429 | 429 | 429 |\n\nWikimedia's edge returns an explicit robot-policy 403 to clients that claim to be browsers and are not, so the shipped Chrome-spoofing header set was locking the platform out of a source class the corpus already contains. The fix shipped in this run is to **identify honestly first and fall back to a browser UA only on a bot-gate status**, with three regression tests — including one asserting that a genuine 500 does *not* trigger a retry, since re-requesting under a new identity would only double the load on a failing origin. Wikipedia now fetches 38,107 words through the platform's own function.\n\nThe other two rows are not header problems and should not be re-litigated as such: HRW blocks httpx's TLS fingerprint (curl passes with identical headers), and B'Tselem blocks the IP regardless of client. Both are recorded in the code comment so the next session does not repeat the A/B.\n\n## What a non-specialist would find remarkable\n\nEverything above is written for people who already know what the system does. This section is not. It is the part of this run worth telling anyone.\n\n**In a real argument, the most important things are usually the ones nobody says out loud.**\n\nThat sentence sounds like a truism. Here it is a measurement, three times over.\n\n**First: one word, counted.** Two law professors published on the same question in the same period, one arguing Israel has a right of self-defence against Hamas and one arguing it does not. The second author's entire case rests on the territory being *occupied* — she uses the word 54 times. The first author never uses it once. Not rarely: never. He is answering a different question, carefully and well, and the two texts do not meet. On a subject where everyone assumes the disagreement is total, part of it turns out to be simply absent.\n\n**Second: both sides left out their own decisive step.** Each author's argument depends on something they never wrote down, because to them it was too obvious to say. One needs it to be true that a fresh attack cannot restart the legal clock once an occupation exists. The other needs it to be true that the occupation is irrelevant to whether the clock starts at all. Neither states their version. That is the thing they actually disagree about, and it appears in neither text — which means anybody reading both would come away thinking the disagreement is somewhere else entirely.\n\n**Third: nothing was called sacred.** This is the conflict the world treats as the most sacred on earth. Between the two texts there is not one word of sacred language — nothing described as sacred, inviolable, non-negotiable, or beyond price. The deepest commitments simply do not enter the argument. They sit underneath it, unspoken, exactly like the crux does.\n\nThree different findings, one shape: **what is on the page is the part that was easy to say.**\n\nAnd this generalises past law. Most arguments that go nowhere — between colleagues, between housemates, between countries — go nowhere for this reason. The sentence each side would need to hear is the one neither side has said, because each assumes it goes without saying.\n\n**The honest part, which is the reason to trust the rest.** When the software looked at the two sentences where this disagreement actually pivots, it recorded them as *agreeing*. Both sentences accurately describe the same court ruling; one author cites it to accept it and the other to reject it, and the software could not see the difference. We are publishing that because a tool that summarises disagreements can quietly smooth them, and the only defence against that is measuring it out loud rather than hoping.\n\n**And one thing this does not mean.** Two lawyers converging on \"civilians must be protected and force must be limited\" says nothing about whether the underlying conflict is a misunderstanding. It is not. Someone who understands your position perfectly and still wants your land was never confused. What the run shows is narrower and still worth having: the *legal* argument between these two texts is smaller and more specific than its public version, and it is now possible to point at exactly where it begins.\n\n## What happened when it went into the live graph\n\nBoth texts are now on deliberus.com, stored through the production path rather than by hand-written database writes: 140 claims, 98 within-text relationships, 10 contested concepts, 22 cross-source edges. Three things showed up only once it was live.\n\n**The concept layer bridged what the claim layer could not.** Both extractions independently produced *self-defence* and *armed attack* as contested concepts, so the two texts now meet at the concept nodes while their opposing claims stay separate — run 3F's H10 (\"the term merges, the dispute does not\") reproduced without being aimed at. This is also the constructive half of J4: the corpus's cross-domain connective tissue is concepts, and here it worked on the first try.\n\n**The auto-link result held exactly.** Eight `SIMILAR_TO` edges were created, all of them between the two new texts, and zero to the other 23 sources. The corpus went from 48 such edges to 56. The prediction and the live behaviour agree.\n\n**J13 — dogfood run 1's F6 split-brain recurred in a new caller, and that is the finding.** The first write attempt stored both texts to the graph and only one to Postgres, dying on \"attached to a different loop\" — the pooled asyncpg engine binds connections to the loop that created them, and a second `asyncio.run()` reuses a connection from a dead loop. July's fix corrected the SSE generator that had the bug; it did not make the engine safe for the *next* caller. The hazard is a property of the shared engine, so every new caller re-earns it, and the failure lands after the graph write has already succeeded — which is precisely the shape that makes it look like success. **Generalising the fix one level up is the real work**: either the engine refuses a foreign loop with a clear error, or a single documented helper owns the loop and callers cannot get it wrong.\n\n**J14 — a hand-extraction fidelity miss worth recording against the ceiling claim.** Every argument scheme name I wrote by hand was plausible and wrong: `argument_from_verbal_classification` where the taxonomy has `verbal_classification`, `argument_from_rule_application` where it has `established_rule`, and `reductio`, which does not exist in the taxonomy at all. Thirteen of fourteen names had to be remapped before storage. The pipeline cannot make this mistake, because its scheme field is a constrained enum. So the manual ceiling is not uniformly above the pipeline: on free-text judgement it is higher, and on anything the schema constrains it is *lower*, because a schema is a competence a careful reader does not have. Any future run-3F-style comparison should score scheme naming as a category where the cheap path wins by construction.\n\n## Artifacts\n\n- `dogfood-run-6/tsagourias-israel-right-of-self-defence.json` — 46 claims, 3 implicit premises, 33 relationships, 4 contested concepts\n- `dogfood-run-6/erakat-no-right-of-self-defence.json` — 88 claims, 3 implicit premises, 65 relationships, 6 contested concepts\n- `dogfood-run-6/live-claim-ids.json` — the mapping from artifact ids to the live graph's claim ids, so every finding above can be checked against deliberus.com\n- `dogfood-run-6/cross-source-israel-palestine-self-defence.json` — 22 edges, profile 14/3/5, preservation 1.00\n- `dogfood-run-6/auto-connect-candidates.json` — the 4 real candidates with their cosines\n- `dogfood-run-6/classifier-spotcheck.json` — 8 production classifier verdicts\n\n## Where these findings landed\n\nJ4 (structural kinship invisible to claim-level cosine) and J14 (a constrained enum outperforming a careful reader on scheme naming) turned out to be the two halves of a larger question the founder asked next: does the nativism dispute — built-in structure against search — apply to this project's own architecture? It does, and the answer reframes the bet away from capability and onto persistence and contestability. See [structure-versus-scale.md](structure-versus-scale.md) and [nativism thread and false dissolution](nativism-thread-and-false-dissolution.md).\n\n## What this run says about the next one\n\nThe comparison protocol from run 3F still cannot execute, and this run sharpens what it should measure when credit returns. The headline test is unchanged — implicit-premise recall — but it now has a specific, scoreable form on this pair: **does the pipeline surface `ti1`/`ti3` from a text that never once uses the word \"occupation\"?** If it does not, the debate's central collision will enter the graph as a `SUPPORTS` edge, and that failure is now demonstrated rather than hypothesised.\n\nTwo items graduate from this run to the buildable list: an instrument for stance flattening, and a strength floor plus prompt warning for `decomposes_into` in auto-connect. One graduates to the founder's list: whether the three taxonomy gaps found in three consecutive runs want three new slots or one explicit *unclassified* value.\n\n"}