{"path":"research/cross-source-premise-run.md","content":"# The Cross-Source Premise Run — Pre-Registration and Results\n\n**Date**: 2026-08-27 · **Status**: pre-registration written **before** the pass ran; results appended below. · **Pass**: `deliberus/extraction/cross_source_premises.py`, operator-invoked via `scripts/propose_cross_source_premises.py`, propose-only (writes nothing without `--write`).\n\n## The thesis this tests, stated so it can fail\n\nThe founder's, phrased for testability: **the premises that decide a disagreement are systematically the ones nobody wrote down, and a machine can surface them ahead of anyone contesting them.**\n\nThree separable parts, and this run touches two:\n\n1. **Systematicity** — the missing premises are not random gaps; they cluster where the decision is made (the crux, the classification of the situation, the deleted counterweight in a weighing). *Partially tested here.*\n2. **Recoverability** — a machine can surface them. **This is what the run tests**, and specifically for the class a within-source pass cannot reach: a premise invisible from inside one text because inside that text it is not missing, it is assumed. Only the opposed pair makes one side's silence legible against the other's insistence.\n3. **Ahead of contestation** — they can be surfaced before any human objects, which is what makes the mechanism scale past a decomposer tier that the engagement gradient puts at a fraction of one percent. *Not tested here; it is an architecture property, not a measurement.*\n\n## Pre-registered outcomes (committed before the pass ran)\n\n**Positive**: a proposed premise that (a) one side's argument requires, (b) the other side would deny, and (c) neither states explicitly.\n\n**Null**, and each is a different failure:\n- **Generic background** — \"words have meanings\", \"the topic matters\". Recoverable but not load-bearing.\n- **Restatement** — a premise one side already asserts explicitly. The necessity filter failed.\n- **Uncontested common ground** — something both sides would grant. Real but not the crux.\n\n**Positive control**: the Israel/Palestine pair, where the corpus already measured the crux unwritten on **both** sides and identified the classificatory premise by counting alone (one side uses \"occupation\" 54 times, the other never). **If the pass misses that pair, the pass is weak regardless of what it finds elsewhere.**\n\n**Registered prediction**: the *classificatory* class (one side's frame word absent from the other) is the most likely hit. The *weighing-counterweight* class is the least likely, because this pass reads stored claims rather than the weighing layer.\n\n## Scope\n\nFour designed adversarial pairs. The graph holds 13 unordered cross-source `ATTACKS` pairs, but nine are incidental single-edge overlaps between AI-topic sources rather than opposed treatments of one question.\n\n| Pair | Register |\n|---|---|\n| Israel/Palestine self-defence | legal-doctrinal (positive control) |\n| Capital punishment | legal-scholarly vs advocacy |\n| Assisted dying | advocacy vs advocacy |\n| Minimum wage | economics reference vs policy journalism |\n\n## Results (2026-08-27)\n\n**Four pairs, four calls, 7 minutes, $0.64 on the Claude backend. Ten premises proposed, nothing written to the graph.** A fifth call was wasted on a truncated source id and returned *\"one side has no extracted claims\"* — the pass confessing rather than silently returning nothing, which is the confession principle working at the smallest possible scale.\n\n### Scored against the pre-registration\n\n| Pre-registered outcome | Result |\n|---|---|\n| **Positive control hits** | ✓ **and exceeded** — see below |\n| Generic background | **zero** |\n| Restatement of an explicit claim | **zero**, and one was explicitly declined |\n| Uncontested common ground | **zero** |\n| Prediction: classificatory class hits first | ✓ |\n| Prediction: weighing-counterweight class least likely | ✗ **wrong**, and wrong in the useful direction |\n\n**The positive control did better than its own prediction.** It recovered the classificatory premise the corpus had found by *counting* (occupation: 54 uses on one side, zero on the other) — *\"Source B never once mentions occupation status… That silence is itself the tell\"* — and then found a **second** premise nobody had identified: whether self-defence law is autonomous from occupation law, which is a question about *legal hierarchy* rather than classification. It also **rejected the surface dispute as not the crux**: both texts argue explicitly about whether a non-state actor can commit an armed attack, so *\"the real unwritten fault line sits one level up.\"*\n\n**Two nulls were actively avoided rather than merely absent.** On the minimum-wage pair the pass wrote: *\"I treated the more famous 'zero-job-loss standard' framing as too explicit to count, since Source B's claims already name and reject it directly — that premise is effectively already on the table, not unwritten.\"* That is the necessity filter refusing a restatement, unprompted.\n\n### The premise classes that emerged, and this is the transferable result\n\nNine distinct kinds across ten premises. Only the first was predicted:\n\n| Class | Instance | Pairs |\n|---|---|---|\n| **Classificatory** | occupation reclassifies violence as internal disorder | Israel/Palestine |\n| **Normative hierarchy** | which body of law governs when two conflict | Israel/Palestine |\n| **Unit of moral analysis** | judge the individual verdict or the system's aggregate output | capital punishment |\n| **Aggregability of irrevocable harm** | can irrevocable harm be offset by aggregate benefit at all | capital punishment |\n| **Sufficiency of a justification type** | does retribution alone justify punishment | capital punishment |\n| **Detectability** | can any paper safeguard see private coercion | assisted dying |\n| **Reference class** | does Oregon predict the UK bill; do small hikes predict a big one | **two pairs** |\n| **Burden of proof under uncertainty** | presume harm or presume valid choice when unverifiable | assisted dying |\n| **Mechanism model** | is the low-wage market competitive or suppressed | minimum wage |\n\n**The reference-class premise appearing in two unrelated debates is the most reusable finding** — it is a recurring class the taxonomy does not have, and it has a mechanical shape: one side generalises from a studied instance, the other insists its case is distinguishable, and neither argues for the transfer.\n\n**And the pattern across all ten is sharper than any one of them: every premise is frame-level, not first-order.** None is a missing fact. Each decides *which facts count* — the unit of evaluation, the reference class, the burden of proof, the governing body of law, the market model. The capital-punishment reasoning says it in its own words: *\"All three premises are decision points about the FRAME of evaluation.\"* That is a claim about **where the unsaid lives**, and it is stronger than the recoverability result it came packaged with.\n\n### What this does and does not establish\n\n**Establishes**: the recoverability half of the thesis, for the class a within-source pass cannot reach by construction. Ten premises, zero junk, on four pairs, for under a dollar.\n\n**Does not establish**: that the premises are *load-bearing* — that needs a human read, and the founder's is the one that counts. Nor does it establish the ahead-of-contestation half, which is an architecture property rather than a measurement. And n=4, all in registers the corpus already holds; the three-runs law says expect the class list to break on the next register.\n\n**Open decision**: whether to store the ten as proposal-grade claims (`--write` stores them `confirmed=false`, provenance-labelled, asserted by nobody). Ten proposals entering the graph is a real change to what the corpus contains, so it is the founder's call and nothing was written.\n\n\n\n## Stored, and one thing the second run revealed (2026-08-27)\n\nFounder authorised storage. **Eleven proposal-grade claims and 50 edges (23 supports, 27 attacks) are now in the graph**, all `confirmed=false`, `provenance='cross_source_premise'`, `claim_kind='implicit'`.\n\n**Eleven, not the ten the dry run showed.** The minimum-wage pair returned two premises on the first pass and three on the second, from identical inputs. That is **extraction non-determinism measured on this pass** — the mechanical-consistency tier the wiggle reading flagged as untested, appearing unbidden. Worth recording rather than smoothing: a decomposition method that returns a different set on re-run has a stability number, and nobody has taken it.\n\n**A guard was made deliberate before storing.** These proposals were already excluded from badge computation, but for two accidental reasons — the edges carry no `scheme`, and `strength = 0.5` contributes zero energy — and no deliberate one. The badge query now filters on `confirmed` explicitly. Measured before the change: 665 edges already carried `confirmed=false` and **zero** of them were scheme-bearing, so the filter was a provable no-op on the existing graph and becomes load-bearing the moment a proposal gains a scheme. Pinned by a regression test. This is the third instance of the \"satisfied by accident\" shape (after the `REPORTS` link and `count_safe_summary`), and the first one caught *before* it could bite.\n"}