{"path":"research/the-residual-error-taxonomy.md","content":"# The Residual Error Taxonomy — What Step-Checking Cannot Catch\n\n**Date**: 2026-08-25\n**Type**: Founder question following the QWM read (*\"What kind of errors could systematically pile up in Deliberus, given the ontology and everything we do to check every inference step?\"*), answered against the corpus's own measured incidents rather than in the abstract. Every class below is anchored to something this project has already observed in itself, which is what separates a threat taxonomy from a worry list.\n**Companion docs**: [fragile-checkers-and-the-verification-bottleneck.md](fragile-checkers-and-the-verification-bottleneck.md) and [grounded-critics-and-quarantined-imagination.md](grounded-critics-and-quarantined-imagination.md) (the week's trilogy this question follows from).\n\n---\n\n## 0. The organizing law\n\n**Checking every inference step guarantees local validity. Systematic error lives in global properties that no single step owns.**\n\nInspection catches *incoherence* — a step that contradicts its neighbors, a claim that fails its own critical questions. It cannot catch **coherent tilt** (every step bent the same way), **coherent absence** (what never became a step), or **coherent drift** (slow change no single comparison notices). A graph can be everywhere-locally-right and globally wrong, the way a map is wrong when every street was individually verified against the same distorted aerial photo.\n\nA second law rides with it: **determinism amplifies systematic error rather than cancelling it.** Random errors wash out across a large graph; a wrong choice in the deterministic strength layer repeats identically on every claim it touches. The deterministic core is the right design (it is what makes challenges land in math instead of in a persuadable judge) — but its parameters are single points with total coverage, which is why the corpus's worst measured defect (attacks counted as supports) lived there.\n\n---\n\n## 1. The eight classes\n\nOrdered roughly by how exposed we are. Each entry: the mechanism, why per-step checking passes it, the measured anchor, and the current defense state.\n\n### 1. Same-hand bias — errors that agree with each other\n\nNearly everything in the graph came from one model family, and most checkers run on that same family. N locally-plausible steps from one generator are one sample repeated: a *consistent* tilt — toward the semantic center, toward legible registers, away from advocacy's heat — passes every per-step check, because consistency is what checks test. The instrument-independence gap named this; the wiggle paper's jury finding priced it (same-model agreement is the weak signal).\n\n- **Anchors**: advocacy flattens most (disagreement preservation 0.82 vs 0.94, dogfood run 2); same-fact-opposite-use classified as SUPPORTS at 1.00 confidence (run 6 follow-up).\n- **Defense state**: *partially designed.* The two-family jury (Claude + Gemini backends) is an afternoon of plumbing; the checker-independence escalation path exists on paper. Nothing runs today.\n\n### 2. Coherent absence — asymmetric ingestion becomes asymmetric strength\n\nThe step-checker's blind spot by construction: an absent premise accumulates no attacks, so the recorded side's strength grows relative to the unrecorded side. Which sources get extracted, which registers, which languages, which *sides of a debate* — every selection at the ingestion boundary becomes a strength differential inside the graph, and **a lean caused by what got ingested is indistinguishable, in-graph, from a lean in the reasoning**. This qualifies the founder's visibility-plus-invitation stance: a visible lean is attributable to contributed reasoning *only if ingestion balance is measured*; otherwise the attribution assumption is doing unexamined work.\n\n- **Anchors**: the crux of a real debate typically unwritten on at least one side (run 3F, run 6); the implicit-premise pass has **never run on ~60% of the extracted corpus** (Temporal path only; the stream path does not call it) and produced **10 premises across the 107 arguments it did see**, under a prompt instructing it to prefer fewer — an exposure audit, not a quality verdict ([self-similar-decomposition-and-claim-ontology.md § How much was ever ATTEMPTED](self-similar-decomposition-and-claim-ontology.md)); the completeness oracle checks *within* recorded structure only; the Swedish comparison-key bug class (language selection effects reaching identity).\n- **Defense state**: *weakest of all.* The ingest-by-debate-cluster rule is the policy defense; the cross-source premise proposer and the counterweight question attack single instances. **No instrument measures what the corpus is a sample of** — per-cluster side balance, register spread, language spread. That instrument is cheap (counting, no LLM).\n- **The same class one level down, named 2026-08-27**: coherent absence also occurs *inside a single claim's decomposition*. A decomposition that drops something the parent asserted fails silently — the completeness oracle only reads what was written down, so a lossy decomposition reports the same completeness as a faithful one. The applied literature calls this dimension **coverage** and gives it two siblings (**coherence**: did the decomposition invent content the parent never asserted; **atomicity**: are separable things still fused). **The instrument now exists** — a recomposition check that rebuilds the parent from the children alone, in a separate call so no model grades its own omissions — and on its first live run it caught an invention: a decomposition that added *\"(e.g. job loss, reduced hours)\"* to a claim that never specified the form of the hardship. → [the-jhu-decomposition-line.md](the-jhu-decomposition-line.md), `deliberus/extraction/decompose_claim.py`\n\n### 3. Identity scattering — claim-sameness as a strength distorter\n\nWhen one proposition exists as several nodes, evidence and attacks scatter across the copies: each looks weakly examined, one twin can gather the supports while its attacks sit on another. Every step correct; the corpus-level accounting wrong. The converse — false merges of distinct senses — manufactures agreement in the layer whose job is preserving distinction. Systematic because re-minting is the default: decomposition performs no lookup of existing claims (measured), so scattering grows with every extraction.\n\n- **Anchors**: SIMILAR_TO at 56 edges across 4,768 claims; the always-mint decision making claim-sameness load-bearing; the polysemy safety condition (`usage_count × sense_count` coming apart).\n- **Defense state**: *known, unsolved, and named the binding constraint* — assembly theory's own hardest problem. No new action here beyond what the corpus already carries; listed because its error mode is systematic, not hygienic.\n\n### 4. House-choice multipliers — the unspecified middles of the deterministic layer\n\nChoices so natural nobody records them (summing because summing is what one does; averaging across edges; one strength for a source's mixed methodologies) sit in the layer with total coverage. Tests authored from the code cannot catch them — a test that pins the implementation cannot falsify the intent — and the defect class lives precisely in layers nobody wrote a spec for.\n\n- **Anchors**: 1,567 attacks counted as supports (the layer nobody specced); additive sibling energy paying for splitting until `necessary` semantics; conjunction vs corroboration indistinguishable; `heterogeneity` extracted and inert; the arithmetic-mean-across-edges house choice, still unexamined.\n- **Defense state**: *reactive but real.* The de-bake / confess / publish triad is the right response *once a choice is noticed*; the ontology constraint layer (approved) is where noticing gets systematized. The generator of the class — unrecorded naturalness — has no detector by definition; the countermeasure is the sorting test run on schema elements, which is exactly what found the is/ought partition.\n\n**Human twin (2026-08-27).** Class 1 is generator and checker sharing a model family. The same\nstructure holds for people: the wisdom of crowds requires errors to be **approximately\nindependent**, and correlated errors do not cancel, they accumulate — aggregation then amplifies\nsystematic error instead of cancelling random error. A bias is by definition the correlated kind;\nthat is what distinguishes it from noise. So recruiting more humans does **not** wash out shared\nbias, it raises confidence in it, and *two people from the same reading community are one detector\nsampled twice*. This promotes register diversity from good practice to the condition under which\naggregation means anything at all. See\n[cognitive-bias-codex-and-human-contribution.md](cognitive-bias-codex-and-human-contribution.md) §4.2.\n\n### 5. Dead instruments believed alive\n\nA shipped detector creates warranted complacency while silently matching nothing. Error then accumulates *behind* the green light, which is worse than having no light: vigilance ended when the instrument shipped. The three-runs law applies to checkers too — an instrument induced from one register goes blind on the next, and blindness and cleanliness report identically.\n\n- **Anchors**: the sacredness brake matched 0 of 1,678 claims as shipped; the frontier size rule was violated for its entire pre-test existence; the flaky security test that tested nothing 13.8% of the time.\n- **Defense state**: *disciplined but recurring.* The fixture rule (test data copied from reality) and measure-on-the-real-corpus-before-shipping are the standing counters; both exist because this class already fired twice. Expect it again with every new detector — the weighing eval's committed-before-detector key is the current best practice.\n\n### 6. Ratification decay — the gate that degrades while nominally intact\n\nPropose-then-ratify assumes the ratifier catches plausible-but-wrong. But machine proposals are fluent and well-formatted — the wrong ones are precisely those selected to survive surface review — the ratifier is currently one person, and the detect-but-defer result (80.7% of errors detected, followed anyway) shows detection alone does not produce override. The audit trail then becomes the false reassurance: every proposal \"was ratified,\" and the record of checking substitutes for checking. Compounding arrives when ratified errors become priors for future machine proposals — the corpus citing itself.\n\n- **Anchors**: detect-but-defer (cognitive-commons read); the ratification-tether threat entry; the ratification debt named in the supply-side analysis.\n- **Defense state**: *designed, unbuilt — and worse than that line implied.* Traced 2026-08-27: **there is no ratify action in the system.** No endpoint confirms an existing proposal; the single exception is `POST /claims/{id}/validity`, which ratifies a *staleness* flag. Everywhere else `confirmed=true` is the **default on human-authored edges**, so the flag records *a human wrote this*, never *a human checked this* — two different acts, conflated. And until a badge filter was added the same day, **nothing computational read it at all**: a human 'ratifying' changed no number anywhere. So the propose-then-ratify architecture has a propose half built five times over (terminus classifier, stance conflicts, staleness daemon, cross-source premises, claim decomposition) and a ratify half built once. The seed-flawed-proposal probe measures a gate that, for four of those five, **does not exist yet**. → [steadying-an-outside-judge.md](steadying-an-outside-judge.md) for what ratification would have to become.\n\n### 7. Frame lock-in — the reuse flywheel's shadow\n\nThe first extraction of a domain fixes its concept vocabulary and decomposition axes; later material is assimilated *into* that frame, because reuse is the whole cost-curve bet — a mapped premise is cheap to reuse and expensive to re-derive. So the better the flywheel works, the more expensive *re-framing* becomes relative to incremental contribution: a structurally enforced conservatism where every step is correct within the frame and the frame itself is never a challengeable object. This is Kuhn's normal-science dynamic rebuilt in software: the paradigm is whatever the early corpus happened to crystallize. Supersession exists for edges; **frame-level supersession — re-decomposing a whole region along a different axis, with the old frame preserved and comparable — has no mechanism**, and the reuse economics push against anyone building one ad hoc.\n\n- **Anchors**: decompose-along-an-axis-only-when-contested presumes the *axes on offer* are adequate (eight are named; the frame chooses among them at extraction time); concept sense-lifecycle states exist but concepts are the frame's *vocabulary*, not its *geometry*; the three-runs law is this class observed at taxonomy scale — and taxonomies got a confession channel while frames did not.\n- **Defense state**: *naked, and previously unnamed.* Nothing in the ontology can currently say \"this region's decomposition is one frame among possible frames.\" The cheapest first move is ontology work, not code: decide whether a frame is a representable object (the sorting test applies — two users disagreeing about a frame should be holdable).\n\n### 8. Confidence laundering through aggregation\n\nProvenance labels live on nodes; aggregation is the operation that erases origin. Strengths, feed rankings, hinge numbers, and syntheses blend machine-minted, human-asserted, and unexamined inputs into single authoritative-looking values — the aggregate inherits the record's authority while blending imagination-tier inputs. Unless every *derived* surface carries the provenance split, laundering is the default, because that is simply what aggregation does.\n\n- **Anchors**: `count_safe_summary` built and never called (`/graph/stats` publishes blended totals); the badge computing over machine-extracted edges ungated (the QWM doc's honest inversion); the gray band covering *unexamined* but not *examined-by-whom*.\n- **Defense state**: *proposed.* The provenance-split strength readout (machine-derived vs human-engaged contributions reported separately, the stance instrument's two-signal pattern) is the counter; wiring `count_safe_summary` is its trivial first step.\n\n### Noted ninth: temporal skew — listed as ruled, not naked\n\nOld, well-connected structure accumulates connectivity advantage; empirical claims decay while value claims don't; strength is age-blind by explicit ruling (staleness flags and ratified validity lapse are filtered *out* of strength on purpose). This is a systematic bias accepted with eyes open — the temporal rung's instruments (dates corpus-wide, the staleness daemon, supersession) are the containment, and revisiting the age-blindness ruling is a founder decision already parked in the temporal-rung docs. Included so the taxonomy is honest about a bias we chose.\n\n---\n\n## 2. Triage\n\n**Least defended, in order**: coherent absence (class 2 — no instrument even in design for ingestion balance), frame lock-in (class 7 — previously unnamed, no mechanism, economics push the wrong way), ratification decay (class 6 — probe designed, unbuilt). Same-hand bias (class 1) sits just behind: its counter is buildable this week.\n\n**Well-covered relative to their danger**: house choices (the triad + constraint layer), dead instruments (the fixture discipline), identity (known and load-bearing), laundering (proposal on the table), temporal (ruled).\n\nThe pattern across the triage: the naked classes are the ones whose errors live *outside the graph* — in what enters it (2), in the shape it crystallized (7), in the human gate around it (6). The well-covered classes live *inside* the graph, where the project's instrument-building instinct naturally reaches. The taxonomy's practical value is pointing the instinct at the boundary.\n\n## 2b. Where class 2 comes from on the operator's side: the over-broad refusal\n\n*Added 2026-08-27 after the founder pointed at an irony — a rule written that afternoon appeared to\nforbid the exact evidence the same document was built out of. Not a ninth class. A named **source**\nfor class 2 (coherent absence), and the one that originates with us rather than with the corpus.*\n\n### The shape\n\n**A refusal issued at one level of abstraction silently forbids things one level below it.** You\nreject a thing correctly, phrase the rejection one notch too wide, and the phrasing removes\nsomething valuable that nobody then misses — because what it removes is *work not done*, and\nnobody files a bug for that.\n\nThe instance that surfaced it: *\"the codex's organising question is useful; its 188 entries are\nnot.\"* The intended refusal was **the taxonomy as schema** — categories a user attaches to a claim.\nWhat the sentence said was that the bias findings are not useful, while **the very document it\nappeared in was built out of them** (hidden profiles, cascades, polarization, correlated error, the\nbias blind spot). The rule and the practice contradicted each other on the same page, and the author\ndid not notice.\n\n### It is a pattern, and the corrections all came from outside the instruments\n\n| Refusal, as first stated | Re-scoped to | When |\n|---|---|---|\n| The sacredness brake stops the descent | **No stop signs** — analysis may go anywhere; what changes is *proportional attunement* and heightened humility, and only a premature verdict is prohibited | founder-ruled 2026-08-25 |\n| A typed residue is where mapping ends | **No residue is a dead end** — the contextual exit is first-class and proposable from priors; the permissive zone's interior is mapped with attribution | founder-ruled 2026-08-20 / 08-24 |\n| Do not borrow assembly theory | Do not borrow its **authority** — the mechanism (construction-with-reuse) survives and is independently attested in Kauffman and Arthur | 2026-08-16 |\n| Do not import the bias taxonomy | Do not import it **as schema** — the findings are design knowledge and already load-bearing | 2026-08-27 |\n\nFour instances, one shape, and **in every case the correction came from the founder rather than\nfrom any instrument here.** That is the diagnostic fact: nothing in the apparatus looks at the rule\nlayer.\n\n### Why it produces class-2 error specifically\n\nAn **under-broad** rule fails loudly: somebody does the thing, it goes wrong, it gets caught. An\n**over-broad** rule fails silently: nobody does the useful thing, nothing goes wrong, and the\nabsence is indistinguishable from the thing never having been worth doing.\n\nThat is coherent absence with an internal cause. Class 2 is normally about what the *corpus* is a\nsample of; this is about what our *own rules* quietly excluded from it. It is strictly harder to\ndetect than the ingestion version, because an ingestion imbalance is at least countable and a\nforbidden line of work leaves no row anywhere.\n\n### The tell, and it is checkable\n\nAcross all four instances the over-broad version names **a noun**; the correct version names **a\nuse**.\n\n> *Don't import the taxonomy · don't borrow assembly theory · don't descend on sacred claims*\n> versus\n> *don't import it as schema · don't borrow its authority · don't publish a premature verdict*\n\n**A refusal whose object is a category rather than a use is over-broad until shown otherwise.** That\nis a one-line check on any rule this project writes, and it would have caught all four.\n\n### A sibling measured on ourselves: sweeps check presence, not truth (2026-08-27)\n\nTwo documentation sweeps ran in one session and **both passed cleanly while four figures in the docs\nwere stale** — sources 25 against a live 31, classified termini 2 against 3, decomposition edges \"all\n131 untagged\" against 132 of 149 with 17 tagged, and a refusal count that drifted 38→39 *inside the\nsession that counted it*.\n\n**The sweeps were not sloppy; they were aimed elsewhere.** A congruence pass asks *is this written\ndown* and *do the docs agree with each other*. Neither question can catch a claim that was **correct\non the day it was written** and has since been overtaken — the docs agree with each other perfectly\nwhile all of them are wrong together.\n\nOne of those four changed a *claim* rather than a number: `support_semantics` was described\neverywhere as **shipped and inert**, which stopped being true when the synthetic stress suite tagged\n17 edges. A reader would have concluded the splitting-inflation defence had never been exercised.\n\n**The fix is structural rather than diligence**: `scripts/graph_facts.py` prints the figures the docs\nkeep quoting, and the docs now cite the command. **A number copied into prose is a future lie with a\ndelay fuse** — right when written, stale by the next session, and invisible to every sweep that asks\nabout presence. Where a reader genuinely needs a rough sense, pin the number *and date it*; otherwise\nname the authority.\n\nRelation to the classes above: this is class 5 (**dead instruments believed alive**) applied to the\ndocumentation layer rather than the detector layer. Error accumulates behind a green light, and the\ngreen light is a passing docs sweep.\n\n### Where the failure actually lives: the headline, not the rule\n\nAn audit of every bolded refusal in the corpus (38 of them) found the corpus in **much better shape\nthan the four instances suggested** — nearly all name a use correctly (*never demand formal structure\nas input*, *never collapse quality and stance into one signal*, *never present AI-generated structure\nas authoritative*). This is not widespread rot.\n\n**One live instance surfaced, and it is diagnostic.** UX principle 9 read *\"Never let the LLM do the\nthinking\"* — which, taken alone, forbids most of what ships, since terminus classification, scheme\ndetection, stance conflicts, synthesis and decomposition proposals are all an LLM thinking. **But\nits body always carried the scope**: *scaffold, don't replace… provoke engagement, not\nrubber-stamping.* The rule was well-formed. The **headline** was not.\n\nThat is the mechanism, and it is narrower and more actionable than \"people write over-broad rules\":\n\n> **Compression drops the scope clause first, because the scope clause is the least memorable part —\n> and the compressed form is the one that travels.**\n\nAll four re-scoped instances have this shape, including the one that started it: the careful version\nexisted in the surrounding paragraphs, and the quotable sentence did not carry it. A rule is quoted\nby its headline, inherited by its headline, and violated by its headline.\n\n**The test, and it takes seconds: does the bolded part *alone* forbid something you actually do?**\nIf yes, the headline is over-broad no matter how careful the body is. Fix the headline; do not rely\non the body. (Principle 9 is now *\"Scaffold, never replace\"* — the correct scope was already sitting\none line below.)\n\n**A headline needs two tests, not one — found the same day by shipping the second failure.** The\nrewrite *\"Scaffold, never replace\"* passes the over-broad test (it forbids nothing we do) and fails\nthe **stranger test**: scaffold *what*, replace *what*? Both objects are unstated, so a reader\nwithout the context gets a metaphor and no referent. Having observed that morning that the stranger\ntest had never been aimed at a rule, the rewrite did not aim it either. So:\n\n1. **Over-broad test** — does the bolded part *alone* forbid something we actually do?\n2. **Stranger test** — does the bolded part *alone* say what it means to someone who does not hold\n   the context?\n\nA headline can pass one and fail the other, and both failures travel the same way.\n\n**And the proportionate response matters here.** Building an instrument for a failure the audit says\nis rare would itself be the over-broad move — the same error one level up. A one-line headline test\nand one rewritten principle is the right size.\n\n### Two instruments already exist and are pointed the wrong way\n\nThis is the operator-reflexivity gap in its most concrete form so far. **A rule is a claim**, and\nthis project owns unusually good machinery for claims — none of it aimed at its own governance.\n\n- **The stranger test.** Extraction Pass 2b decontextualises every claim so that someone who never\n  saw the source understands exactly what is asserted. The offending sentence passed *the author's*\n  reading, because the author held the surrounding context; it failed a stranger's. That is exactly\n  what the stranger test catches, and no rule in this corpus has ever been put through it.\n- **The separability floor.** *Decompose until nothing separable is still bundled* is the project's\n  own decomposition target. Run on that sentence it finds **three separable objects fused into one\n  refusal** — taxonomy-as-schema, findings-as-knowledge, arrangement-as-evidence — of which only the\n  first was meant.\n\nNeither needs building. Both need aiming.\n\n### The asymmetry, and its one exception\n\n**Rules should err narrow.** The instinct when writing a safety rule is the opposite — forbid\ngenerously, since the cost of permitting harm is worse than the cost of forbidding good. That\ninstinct is correct **only where the hazard is irreversible**, and this project already carries the\ntest for that (*classify errors by reversibility*). Everywhere else the over-broad rule is the more\nexpensive error, because its damage is invisible and compounds silently while the narrow rule's\ndamage announces itself.\n\nNote the relation to a neighbouring failure the corpus already names: **leveling dissolution** is a\n*reduction* that erases a distinction (*after the dissolution, does anything about what to do\nchange?*). This is a *refusal* that erases a distinction. Same casualty, different instrument — one\nargues the distinction away, the other legislates it away.\n\n## 3. What this adds to the threat model, and what it merely reorganizes\n\n**New entries in substance**: asymmetric-ingestion-becomes-asymmetric-strength (class 2's sharp form — it qualifies a founder-ratified stance, so it must be stated where that stance is stated), and frame lock-in (class 7 — the reuse flywheel's shadow, nowhere previously in the corpus). **Reorganized rather than new**: classes 1, 5, 6, 8 gather existing threat-model material under the step-checking law; classes 3, 4 restate known binding constraints in their systematic-error aspect; the ninth is a ruling, restated for honesty.\n\n## 4. Buildables (proposals — founder-paced, referenced not restated)\n\n1. **Ingestion-balance dashboard** (new): per-debate-cluster side balance, register spread, language spread — counting only, no LLM. The instrument that lets visibility-plus-invitation keep its attribution assumption honestly.\n2. **Frame-as-object ontology question** (new): run the sorting test on \"decomposition frame\"; if two users can disagree about a frame while the system keeps working, it belongs in the graph — which implies frame-level supersession as a future edge-lifecycle sibling.\n3. **Seed-flawed-proposal probe** — already designed in the ratification-tether entry; this doc adds urgency ranking only.\n4. **Two-family jury, provenance-split readout, multi-hop flattening check, mechanical-consistency eval tier** — already proposed in the trilogy docs ([fragile-checkers](fragile-checkers-and-the-verification-bottleneck.md) § 6, [grounded-critics](grounded-critics-and-quarantined-imagination.md) § 7); classes 1, 8, and 5 are their justification restated.\n\n**Cross-references**: the trilogy docs above · [the-scrutiny-gap.md](the-scrutiny-gap.md) (classes 1, 5) · [cognitive-commons-and-deliberus.md](cognitive-commons-and-deliberus.md) (class 6) · [lowering-the-cost.md](lowering-the-cost.md) + [assembly-theory-and-the-reuse-mechanism.md](assembly-theory-and-the-reuse-mechanism.md) (class 7 is their shadow side) · [incentives-analysis.md](incentives-analysis.md) § 6b (class 2 at the supply-side boundary) · [self-similar-decomposition-and-claim-ontology.md](self-similar-decomposition-and-claim-ontology.md) (class 3) · [strength-layer-audit.md](strength-layer-audit.md) + [what-belongs-in-the-ontology.md](what-belongs-in-the-ontology.md) (class 4) · [supersession-and-bitemporal-lifecycles.md](supersession-and-bitemporal-lifecycles.md) (the ninth).\n"}