{"path":"research/slicing-stability-first-run.md","content":"# Slicing Stability, First Run: Does the Score Survive the Filing?\n\n**Date**: 2026-08-22 · **Occasion**: open problem #1's own text promised this study as \"our planned defense, not yet run\" — the founder asked for the most pressing backlog item, and an unrun promise on the page shown to first outside readers ranked top among the executable ones. · **Instrument**: `deliberus/slicing_stability.py` + `scripts/slicing_stability_study.py` — pure recomputation through the shipped badge machinery under alternative filings; the graph is never written. · **Design source**: [evidence-division-and-the-foundation.md](evidence-division-and-the-foundation.md) §5 (\"route the same corpus under k reasonable policies, compare badge distributions — measurable, and nobody has measured it\").\n\n## The attack, restated\n\n*\"Your scores reflect your filing choices, not the evidence.\"* Assembly theory took exactly this hit — \"your number is an artifact of how you sliced things\" — and had no answer, because it never ran the test on itself. This is us running it on ourselves.\n\n## The four filings\n\nEach variant is a choice a reasonable person could have made with the same material:\n\n| Filing | The person it represents |\n|---|---|\n| **current** | the filing as it stands (recursive badge, shipped semantics) |\n| **flat** | the person who never decomposed: every child's evidence filed directly on the whole, no parts |\n| **joint_flip** | the person who filed the same tree but read the joins the other way — untagged parts as jointly required, tagged-necessary parts as merely corroborative (the stool-vs-pile question, answered oppositely) |\n| **fair_share** | the person who counted shared evidence once across siblings (the built, unadopted discount) |\n\n## The result: the score survives, and the exceptions name their own cause\n\n**87 decomposed mothers** (every claim with live decomposition children in the graph, synthetic and real alike), each recomputed under all four filings:\n\n- **Verdict stability: 90.8%** — for 79 of 87 mothers, every filing that produces a verdict produces the *same* verdict.\n- **Median strength spread: 0.000.** The typical decomposed claim's number does not move at all under refiling.\n- **Exactly two claims are numerically sensitive** (spread ≥ 0.05), and both for the **same reason**: `Rhetorical justifications for executions` (spread 0.431 — strength 0.491 under the current reading, 0.060 under flipped joins) and `Urban insulation from minimum wage` (spread 0.222). Both are real corpus claims whose decomposition edges carry **no recorded join interpretation** — so the additive and the jointly-required readings are both live, and they disagree substantially. The sensitivity is not diffuse filing-arbitrariness; it is one unrecorded choice, localized to two trees, with a known remedy: record the interpretation (`support_interpretation`, shipped 2026-08-17) on exactly those edges.\n- **Verdict-vs-silence is the dominant non-effect**: under the flat filing, 86 of 87 mothers go mute (NO DATA / UNEXAMINED) rather than disagreeing — at the corpus's current stage (critical questions largely unanswered), **decomposition is what gives the graph a voice at all**: the count-based child signals speak where a flat filing has nothing scored to say. A rival filing that cannot speak is not a rival verdict.\n- **Fair-share changed nothing anywhere measured** — expected, and itself evidence: the one known shared-evidence case (the fasting meta-analysis) was dissolved by evidence decomposition plus edge supersession earlier the same week, which is the founder's most-apparent-sharing-dissolves correction visible in an instrument.\n\n## What this does and does not establish\n\n**It establishes** that at the current corpus, the published strengths are not artifacts of filing choices — the attack's premise fails on measurement for 98% of decomposed claims, and the residual 2% is attributable, flaggable, and fixable rather than mysterious. That is the \"visible sensitivity flag\" the open problem asked for, delivered as a list with causes rather than a per-page banner (the banner becomes worth building when the list stops being two items long).\n\n**It does not establish** stability under *future* engagement: the corpus's critical questions are largely unanswered, so most strengths sit near neutral, and spreads will grow as real answers move numbers off 0.5 — which is why the study ships as a **re-runnable instrument**, not a one-time certificate. Nor does it yet cover per-piece evidence refiling (moving one doctor's note between a part and the whole, rather than all-or-nothing) — the flat filing is that family's extreme point; the finer-grained members join when the corpus has enough per-piece evidence to make them meaningful. And the standing scope-caution applies: every one of these 87 mothers is machine- or operator-filed; no interested party has yet filed anything adversarially.\n\n## Cross-references\n\n[evidence-division-and-the-foundation.md](evidence-division-and-the-foundation.md) §5 (the falsifier this operationalizes) · [decomposition-axes.md](decomposition-axes.md) (the boundary-choice attack on assembly theory that motivated the test) · [strength-layer-audit.md](strength-layer-audit.md) + [support-semantics-and-the-conjunction-problem.md](support-semantics-and-the-conjunction-problem.md) (the join-interpretation machinery the two sensitive claims need) · [synthetic-stress-suite-and-ontology-reflection.md](synthetic-stress-suite-and-ontology-reflection.md) (the week's strength-layer settling that made this measurement meaningful) · `../open-problems.md` #1 (the promise this fulfils)\n"}