{"path":"research/derivability-of-missing-considerations.md","content":"# How Derivable Is a Missing Consideration? — ten real arguments, one ladder\n\n**Date**: 2026-09-02 · **Type**: founder-challenged claim tested on real material · **Status**:\nfinding, unruled. Founder verbatim: *\"'since a genuinely missing consideration is by definition not\nderivable from the sentence' — Now I'm feeling dubious about this. Reason about ten different random\nreal example arguments found online.\"*\n\n**Verdict up front: the founder's doubt is correct and the sentence is withdrawn.** A missing\nconsideration is not derivable from the sentence *alone*; almost all of them are derivable from the\nsentence **plus** what its words mean, **plus** the critical questions of its inference type, **plus**\nordinary domain knowledge, **plus** already-articulated alternative frames — every one of which a\nfrontier model holds. The residue that no machine can reach is real but **small, and it is mostly\nnot considerations at all — it is evidence** (what a thing costs from where a particular person\nstands, a local fact nobody wrote down). This calibrates the *size* of the two human-reserved error\nclasses; it does not touch their existence. **⚠ Read § 10 before quoting the zero**: the grader was a\nmodel, and a model cannot see its own corpus's absences, so the tally bounds the reserved classes\nfrom one side only.\n\n---\n\n## 1. The ladder — six rungs of \"derivable\"\n\nThe word \"derivable\" hid a gradient. Naming the rungs is what makes the test runnable:\n\n| Rung | Derivable from… | Who reaches it | Example |\n|---|---|---|---|\n| **0 — literal** | the sentence's own words | the shipped `check_coverage` (parent-vs-parts) | \"credible as a labor economist\" commits to *labor economist* |\n| **1 — conceptual** | what the predicate means | any competent reader; a model trivially | \"credible expert\" → publications, position, field match |\n| **2 — scheme** | the inference type's critical questions | the shipped CQ scaffolding | consequences → *countervailing consequences?*; causal → *alternative cause?* |\n| **3 — domain** | what a practitioner in the field knows | a frontier model, most of the time | the Sydney Diet Heart trial's trans-fat confound |\n| **4 — lens** | a different, already-articulated frame | a model *prompted with that lens*; a human who lives in it | the cliodynamic frame on a narrative history |\n| **5 — corpus-absent** | nothing the machine can read | a human, and only some humans | what SiS placement feels like from inside |\n\nRung 0 is what my sentence meant by \"derivable from the sentence.\" Rungs 1–4 are what the founder's\ndoubt was pointing at. Rung 5 is the human-reserved class as the harness conviction actually words it\n— *\"the machine cannot know what its corpus silently omits\"* — and the test below is about how\nmuch lands there.\n\n## 2. The ten arguments (plus three), each with its strongest missing considerations, each rung-graded\n\nAll fetched 2026-09-02; the rung grading is my judgment and is the thing to argue with. \"Missing\"\nmeans a consideration a strong critic would raise that the argument does not address.\n\n**1. r/jobs (Aug 2026)** — *Low unemployment masks a broken market; either jobs return or we need a\nmild recession.* Missing: the household survey does not depend on benefit rolls, so \"they drop off\nthe counts\" is a methodology error (**3**); U-6 already counts the discouraged and underemployed\n(**3**); who bears a \"mild recession\" — the countervailing-consequences question (**2**); the\nframe's assumption that a job market has one shape — gig work as form-change rather than\nconcealment (**4**, the labour-form lens). Rung 5: none.\n\n**2. r/AmItheAsshole (Aug 2026)** — *A degree is an investment, so I fund only a practical major.*\nMissing: the arguer's own career (teacher, healthcare) fails the return test he applies — a\nreductio a commenter found in seconds (**3**, a consistency check on the arguer's premises); STEM\ngraduates are also unemployed in this market, so the practical/impractical split is empirical and\ncontested (**3**); the wife co-owns the money (**3**); *investment* versus *gift* is the frame the\nwife and family argue from (**4**). Rung 5: the daughter's actual talent and the family's dynamics\n— **evidence**, not a consideration.\n\n**3. Nigerian Presidency on the fuel subsidy (Aug 2026)** — *Restoring the subsidy would sabotage\nlocal refining, crush private-sector confidence, and drain foreign exchange.* Missing: the\ndistributional cost of removal on households — the very question the opponent raises and the\nstatement never answers (**2**); targeted versus blanket subsidy as a third option (**3**); the\n\"U-turn\" charge is ad hominem by inconsistency and bears on the policy only through its own\ncritical question, *does the inconsistency bear on the truth?* (**2**); what actually happened to\nthe savings — a checkable public-record question (**3**). Rung 5: the household-level reality in\nNigerian cities — thinly documented in text, so partly corpus-absent: **the tail**.\n\n**4. \"The Seed Oil Bible\" (Substack, Apr 2026)** — *Chronic high-linoleic exposure remodels tissue\ninto an oxidation-prone, less resilient state.* Missing: the Sydney trial's margarine contained\ntrans fats, a known confound the piece does not mention (**3**); large cohorts associate linoleic\nacid with *lower* mortality (**3**); \"tissue resilience\" is defined so that no lipid or inflammatory\nmarker can refute it — *what would count as evidence against?* (**2**); the author's own move\n(\"we are asking a different question\") is a frame shift, and its unfalsifiability is visible from\nthe mainstream frame it walked away from (**4**). Rung 5: none — every consideration is public.\n\n**5. Ben Nadel, InVision (2020)** — *Merging microservices back into the monolith right-sizes my\nteam's domain.* Missing: staffing rather than architecture as the remedy for a shrinking legacy\nteam (**3**); n=1 and the generalization CQ (**2**); the modern-platform teams' cost of the merge —\nanother stakeholder's frame (**4**). Rung 5: InVision-internal facts — evidence.\n\n**6. Diane Lee, \"Remote work is destroying productivity\" (Mar 2026)** — *At scale, remote work\nerodes focus, coordination and culture; force the return.* Missing: the productivity RCTs, where\nhybrid measures neutral-to-positive (**3**); selection — who leaves under a forced return, the\nworkers with options (**3**); the disability-employment gain, which a same-side author (Strain,\nbelow) names (**3**); \"motion is not momentum\" is a claim needing measurement, the evidence CQ\n(**2**); the whole piece is written in the company's frame — the worker-welfare lens sees commute\nas unpaid labour (**4**). Rung 5: none.\n\n**7. Michael Strain, AEI, \"Just Say No to Remote Work\" (Jul 2026)** — *Remote work explains\ntwo-thirds of the youth unemployment gap, so young workers should refuse hybrid days.* Missing:\nthe level shift — the study concerns employers' hiring for remote *roles*; an individual declining\nhybrid days does not change the role's remoteness, so the practical-reasoning CQ *does the means\nreach the goal?* fails (**2**, and any labour economist sees it: **3**); the advice about arriving\nbefore the supervisor is a separate norms claim riding on the evidence claim (**2**); housing cost\nnear the office from the young worker's frame (**4**). Rung 5: none.\n\n**8. Ben Le Fort, \"Think renting is throwing money away?\" (Apr 2026)** — *Renting and investing\nthe difference frequently outperforms owning.* Missing: leverage — a mortgage is leveraged exposure\nto appreciation and the cited simulation's treatment of it is the crux (**3**); tax asymmetries,\nuntaxed imputed rent, capital-gains exclusion (**3**); rent-inflation risk against a fixed nominal\nmortgage (**3**); the discipline assumption, which the author names himself (**0**); the security\nframe versus the finance frame, which he also reaches — \"a lifestyle choice\" (**4**). Rung 5: the\nindividual reader's discipline and life plan — evidence.\n\n**9. Stefan Holgersson, DN Debatt, \"Gör om SiS från grunden\" (Aug 2026)** — *Replace SiS with a\nnew agency and small state homes.* Missing: whether the failures are organizational or\nstaffing/competence failures a new sign cannot fix — which the author pre-empts (**0/3**);\ntransition risk for children placed today (**2**); evidence that smaller units reduce violence,\nthe Nordic institutional-care literature (**3**); the state/municipality responsibility split\n(**3**); the children's own perspective as a frame (**4**). **Rung 5, and real**: what placement\nfeels like from inside — testimony of placed children, which exists in almost no text a model\nreads. The consideration *\"does the reform improve the children's experience?\"* is rung 2; the\n**content** of that experience is rung 5. **This is where the reserved class actually lives.**\n\n**10. Peter Heather, History & Policy, \"The fall of the Roman west\"** — *External factors (Persia,\nthen the Hunnic-driven migrations) were the prime mover.* Missing: Turchin's structural-demographic\ndriver, elite overproduction — a rival account any historian knows (**3**, and a rival frame:\n**4**); the East-survived counterfactual cuts both ways since the East also faced Persia — the\nargument-from-difference CQ (**2**); climate and plague (Harper) (**3**). Rung 5: none.\n\nThree extras, same shape: **Turchin** (elite overproduction as the law) is missing Heather\nsymmetrically (**3/4**) and the falsifiability of \"laws of history\" (**4**); the **Hacker News\nnuclear thread** is missing nothing — the *crowd* surfaced time-to-build, full lifecycle for\nrenewables, regulation, and the portfolio reframe that dissolves \"best bet\" (**2–4**, all supplied\nby other commenters, which is lens rotation done by a population); the **EV purchase piece** is\nmissing total cost of ownership (**3**), the scope of \"most people\" (**1**), and its depreciation\nfigures are dated — a staleness case (**3**, time-indexed).\n\n## 3. The tally, and what it says\n\nAcross 13 arguments and 47 graded considerations: **rung 0: 2 · rung 1: 1 · rung 2: 13 · rung 3:\n22 · rung 4: 9 · rung 5: 0 considerations, 4 pieces of evidence** — *as graded by a model; see § 10 for why\nthat zero is structurally unable to be anything else.*\n\nThree things follow, in descending order of how much they change.\n\n**(a) The sentence-level check is the wrong yardstick for \"missing.\"** `check_coverage` (rung 0)\nfound 2 of 47. It answers *did the parts drop something the parent said* — which is a real\nquestion, the losslessness question — but it is not the completeness question. **Completeness is\nrungs 2–4, and those are the CQ scaffolding, the implicit-premise pass, and lens rotation.** The\nmachine reaches all three; the corpus already ships the first as 64% of the graph.\n\n**(b) The human-reserved residue for CONSIDERATIONS is close to empty on ordinary debates; the\nreserved residue is EVIDENCE.** Every \"nobody here has said X\" that a strong critic would raise was\nin the public corpus. What no model could supply was *what it costs from where I stand* (the SiS\nchildren; Nigerian households) — which is the fourth act in\n[what-human-judgment-is-for.md](what-human-judgment-is-for.md), impact testimony, the attunement\npole's own contribution. **The distinction that does the work: a consideration is a dimension that\nbears on the conclusion; evidence is the ground for it.** Machines derive dimensions; humans\nsupply much of the ground.\n\n**(c) Where rung 5 does appear, it is predictable: it lives in the tail.** Underrepresented\nregisters — placed children, households in a thinly-documented economy, the minority rendering —\nare exactly what `saturation-and-the-long-tail.md` § 3b names as the value-bearing tail and what\nthe ingestion-balance instrument was proposed to measure. Coherent absence is not diffuse; it\nconcentrates where text is thin. That makes it targetable: register diversity, cluster-balanced\ningestion, and impact testimony as first-class evidence are the defenses, and they already exist as\nproposals.\n\n**One honest caveat on rung 4.** *Never ask a mind to audit its own frame* still holds in narrowed\nform: a model prompted with a different lens produces the corpus's *rendering* of that lens — a\nsample of the corpus's frame, not a differently positioned observer. It reliably finds the\nwell-articulated alternative frames (every rung-4 item above is a famous one) and cannot find a\nframe nobody has written down. The human of that frame remains the stronger instrument; the machine\nis a cheap first pass.\n\n## 4. What this does to the deflation defense (weakest-link doc § 9d, corrected)\n\n§ 9d sorted a required part by whether the parent's sentence *entails* it. That was rung 0, and\nthe ladder shows it is the wrong cut: the honest \"nobody has established X\" is usually rung 2–3,\nand a rung-2–3 requirement is exactly as machine-derivable as it is legitimate. The corrected sort:\n\n- **Inside the derivable requirement space (rungs 0–4)**: the machine can pre-populate these\n  for *every* claim, uniformly, as latent potentials — which the CQ scaffolding already does for\n  rung 2 and the latent-scaffolding ruling makes cheap. A requirement the machine would have\n  derived anyway earns nothing when a human asserts it selectively: **selective scrutiny stops\n  paying, because the scrutiny was already there for everyone.** That is payoff removal on the\n  deflation attack, in the same shape min was on the inflation attack.\n  **⚠ Corrected 2026-09-03 (drowning doc § 8): \"inside the space\" does NOT mean the cap applies\n  automatically.** The machine's own derivable list is roughly half junk (CQs-Gen 2025), so a\n  materialized junk question at 0.5 would cap every whole at neutral forever. The cap weight is\n  the edge's *relevance strength* — presumed-but-rebuttable and modest for a scheme's standard\n  question, argued for anything else, closable by a *not applicable* answer that is itself\n  attackable. This also resolves the register's fork 2 (latent requirements in the cap?): they\n  count by relevance strength, never by mere existence.\n- *(Built 2026-09-17 in this corrected form — `support_semantics.requirement_weight`; the weakest-link doc § 9d carries the build note. The pre-population substrate is still unbuilt, so the presumption is per edge provenance rather than per rung: confirmed parts 0.5, unconfirmed proposals 0.)*\n- **Outside it (rung 5)**: a proposed requirement, which caps nothing until argued and then caps\n  through its argument's strength — unchanged from § 9d, but now correctly scoped to the genuinely\n  novel residue rather than to everything a sentence fails to entail.\n- **The saboteur's \"moon is cheese\"** fails the derivability judge (not in the space under any\n  rung) and lands as a proposal that must argue — a single-shot judgment on the record, so the\n  judge-wiggle result does not reach it. A *plausible* spurious requirement passes the judge, and\n  then it is simply a legitimate critical question that anyone could have asked.\n\n## 5. What it does to the lift, and one new design fork\n\nThe lift's coverage measure is therefore not *do the children exhaust the parent's sentence* but\n**what fraction of the parent's derivable requirement space is established** — weighted by\ncentrality (the hinge already ranks it). The raised-not-total up-weight is the residual for rung 5,\nand it now has a name: *the price of the tail*.\n\n**New fork, unruled**: should unanswered *latent* derivable requirements participate in the\nweakest-part **cap** (an unanswered critical question at 0.5 would then cap every claim at neutral\nuntil answered — honest but severe), or only in the lift's coverage measure (a claim rises as its\nderivable space fills and is capped only by requirements someone has materialized)? The edge layer\nalready takes the first position for CQs on inferences (an edge with unanswered CQs contributes\nzero). Consistency argues for the first; the pebble worry argues for the second. Founder call.\n\n## 6. Resource: *Bayesian Decision-making Algorithms* (Terenin, 2026, work in progress)\n\n`https://bayesianalgorithms.com` — Alexander Terenin's public-draft monograph (compiled\n2026-08-28; only the introduction and Chapter 2 are up). Chapters: expected improvement, **Gittins\nindices**, optimism / upper confidence bounds, entropy search and information-based algorithm\nexecution, Thompson sampling — explore-exploit decision-making under a stochastic model.\n\nWhy it belongs beside this doc and the ask-or-act thread\n([active-inference-context-acquisition-and-deliberus.md](active-inference-context-acquisition-and-deliberus.md)):\n\n- **Pandora's Box is the descent.** Weitzman's problem — boxes with known priors, each openable at\n  a cost, take the best once you stop — is exactly a claim's latent requirement space: each\n  unanswered requirement is a box, answering it costs attention, and the Gittins/reservation value\n  ranks which to open next *non-myopically*. `worth_asking` and the hinge are the greedy one-step\n  version; Gittins is the exact policy for independent boxes.\n- **The reservation value is the stopping rule the lift needs.** Stop opening when the best known\n  value exceeds the next box's reservation value — which gives the raised-not-total up-weight a\n  principled form: the expected strength given the opened requirements and the priors over the\n  unopened ones, rather than a hand-set constant.\n- **Entropy search is `worth_asking` done properly**: choose the question that most reduces\n  uncertainty about the *target* quantity (the root's strength), not the question with the largest\n  local swing.\n- **Thompson sampling** is the argument for randomizing which gaps get surfaced, against the\n  curiosity-gradient trap where every reader descends the same head.\n\nFiled as a TODO loop: evaluate reservation-value ranking once latent requirements exist. Not\ncited in any argument until the relevant chapters are public.\n\n---\n\nCross-references: [weakest-link-arithmetic-and-the-merge-hunch.md](weakest-link-arithmetic-and-the-merge-hunch.md)\n§ 7, § 9d · [what-human-judgment-is-for.md](what-human-judgment-is-for.md) § Act one · [the-harness-conviction.md](the-harness-conviction.md)\n(the two reserved classes, whose *size* this calibrates) · [saturation-and-the-long-tail.md](saturation-and-the-long-tail.md)\n§ 3b (where rung 5 lives) · [the-residual-error-taxonomy.md](the-residual-error-taxonomy.md)\n(coherent absence, class 2) · [latent-scaffolding-and-the-type-token-split.md](latent-scaffolding-and-the-type-token-split.md)\n(the pre-population substrate) · [scheme-set-exhaustiveness.md](scheme-set-exhaustiveness.md)\n(rung 2's generative basis).\n\n## 7. Light research check (2026-09-03, founder-commissioned) — what the literature adds to the ladder\n\nSix sources, read at abstract-plus-key-sections depth. They confirm the recall half of the finding\nand add four sharpenings the ladder lacked.\n\n**Sharpening 1 — the machine finds the missing considerations, but half of what it offers is\njunk, and it cannot grade its own.** The 2025 Critical Questions Generation shared task (ArgMining\n2025; Calvo Figueras & Agerri) had LLMs generate Walton-style critical questions for real debate\ninterventions. The winning system's questions were ~57–59% *useful*; a strong generation subset\ncame out **44% useful, 23% unhelpful, 15% invalid** (TriLLaMA); the 2024 zero-shot baseline was only\n28% valid. And every LLM *classifier* over-rated its own candidates — **75–85% labelled useful\nagainst 44% actually useful** — same-hand bias measured in the wild. So rung 2 is reachable at\nrecall and poor at precision, which is exactly why pre-population must be *latent* (the\nlatent-scaffolding ruling) and why what enters the cap needs a gate the generator does not hold.\n\n**Sharpening 2 — models detect absence and then do not act on it.** Fan et al. 2025\n(arXiv:2504.06514, *Missing Premise exacerbates Overthinking*) find reasoning models usually\n*suspect* early that a problem lacks a necessary premise, then \"do not dare to abstain\" and\ngenerate thousands of tokens anyway; non-reasoning models abstain better. Consequence: a\n`does_not_fit` / `missing_premise` slot must be an explicit output field, never left to the model's\ninitiative — the confession principle with an external mechanism behind it. (Detection is rung 2–3\nterritory; *acting* on detection is the training recipe's gap.)\n\n**Sharpening 3 — the tail is documented, by name.** Santurkar et al. 2023 (*Whose Opinions Do\nLanguage Models Reflect?*, ICML) find LM opinions misaligned with the US public \"on par with the\nDemocrat–Republican divide on climate change\", skewed liberal/educated/wealthy, with groups **poorly\nrepresented by every model tested** (65+, widowed, high religious attendance), and that RLHF-tuned\nmodels **collapse to the modal view of a group** — caricatures, >99% on one option where humans\nspread. Steering toward a group helps modestly and resolves none of it. Tao et al. 2024 (PNAS\nNexus) replicate at country scale: all GPT versions sit with English-speaking Protestant Europe on\nthe Inglehart–Welzel map; cultural prompting improves alignment for 71–81% of countries and\n**does not close the gap**. This is rung 4's caveat with numbers (a prompted lens is the corpus's\nrendering of that lens) and rung 5's address: the corpus-absent lives where these papers say the\nmodels are blind.\n\n**Sharpening 4 — \"positional\" has a philosophical literature, and it splits the residue in two.**\nStandpoint epistemology (Toole 2023, *J. APA*; Saint-Croix's formal model; a 2026 *Philosophical\nStudies* paper on zetetic deference) distinguishes (a) the **evidential advantage** of a social\nlocation — access to \"what-it's-like\" evidence, which the occupant alone can *gather* but which\n**is shareable through testimony, story-telling and analogy** — from (b) a **standpoint**, an\nachievement reached through consciousness-raising, likened explicitly to expert training. And it\nnames a second form of deference beside epistemic: **zetetic deference** — letting the positioned\nshape *which questions get asked* (the Flint water example: scientists took residents' concerns as\nthe starting point of inquiry). Read onto the ladder: the human residue is not one thing but two —\n**positional evidence** (act four, impact testimony) and **positional question-setting** (the\nframe class, rung 5's second half). The first the graph can hold as attributed evidence; the\nsecond is what \"this region is framed wrong\" was reaching for, and standpoint theory says it is\n*earned collectively*, which makes a co-present session of people from one location an instrument,\nnot a nicety.\n\n**Two smaller confirmations.** Implicit-premise recovery by LLMs is good and improving: a single\nLLM already beats the prior state of the art on the SemEval-2018 warrant task and a two-agent\ndebate reaches 0.87 accuracy, with the finding that **forcing agents to defend assigned stances\ndegrades results** (Ku et al., ArgMining 2025) — a design note for the two-family jury and the\nlens-rotation pass: let them choose, never assign. And counter-argument work agrees that finding\ncandidates is easy while **choosing which premise matters is hard**: GPT-4 agrees with experts on\nthe premise to attack only ~80% of the time (Ozaki et al., NAACL-SRW 2025; Alshomary et al. 2021).\nSo the hinge and `worth_asking` are doing the genuinely difficult part, and their verdicts should\nstay proposals.\n\n**Net effect on the canon proposal (§ 3b of the register).** The line should point at *two* human\nacts, not one: *here is what this costs from where I stand* and *here is what you should be\nasking* — evidence and question-setting from a position. \"What argument is missing here\" stays\na human-performable, low-rung act, but the literature agrees it is no longer the human-reserved\none.\n\nSources: Calvo Figueras & Agerri (CQs-Gen 2025 overview; 2024 CoNLL); Del Favero et al., Ramponi\net al., Turkstra et al. (ArgMining 2025 system papers); Fan et al. arXiv:2504.06514; Santurkar et\nal. ICML 2023; Tao et al. PNAS Nexus 2024; Toole, *J. APA* 2023; Saint-Croix (PhilArchive);\n*Philosophical Studies* 2026 \"Zetetic and epistemic deference in standpoint epistemology\"; Feng &\nHunter arXiv:2603.06114 and arXiv:2608.18821; Ku et al. ArgMining 2025; Alshomary et al. Findings\nACL 2021; Ozaki et al. NAACL-SRW 2025. Read-depth: abstracts and key sections; verify before citing\noutward.\n\n## 8. Can we simulate the tail to counteract the skew? (2026-09-03, founder question)\n\n**Short answer: yes for questions, no for testimony — and the literature says exactly where the\nline runs.**\n\n**Where simulation works.** Argyle et al. 2023 (*Out of One, Many*, silicon sampling) showed a model\nconditioned on *real* respondents' backstories reproduces the pattern of associations in US survey\ndata with a mean Cramér's V difference of 0.026 — and the method's whole point is *correcting skewed\nmarginals by conditioning on real data*. That is the head: groups the corpus already documents,\nanchored on real records. Tao et al. 2024 add that cultural prompting improves alignment for\n71–81% of countries.\n\n**Where it fails, in three named ways.** Wang, Morgenstern & Dickerson (*Nature Machine\nIntelligence* 2025; 3,200 participants, 16 identities, 4 models): (1) **misportrayal** — a model\nprompted as a group renders what *out-group* members say about it, because training text rarely\nlinks author identity to text (their example: a \"person with impaired vision\" persona that opens\n\"while I may not be able to visually observe…\"); (2) **flattening** — every model on every question\nwas less diverse than humans, GPT-4 covering 3 of 5 answer options across 100 samples, which erases\nexactly the within-group variation the tail is made of; (3) **essentializing** — the identity\nprompt itself reduces a person to fixed traits. Cheng, Durmus & Jurafsky 2023 (*Marked Personas*)\nmeasured that GPT-3.5/4 personas of marked groups contain *more* stereotypes than human-written\nones with the same prompt, including seemingly positive ones (the \"strong, resilient\" trope).\nCummins 2025 (arXiv:2509.13397) shows the fidelity of silicon samples swings from r = .23 to .84\nacross defensible configuration choices even for well-represented respondents, and states the tail\ncase outright: silicon samples risk *\"an illusion of representation while in fact speaking over\nthose very same populations.\"* Wang et al.'s inference-time mitigations *reduce but do not remove*\nthe harms; their recommendation is supplement, never replace.\n\n**The Deliberus reading.** The skew is in the training data, so a persona prompt draws on the same\nskew: simulation is strongest where the corpus is rich (where you needed it least) and weakest\nwhere it is thin (where you needed it most). And the specific failure — fluent, confident, flat —\nis the one most likely to *pre-empt* the real voice: a graph that already \"has\" the child's\nperspective stops looking for the child. That is the interpassivity warning and the T-cell warning\nfrom the corpus's own docs, arriving with numbers.\n\n**Design that makes simulation net-positive (proposal, unruled).** Eight constraints; the first\nthree are the load-bearing ones.\n\n1. **Simulate to ask, never to testify.** A simulated lens produces *questions and candidate\n   considerations* — \"what would a placed child ask about this reform?\" — entered as latent\n   requirements. It never produces claims attributed to a group. Rung 4 output only; rung 5\n   stays empty until a person fills it.\n2. **Its own provenance kind.** `simulated` beside asserted / minted / reported, on every count\n   and every surface. A reader must never mistake a simulated consideration for a heard one.\n3. **Cap-only, never lift.** A simulated requirement may keep a whole from reading *established*\n   (an open marker: \"the affected position has not been heard on this\"), but carries zero evidence\n   weight and can never raise a strength. Simulation lowers false confidence; it may not create\n   confidence.\n4. **Anchor on real testimony wherever any exists** — Argyle's move. The online-testimony backfill\n   (register band D) supplies anchors; an *unanchored* simulation is labelled so and treated as a\n   louder alarm, not a weaker voice.\n5. **Distributions, not modes.** Ask for the range of positions *within* the group and the\n   disagreements among them; sample repeatedly; use two model families. This targets the\n   flattening failure directly, since it is measured as the largest.\n6. **Simulation is a gap detector, not a gap filler.** Its most valuable output is the finding\n   *\"no real testimony from this position exists in the graph\"* — an ingestion-balance alarm and a\n   recruitment target. The instrument points at the hole; it does not fill it.\n7. **Override and refusal.** A real member of the position can disavow the simulated rendering\n   (act three, extended to a rendering of one's group — held with the identity-claims question);\n   and the pass may *decline* to simulate identities where the stereotyping literature is explicit\n   (Cheng et al.'s \"critical refusal\").\n8. **Calibrate per position.** Track how often simulated considerations are later confirmed by\n   real testimony from that position. Where the rate is high, simulation is a usable first pass;\n   where it is low, stop simulating and recruit. This is the human-vs-daemon survival instrument\n   (analogy daemon § 7) applied to lenses.\n\nNet: simulation is the lens-rotation pass with a muzzle — it may ask on the tail's behalf and may\nnever answer for it.\n\nSources: Argyle, Busby, Fulda, Gubler, Rytting & Wingate, *Political Analysis* 2023; Wang,\nMorgenstern & Dickerson, *Nature Machine Intelligence* 2025 (arXiv:2402.01908); Cheng, Durmus &\nJurafsky, ACL 2023; Cummins, arXiv:2509.13397 (2025); Tao et al., *PNAS Nexus* 2024; Santurkar et\nal., ICML 2023. Read-depth: abstracts and key sections.\n\n## 9. Which human contributions belong? (2026-09-03, founder: *\"I still feel a tension re the saboteur etc. How can we know which human contributions belong or don't?\"*)\n\n**We do not know, and the design does not need to.** Under the blessed rulings there is no\nmembership test: *contributions are added, never overwritten — nothing needs permission to\nexist.* \"Belong\" is therefore the wrong frame; the right one is **what a contribution may DO**.\nExistence is free; effect is earned. And effect is earned differently for each of the three\nthings a human can add:\n\n| Human contribution | How it earns effect | Where the saboteur is caught |\n|---|---|---|\n| **A claim or argument** | strength from the evidence and attacks around it | by being attacked; a bad argument survives nothing |\n| **A requirement** (\"W needs X\") | inside the derivable space: pre-populated for everyone, so asserting it adds nothing; outside: caps only through its own argued strength | the argument for necessity is attackable; \"moon is cheese\" never gets an argument |\n| **Impact testimony** (\"here is what this costs from where I stand\") | as evidence through the *witness testimony / position-to-know* schemes, whose critical questions are exactly: *is the witness in a position to know? is the testimony consistent with other witnesses? does the witness have a stake in deceiving?* | by people who share the position |\n\n**The sharp part, stated honestly.** Section 3 found that lived experience is the one thing the\nmachine cannot supply. The same fact read the other way: **it is the one thing the machine cannot\ncheck.** A consideration in the derivable space can be verified against the corpus; testimony\nfrom a position is corpus-absent by definition, so fabricated testimony is the saboteur's best\nremaining move — *\"I was in a SiS home and it was fine.\"* That is where the founder's tension\nlives, and it is real.\n\n**What defends, in order of strength.**\n\n1. **Position is itself a claim.** *\"This account is a person who lived in a SiS home\"* is a factual\n   claim like any other (the founder's held direction: endorsement and identity claims are\n   ordinary claims, no verification ticks). It carries evidence or it does not; testimony from an\n   unevidenced position is labelled so — the gray band for witnesses.\n2. **Testimony is evidence for a premise, never a verdict.** Its reach is one edge with a scheme\n   and critical questions. It cannot cap a whole (that needs a requirement) and cannot lift one\n   (that needs coverage). Fabricated testimony moves one premise, by one edge's worth.\n3. **Corroboration with the novelty dial.** One uncorroborated voice moves a premise little; many\n   independent voices from the position move it more. Standpoint theory says a standpoint is\n   *achieved collectively* (consciousness-raising), which matches: a lone claimed position is\n   weak evidence, an articulated community is strong. To fake that, the saboteur must fake many\n   independent voices — which is the sybil attack, and the corpus has said plainly that no\n   layer defends against it. The defense there is infrastructural (identity costs, pseudonymity\n   tiers) and social, not epistemic.\n4. **Positional contributions are verified positionally.** This is the closing principle, and it\n   follows from everything above: the machine verifies the derivable space; **only people who share\n   a position can verify testimony from it.** *\"That is not what it is like\"* from another\n   care-leaver is the check no model and no outsider can run. So the co-present session, the\n   islands, the recruitment of many somewheres are not only how positional evidence gets *in*;\n   they are how it gets *checked*. A graph with one voice from a position holds unverifiable\n   testimony and should say so; a graph with a community from that position has its own\n   verification layer.\n\n**What this does to the saboteur worry.** The saboteur cannot make a claim belong; he can only\nmake one exist, which everyone can. To make it *do* anything he must argue (attackable), be\ncorroborated (needs many voices), or be believed as a witness (needs a position claim, checked by\nthose who share it). Each route has a cost that rises with the effect sought, and each leaves a\nrecord. What remains undefended is exactly what the corpus already lists as undefended: a\nwell-resourced strategy-class actor fabricating many independent-looking positions. The honest\nline is not that the system knows which contributions belong. It is that the system never has to\ndecide, and makes every route to influence pass through people who can tell.\n\n## 10. The grader's blind spot (2026-09-05, founder critical-examination pass)\n\nThe founder's challenge — *\"None at all on entry?\"* — was aimed at the conservative design, but the\nsame doubt lands harder here: **none at all in rung 5?** The zero was produced by a model grading\nwhether a consideration is *corpus-absent*. A model cannot recognise a consideration its corpus\nlacks: either it does not generate it, and so never lists it as missing, or it generates it, and so\nit was never corpus-absent. **The test is biased toward zero by construction** — the\n*never-ask-a-mind-to-audit-its-own-frame* principle violated at the centre of a finding about that\nvery principle. The four items graded as unreachable evidence were the same grader's judgment of\nwhat it lacked, so they carry the same bias.\n\n**What survives.** The considerations listed here matched what human commenters in the same threads\nactually raised — the tuition thread's teacher reductio, the Hacker News thread's portfolio\nreframe — which is an external check on *recall at rungs 2–4* and says nothing about rung 5. So the\ndefensible statement is: **on thirteen ordinary public debates, a model raised every consideration\nthat human commenters also raised, and the only things it could name as beyond its reach were\nevidence from a position.** Whether a positioned human would raise considerations the model cannot\ngenerate is untested by this design, and untestable by *any* model-graded design.\n\n**The instrument that would test it** (proposed, unruled): hand the same thirteen arguments to\npeople from the affected positions — a care-leaver for the SiS reform, a Lagos household for the\nsubsidy, a renter in a rent-controlled city — and count what they raise that the model did not.\nThat count is the first honest measurement of the reserved classes' size; everything above is a\nlower bound on the machine's reach, not an upper bound on the human's.\n\n**Consequences propagated the same day**: the phrase *\"0 corpus-absent\"* is softened on every\nsurface that carried it (CLAUDE.md topic row and frontier sentence, the canon principle's evidence\nclause, `what-human-judgment-is-for.md`'s calibration note, the TODO register, the session tables);\nthe canon re-pointing of 2026-09-03 stands as provisional on the softer claim, which still supports\nit, and the founder may revisit it. Full examination:\n[session29-the-derivability-arc-and-the-conservative-design.md](session29-the-derivability-arc-and-the-conservative-design.md)\nPart III.\n"}