{"path":"research/scientific-method-two-steps-and-deliberus.md","content":"# The Scientific Method in Two Steps — What It Means for Deliberus\n\n**Date**: April 25, 2026\n**Type**: External-source comparative analysis + epistemological synthesis\n**Source article**: \"The Scientific Method\" by Jamie Simon (UC Berkeley / Imbue), learningmechanics.pub, April 19, 2026\n**Status**: Analysis complete — proposes testable metrics for Deliberus's bold bets and reframes scheme-bounded decomposition as field-specific Step B technique\n\n---\n\n## Executive Summary\n\nJamie Simon argues the scientific method consists of exactly two steps: (A) figure something out, and (B) check and make sure you're not wrong. Both are non-negotiable; either order works; field-specific auxiliary techniques (preregistration, error bars, blinding) are useful ways to do Step B but none is essential. Missing Step A gives you engineering; missing Step B gives you philosophy; science requires both. \"Science is an edifice that builds on itself — each piece of brickwork must be quite solid to support future building.\"\n\nThis maps onto Deliberus at every scale: the extraction pipeline (scout = A, self-eval = B), user contributions (claim = A, CQ answer = B), the sorry model (A with B explicitly pending), and the civilizational level (convergence thesis = A, live graph producing empirical results = B). The edifice metaphor is structurally identical to the `@[simp]` flywheel. Simon's \"interesting nonrigorous > uninteresting theorem\" validates sorry-driven development. His warning against formalism without insight matches the graveyard of 20 argumentation platforms.\n\nThe article's principal contribution to Deliberus is a framework for honest self-assessment: which parts of the project are Step A (the 60+ research docs, the convergence thesis, the alignment infrastructure claim) and which are Step B (the live pipeline, QEM, community correction UX). Until both steps meet at sufficient scale, the philosophical superstructure remains hypothesis, not science. The secondary contribution is a reframing of Walton's schemes as field-specific Step B auxiliary techniques — checking protocols calibrated to argument-type-specific failure modes.\n\n---\n\n## The Two-Step Model Maps Onto Deliberus at Every Scale\n\n### Extraction Pipeline Level\n\nStep A = the scout, focused extraction, decontextualization, classification, relationship detection, and scheme classification passes — figuring out the argument structure latent in source text. Step B = the self-evaluation pass, CQ generation, auto-connect, QBAF/QEM computation, and the human confirmation step — checking the extraction is not wrong.\n\n### User Contribution Level\n\nStep A = contributing a claim, decomposing a value premise, answering a CQ, clarifying a contested concept sense. Step B = community challenge via CQ-answering, voting, evidence submission, concept lifecycle stabilization, and QEM badge propagation. The sorry marker is the explicit notation that says \"Step A is done here; Step B is still pending.\"\n\n### Epistemic Architecture Level\n\nStep A = \"No Copout Axioms\" / the Socratic function — figuring out what is underneath each premise layer. Step B = QEM gradual semantics, scheme-bounded CQ verification, bridging signal — checking that the decomposed structure is solid enough to build on. The self-similar decomposition principle says this A-to-B cycle is recursive: each level of decomposition reveals the next level's sorry markers (Step A again), which then need checking (Step B again).\n\n### Civilizational Level\n\nStep A = the convergence thesis, the bridging concept, the worldview-filter design, the alignment infrastructure claim — bold conjectures about what the graph could reveal. Step B = the live system at deliberus.com, where actual extractions, actual CQ answers, actual bridging scores, and actual cross-extraction auto-connect edges either confirm or refute these conjectures empirically. \"Building IS the cheapest validation\" is Simon's point that experiments are cheap and should be run.\n\n---\n\n## \"Science as Edifice\" IS the @[simp] Flywheel\n\nSimon: \"Science is an edifice that builds on itself. It usually consists of so many layers that each piece of brickwork must be quite solid to support future building, and each brick must be crafted with future bricks in mind.\"\n\nThis is structurally identical to the `@[simp]` flywheel. Each well-vetted claim is a brick. Its solidity is measured by QEM strength (green badge = both steps done; amber = partially checked). It supports future claims via `SUPPORTS`/`DECOMPOSES_INTO` edges. The flywheel is anti-rivalrous accumulation: 93 cross-extraction auto-connect edges across 7 extractions (Session 8) means each new source links immediately into the accumulated structure.\n\nThe color-coded argument health from P5 IS Simon's edifice principle made visual:\n\n| Color | Simon equivalent | Meaning |\n|---|---|---|\n| Green | Both steps done; brick is solid | Well-supported, evidence + community vetting |\n| Blue | Step A done, Step B ready | Claim stated, prerequisites met, awaiting evidence |\n| White/green border | Step A done, Step B blocked | Some supporting premises are still sorry |\n| Orange | Step A done, Step B far away | Deep dependencies unmet |\n\n---\n\n## \"Interesting Nonrigorous > Uninteresting Theorem\" Validates Sorry-Driven Development\n\nEvery argumentation platform in the graveyard demanded complete formalization before a contribution could enter the system. That is the \"uninteresting theorem\" failure mode: rigorous structure nobody contributes to because the cost of entry is too high.\n\nDeliberus's sorry marker IS the \"interesting nonrigorous result.\" It sketches the structure (Step A) with explicit gaps where Step B is still needed. A claim with sorry markers is interesting (reveals structure, creates connection points, invites contribution), nonrigorous (premises unverified, CQs unanswered), and more valuable than a perfect claim that never gets contributed.\n\nThe sorry model privileges useful structure over complete rigor, and lets rigor accumulate gradually via community contribution — exactly the dots-on-curves discipline Simon describes.\n\n---\n\n## \"No Rules for Step A\" IS \"Structure Is Output, Never Input\"\n\nSimon: \"There are no rules whatsoever as to how you do the first step.\"\n\nThis maps precisely onto Deliberus's founding UX convictions: P3 (voice as first-class input), P7 (\"Help me think about X\"), founding conviction number 4 (\"Structure is OUTPUT, never INPUT\"). The user's Step A has no format requirements. They speak, type, paste a URL, upload a PDF.\n\nSimon: \"There is only one rule with the second step: you have to do a good job checking.\"\n\nThis is the QBAF/QEM/CQ/community-verification architecture. Once a claim enters the graph, it IS subject to typed checking: scheme-bounded critical questions, voting, evidence binding, gradual-semantics computation. Step B has rules. Step A does not.\n\n---\n\n## Auxiliary Techniques Are Field-Specific — This Reframes Walton's Schemes\n\nSimon distinguishes between the essential two steps and field-specific auxiliary techniques. Preregistration, error bars, blinding, peer review — these are \"ways to do Step B\" necessary in some fields but not others.\n\nDeliberus's equivalent: **Walton's 96 argument schemes plus CQs are field-specific auxiliary techniques for Step B in argumentation.** An argument from expert opinion gets CQs about credibility, domain match, and peer agreement — not because those are universal requirements, but because they are the right checking procedures for that specific argument type. An argument from analogy gets CQs about structural correspondence and disanalogy.\n\nThis is a cleaner justification for scheme-bounded decomposition than \"Walton's taxonomy is comprehensive.\" The Simon frame says: these are domain-calibrated verification techniques, chosen because they match the failure modes of each argument type, just as preregistration matches the failure modes of hypothesis-rich/evidence-poor fields.\n\n---\n\n## The \"Hail Mary\" Principle and Deliberus's Bold Bets\n\nSimon celebrates bold, testable hypotheses that attempt large conceptual leaps — even when disproven. His example: Information Bottleneck Theory was \"conclusively disproven soon after its proposal\" but deserves applause for being \"bold, testable hypothesis and honest attempt to figure something out.\"\n\nDeliberus's boldest bets are Hail Maries in this sense:\n\n| Bet | Why it is bold | How it could be disproven |\n|---|---|---|\n| **Convergence thesis** | Values converge at depth across all worldviews | Decomposition to maximum depth reveals irreducible value divergence on shared-bedrock candidates |\n| **Bridging signal** | Reasoning quality across disagreement is computable and novel | Bridging scores do not predict opinion change or productive engagement |\n| **No Copout Axioms** | There is no permanent bottom to decomposition | Users consistently reach genuine bedrock at shallow depth |\n| **Alignment infrastructure** | Deliberus IS alignment, not just a tool for alignment research | Graph's value structure is too noisy or culturally biased to serve as AI verification target |\n| **Sorry-model adoption** | Explicit incompleteness invites contribution rather than repelling users | Users avoid sorry-marked claims; contribution rates no better than unstructured forums |\n\nEach is valuable precisely because disconfirmation is informative. Simon: \"In our field, there are few ideas and much energy available to test them, so I would like to see more bold guesses of this type.\"\n\n---\n\n## \"Every Contribution Should Understand Itself as Part of a Project That Does Both Steps\"\n\nA sorry marker is a Step A contribution that explicitly marks Step B as still pending. The sorry marker IS the structural self-awareness that \"this contribution is part of a project that does both steps, and I've done one of them.\" The invitation cards (\"Answer this CQ,\" \"Add evidence,\" \"Decompose this premise\") are the system saying \"here is where Step B is needed.\"\n\nIndividual user contributions need not complete both steps. But the graph ensures that every contribution is part of a structure that demands both.\n\n---\n\n## \"Checking Usually Contains the Act of Application\"\n\nSimon: \"the most convincing way to check a claim is often to operationalize it.\"\n\nDeliberus applies this: the most convincing check of the extraction pipeline is running it on real content (7 extractions, 318 claims, 93 auto-connect edges). The most convincing check of the convergence thesis is decomposing actual claims across actual worldviews in the live graph. The most convincing check of the sorry model is whether real users fill in sorry markers (correction UX shipped Mar 31).\n\n\"Building IS the cheapest validation\" is Simon's \"checking usually contains the act of application\" applied to the project itself.\n\n---\n\n## The Feynman/Popper to Tao/Lean Alignment\n\n| Simon's source | Deliberus's parallel | Shared insight |\n|---|---|---|\n| Feynman: check against experiment, not theory | Tao's verification thesis: verification is the bottleneck | Step B is the rate-limiter |\n| Popper: conjectures must be falsifiable | No Copout Axioms: every premise must be decomposable further | Nothing gets a permanent pass from checking |\n| Kuhn: paradigm shifts via anomaly | Concept lifecycle: underdefined, emerging, bifurcated, locally stable | Knowledge progresses through structural transitions |\n| Yudkowsky: rationality techniques | Lean's sorry mechanism: explicit incompleteness as contribution interface | Make the gaps visible so others can fill them |\n\n---\n\n## Where the Frameworks Genuinely Diverge\n\n### Single Truth vs Contested Concepts\n\nSimon assumes a single truth that experiments can reveal. \"Check you're not wrong\" presupposes a fact of the matter. In Deliberus's domain, \"freedom\" means three different things to three different communities, and none is \"wrong.\" Step B in argumentation becomes: check that premises support the conclusion (formal), check that premises are themselves well-supported (recursive), check that you are not using a word differently than your interlocutor (semantic), and check that reasoning holds across worldview boundaries (bridging).\n\n### Reflexive Subject Matter\n\nWhen a physicist figures out how neural networks converge, the networks do not change. When Deliberus reveals that two communities use \"freedom\" differently, the act of revealing that changes the communities' relationship to the disagreement. The subject matter is self-modifying under examination — why Deliberus needs the worldview filter, the attunement pole, and the anti-static design.\n\n### \"Convince Other People\" vs \"Convince People Who Disagree With You\"\n\nSimon: \"do a good job checking, ideally good enough to convince other people.\" In natural science, those people share empirical standards. In Deliberus's domain, they may hold different worldviews. The bridging signal IS Simon's \"convince other people\" applied to the hardest case: convince people who disagree on the conclusion that the reasoning is structurally sound.\n\n### Formalism in Nascent Fields\n\nSimon cautions: \"rigor and formality are most useful when the class of objects one wishes to describe is very large and of unknown character.\" He says deep learning is NOT such a field.\n\nDeliberus's domain — human arguments, values, contested concepts — IS such a field. The class of objects is enormous and of unknown character. Pathological cases are prevalent (semantic agreement masking real disagreement). This validates Deliberus's scheme-bounded decomposition and typed edge structure. But Simon's meta-caution still applies: Deliberus must avoid building formal ontology that nobody uses.\n\n### Philosophy vs Engineering vs Science\n\nApplied honestly to Deliberus:\n- The 60+ research docs, convergence thesis, civilizational vision, bridging concept — primarily **Step A**\n- The extraction pipeline, QEM computation, correction UX, auto-connect — primarily **Step B infrastructure** (engineering)\n- The live system producing actual claims, actual CQ answers, actual auto-connect edges — where **both steps meet**\n\nSimon would say: Deliberus's philosophical superstructure is valuable but \"can rarely be built upon\" until the checking infrastructure empirically validates it. The research docs are hypotheses; the graph is the experiment. The Apr 3 atomization frontier acknowledged this: \"Deliberus is live but the results remain blurrier than the self-similar decomposition principle implies they could become.\" That gap between Step A (the principle) and Step B (the empirical pipeline quality) IS the current frontier.\n\n---\n\n## What Deliberus Should Learn\n\n### Define Testable Metrics for the Bold Bets\n\nSimon: \"count the number of interesting, easily verifiable quantitative claims.\" Deliberus should define:\n\n- Number of contested concepts with stabilized lifecycle states\n- Number of sorry markers resolved (Step B completions)\n- Number of bridging claims detected\n- Number of cross-extraction auto-connect edges\n- Ratio of user-confirmed vs user-rejected extraction claims (pipeline accuracy)\n- Whether concept disambiguation measurably reduces downstream disagreement\n\nThese are the \"dots\" on Deliberus's \"curves.\"\n\n### State the Convergence Thesis as Falsifiable\n\nThe convergence thesis is currently a philosophical conviction backed by MGE and Parfit. Simon's Hail Mary principle says: state it as explicitly testable. In topic domain X, decompose claims from worldview groups A and B to depth N. Measure whether the deepest shared premises converge or diverge. Specify what counts as evidence for and against.\n\n### \"Study Only What You Can Describe Well\"\n\nBe cautious about downstream claims (AI alignment, civilizational transformation) until upstream claims are empirically solid: Does the pipeline produce claims humans confirm? Does CQ generation reveal premises humans agree were hidden? Does disambiguation resolve disagreements humans agree were semantic?\n\n### Embrace the Engineering\n\nSimon notes Step B alone is \"usually engineering\" and \"not necessarily problematic.\" The extraction pipeline, QEM computation, and correction UX are engineering — and that is valuable. Not every part of Deliberus needs to carry the weight of the civilizational vision. The engineering is what makes the philosophy testable.\n\n---\n\n## The Deepest Connection\n\nBoth Simon and Deliberus describe anti-rivalrous accumulation where each solid contribution increases the value of all subsequent contributions. Simon calls this \"science.\" Deliberus calls this \"the civilizational deliberation graph.\" The two-step model applies to both, and its application reveals with unusual clarity which parts of the project are Step A, which are Step B infrastructure, and where the two must still meet empirically.\n\nSimon writes from natural science, where subject matter is inert and verification standards are shared. Deliberus operates on contested human reasoning, where subject matter is reflexive and verification must work across worldview boundaries. But the shared frame — figure something out, check you are not wrong, build an edifice solid enough for others to build on — is the frame both projects inhabit. Deliberus's contribution is extending that frame to a domain where \"check you're not wrong\" means not just empirical falsification but semantic disambiguation, scheme-bounded CQ verification, cross-worldview bridging, and recursive decomposition to depth.\n\n---\n\n## Cross-References\n\n- [depth.md](../depth.md) — No Copout Axioms as the Socratic version of \"nothing gets a pass from Step B\"\n- [convergence.md](../convergence.md) — the convergence thesis as the boldest Hail Mary, needing explicit falsification conditions\n- [ux-principles.md](../ux-principles.md) P4 (sorry model), P5 (blueprint), P7 (single-player entry) — sorry-driven development as \"interesting nonrigorous over uninteresting theorem\"\n- [analysis-and-attunement.md](../analysis-and-attunement.md) — the attunement axis that natural science does not need but contested reasoning does\n- [conceptual-threads.md](../conceptual-threads.md) Thread 1 (Structure-Adoption Paradox) — the graveyard as the \"uninteresting theorem\" failure mode\n- [lean-deliberus-analogies.md](lean-deliberus-analogies.md) — the `@[simp]` flywheel as Simon's edifice principle; sorry markers as Step A with Step B pending\n- [scheme-bounded-decomposition-and-evidence-as-subgraph.md](scheme-bounded-decomposition-and-evidence-as-subgraph.md) — Walton's schemes reframed as field-specific Step B auxiliary techniques\n- [qbaf-gradual-semantics-research.md](qbaf-gradual-semantics-research.md) — QEM as the gradual-semantics implementation of Simon's \"solid enough to build on\"\n- [session10-empirical-collaboration-and-atomization-frontier.md](session10-empirical-collaboration-and-atomization-frontier.md) — the frontier where Step A and Step B currently meet\n- [self-similar-decomposition-and-claim-ontology.md](self-similar-decomposition-and-claim-ontology.md) — recursive A-to-B cycle at every decomposition level\n- [wiki-that-writes-itself-and-productive-friction.md](wiki-that-writes-itself-and-productive-friction.md) — companion analysis of the Extended Brain article, sharing the hormesis/productive-friction frame\n"}