{"path":"convergence.md","content":"# Convergence: At the Deepest Level, How Much We Share\n\n**When you decompose any human value system far enough, worldviews converge on shared bedrock.**\n\nThis is the project's boldest claim, and it rests on a prior and more modest one: that the descent is possible at all, because what blocks it is a cost rather than a mystery ([Opacity Is a Cost, Not a Mystery](vision.md#opacity-is-a-cost-not-a-mystery-aug-2026)). Convergence can lose. Whether the descent is traversable at all is what makes the question askable either way.\n\nThe Muslim, the Buddhist, the libertarian, the socialist disagree about policy, definitions, empirical facts, cultural framing. But drill past those middle layers and what remains is shared: consciousness matters, suffering is real, shared humanity is foundational. The disagreements that consume public discourse live in the middle layers — not at the bottom. Cross-cultural research supports this: machine-reading analysis of 256 societies finds the same seven moral structures present across all of them (Curry et al., 2019; Alfano et al., 2024). Derek Parfit argued that consequentialism, Kantianism, and contractualism are \"climbing the same mountain on different sides.\" The [MGE paper](#empirical-support-the-mge-paper) shows that when values are structurally elicited, participants overwhelmingly converge on the directionality of value relationships.\n\nThis is also why the convergence thesis matters for **AI alignment**. Current alignment approaches — RLHF, Constitutional AI, inverse reinforcement learning — treat values as inputs to extract or approximate. None provides the persistent, structured infrastructure where the *reasoning behind* values is decomposed, debated, and refined. If values converge at depth, that convergence is the natural alignment target — not \"what most people voted for\" but what emerges when reasoning is made transparent. An AI system verifying its ethical reasoning against a Deliberus graph would work like [Lean's type-checker](#the-lean-analogy-at-its-deepest) for mathematics: not proving conclusions right, but ensuring the reasoning structure is sound, the critical questions are answered, and the gaps are visible. This is alignment through shared foundations, not imposed constraints. See [research/deliberus-as-alignment-infrastructure.md](research/deliberus-as-alignment-infrastructure.md) for the full case.\n\nThe same thesis has an economic implication: if most disagreement lives in the middle layers rather than at bedrock, then the space of possible non-zero-sum cooperation is larger than surface politics makes it appear. Decomposition does not force agreement; it identifies which conflicts are genuinely irreducible and which are artifacts of compressed language, missing information, or false polarization. The cooperation synthesis treats convergence as an empirical prediction: lower the cost of mutual understanding, and many apparent zero-sum games should reveal positive-sum structure. See [research/non-zero-sum-economics-and-civilizational-cooperation.md §9](research/non-zero-sum-economics-and-civilizational-cooperation.md).\n\n**Fredrik's voice in plain Swedish** (Apr 30, 2026, PauseAI Sweden self-introduction, verbatim): *\"Det är iallafall hoppet och min intuition att det är möjligt, till och med när det gäller värderingsfrågor, eftersom jag tror att det mänskliga hjärtat och samvetet är i hög grad samstämmigt om man nystar tillräckligt djupt i var våra ställningstaganden kommer från; att det mesta som skiljer våra värderingar åt är våra världsbilder snarare än våra hjärtan. Visst, det finns psykopater/sociopater osv men det är alltid bara några få procent av befolkningen, som bara överlever / går under radarn i den utsträckningen som jag förstår det. Och beror ofta på trauma snarare än DNA tror jag.\"* — The hope and intuition is that convergence is possible *even on questions of value*, because the human heart and conscience are highly concordant when you trace where positions come from; what differs across people is their *worldviews*, not their hearts. The trauma-not-DNA caveat for the few-percent who appear genuinely outside this is Fredrik's own — load-bearing for the thesis (it pre-empts the \"but psychopaths exist\" rebuttal without softening the convergence claim). Full context: `.private/pauseai_sweden_introduction_2026_04_30.md` (gitignored).\n\n---\n\n## Why We Haven't Seen This Yet\n\nNo tool has ever decomposed human reasoning at this depth across worldviews. Every argumentation platform stopped at the surface — pro/con trees, binary votes, 500-character claims. The [\"acid\"](depth.md) (Dennett's Universal Acid — reason dissolving assumptions regardless of topic) has never been applied systematically to the middle layers where the confusion lives.\n\nDeliberus, with LLM-powered extraction + scheme-bounded decomposition + the Socratic [\"No Copout Axioms\"](depth.md) function, is the first system designed to reach the depth where convergence becomes visible. The technology didn't exist before: Walton's 96 argument schemes provide bounded decomposition templates, critical questions ARE premises statable in language, and LLMs absorb the structuring burden that killed every predecessor. The [March 2026 convergence moment](vision.md) — five papers on AI-augmented deliberation in a single month — confirms the field is arriving at this space. But nobody else builds persistent, evolving argument graphs with scheme classification and critical question generation.\n\n**Self-similar decomposition enables convergence testing** (April 2026): Testing the convergence thesis requires drilling through layers. The self-similar decomposition principle ensures that even the system's own intermediate representations (CQ polarity assertions, bundled claims) are decomposable — each level peeling away one more layer of the \"middle\" where confusion lives, moving closer to the shared bedrock the thesis predicts. The sorry model is recursive at every level. (See [research/self-similar-decomposition-and-claim-ontology.md](research/self-similar-decomposition-and-claim-ontology.md))\n\n---\n\n## Empirical Support: The MGE Paper\n\nThe Moral Graph Elicitation paper (Klingefjord, Lowe & Edelman, Meaning Alignment Institute, 2024; [arXiv:2404.10636](https://arxiv.org/abs/2404.10636)) provides the strongest existing evidence. Their method: an LLM interviews participants about their values in specific contexts, then maps directional edges — \"in this context, value A is wiser than value B.\"\n\nThe key finding:\n\n> \"Moral graph edges represent broad agreement amongst participants that one value is wiser than another for a particular context, and **participants overwhelmingly converge on the directionality of these transitions.**\"\n\nThis empirical result confirms the core hypothesis: convergence is not imposed — it emerges from structured decomposition. No one told the participants what the right directions were. The convergence emerged when the reasoning was made structurally legible. That is the whole thesis in miniature. The MGE method is structurally identical to what Deliberus envisions at scale, proving that \"bridge\" connections are a discoverable feature of the value-space.\n\n---\n\n## What Would Disprove This\n\nThe convergence thesis is falsifiable, and stating how is part of stating it honestly. If Deliberus reaches the depth at which decomposition becomes stable, and what we find there is not shared bedrock but irreducible structural parallax, that is a genuine discovery too — and the system will have *made it visible*. \"We share more than divides us\" and \"we share less than we hoped\" are both real answers, and the graph has to be able to report either one without collapsing into the other. Deliberus is not an argument for the convergence thesis. It is infrastructure for settling the question one layer at a time, in public.\n\nThe Apr 2026 version of the claim is that structured decomposition can *tell the difference* between genuine convergence and semantic confusion — which is the claim worth making now. The bedrock itself is what the system is for.\n\n**The accounting is what makes it a wager (Jul 2026, from the philosophical-foundations stress test).** A falsifiability attack worth stating at full strength: *a wager whose accounting treats every ceiling as unfinished decomposition cannot lose, so it is not a wager.* Fogelin's 1985 thesis (\"deep disagreements cannot be resolved through the use of argument, for they undercut the conditions essential to arguing\") plus Chang-style parity and permissivism termini are principled convergence *ceilings*, not pending work. If the residue map scores every ceiling as \"decompose further,\" the convergence claim becomes unfalsifiable by construction, and an independent synthesis pass names the same risk in Lakatosian terms (the wager has survived by reformulation, and that series needs a ledger). The honest response is a correction to the accounting, not a defense: **hinge-clash, parity, and permissivism residues count AGAINST convergence, not toward \"work remaining.\"** A residue map that can score the wager as *losing* is a more credible instrument than one that cannot, and this is what lets legibility stay the guaranteed floor while convergence stays a real, failable bet. Full treatment: [research/red-team-synthesis-2026-07.md](research/red-team-synthesis-2026-07.md), [research/philosophical-foundations-stress-test.md](research/philosophical-foundations-stress-test.md).\n\n**A terminus is dated, and the map does not say so (Aug 2026).** A classified terminus is already held as a *fallible fixed point under the current move-set*; the temporal reading adds **under the current evidence base**. An empirical residue that dissolved in 2015 can re-open in 2030, and nothing in the residue map represents a terminus expiring — so the published fraction is a snapshot presented as a state. The general form of the gap: strength computation is blind to premise age, treating a 2010 empirical premise exactly like a 2026 one (Alexander's own caveat about ten-year-old prediction-market data is the small case; Rowson's converging timescales are the general one). Evidentiary claims want a staleness property distinct from contestation — *nobody is arguing with this; it has aged out of its evidence base.* Analysis: [research/fractal-scales-and-temporal-frame.md](research/fractal-scales-and-temporal-frame.md).\n\n**The same question, asked of forecasting (Aug 2026).** Scott Alexander's \"Does Forecasting Have Room At The Top?\" measures prediction's irreducible remainder — aleatoric versus epistemic uncertainty — by triangulating the ceiling with independent anchors rather than asserting it. That is the convergence wager rotated ninety degrees: the residue map does for disagreement what his anchors do for prediction, and typed residues answer a question Brier scores cannot (*what kind* of wall was hit, not merely that one was). Whether to adopt \"the residue map measures the aleatoric fraction of disagreement\" as public framing is an open founder fork — the frame is instantly legible to this audience, but risks miscasting determinate normative residues as chance. Full read: [research/forecasting-room-at-the-top-and-deliberus.md](research/forecasting-room-at-the-top-and-deliberus.md).\n\n**The residue taxonomy is OUGHT-biased (Jul 10, from the deep-descent experiment).** The four types (fittingness, structural, axiom-choice, permissive-zone) were induced entirely from value debates (death penalty, minimum wage, assisted dying, white lies), so they are all ought-shaped. Hand-descending a deep *is*-claim breaks this: \"physical pain is bad\" wears ought-grammar but is largely a claim about consciousness, and it bottoms out on the hard problem (is felt valence ontologically fundamental or reducible), a terminus that is none of the four (not empirical-dissolving, not fittingness, not axiom-choice, not permissive-zone). Two honest responses, a founder call: add an **is-bedrock / open-empirical-metaphysical** type (a Deutsch-soluble question we lack the explanation for, held with its competing readings), or state that the residue map covers normative disagreement only and factual-metaphysical descents are out of scope by design. Left unaddressed, the classifier mis-types such leaves as `fittingness` or `undecided`. Same message as the philosophical stress test's demand to split the permissive zone four ways: a taxonomy induced from a handful of value debates is too coarse for the space of real descents. The companion descent (the white lie) cut the other way and *shrank* an everyday value clash to a small fittingness-parity leaf, one point for the wager on object-level material. Full treatment: [research/deep-descent-experiment.md](research/deep-descent-experiment.md).\n\n**The wager has a precondition, and it is a disposition rather than a mechanism (Aug 2026).** Everything above concerns what the graph finds at depth. A separate question is who descends, and the psychology literature answers it unfavorably: Kahan's programme found that *every* measure of reasoning proficiency he tried made polarization on contested questions **worse**, because comprehension capacity serves motivated reasoning rather than correcting it. Deliberus is comprehension capacity, industrialized. The one disposition that reversed the pattern in his data was curiosity, which changed what people chose to read — science-curious partisans preferred the article that contradicted them, by 20 to 44 percentage points. So convergence is not merely a bet about bedrock; it is a bet about bedrock *conditional on curious readers*, and handed to motivated ones the same machinery produces better weapons. This does not weaken the falsifiability accounting above, it adds a variable the accounting does not currently isolate: a residue map built from incurious use would report ceilings that are really refusals. The picture is less fatalistic than it first reads, because the disposition is partly manufacturable — the curiosity research reports that trait curiosity largely reduces to appraising one's own ability to understand as high, which is exactly the quantity infrastructure can raise. So the precondition describes a loop rather than an external dependency: lowering the cost of understanding partly generates the willingness to pay it. What it does not do is change anyone's motive, so the honest version keeps both halves. Full treatment: [research/curiosity-as-growth-fuel.md](research/curiosity-as-growth-fuel.md).\n\n---\n\n## Parfit's Triple Theory\n\nDerek Parfit's *On What Matters* (2011) argues that the major ethical traditions are \"climbing the same mountain from different sides\": when properly developed, consequentialist, Kantian, and contractualist reasoning converge on a shared \"Triple Theory.\" This is not settled philosophy, but it is the right kind of signal for Deliberus: apparently incompatible ethical frames may share deeper structure once their compressed slogans are decomposed. See [research/trust-economics-and-false-polarization.md §9](research/trust-economics-and-false-polarization.md).\n\n---\n\n## The Lean Analogy at Its Deepest\n\nIn Lean, the axioms (type theory) are trusted, shared, and universal. Every proof in Mathlib verifies against the same foundations. Trust comes from the kernel's simplicity — ~5,000 lines of C++.\n\nIf the convergence thesis holds, Deliberus has the same structure:\n\n| Lean | Deliberus |\n|------|-----------|\n| Axioms: type theory (trusted, universal) | Axioms: human values (convergent at depth) |\n| Kernel: small, auditable | Value bedrock: consciousness, suffering, shared humanity |\n| Agents verify proofs against axioms | Agents verify reasoning against converged values |\n| Trust: kernel's simplicity | Trust: values' universality |\n\nAI alignment through this system isn't \"checking against what humans decided\" (culturally relative, majority-rule). It's verifying reasoning against the **natural convergence of human values at depth** — genuinely universal foundations, not imposed constraints.\n\n---\n\n## Tao's Verification Thesis\n\nTerence Tao's \"Mathematical Methods and Human Thought in the Age of AI\" ([arXiv:2603.26524](https://arxiv.org/abs/2603.26524), March 29, 2026) validates this architecture from mathematics:\n\n> \"AI tools should not be viewed purely through the technical lens, but through the macroscopic humanitarian lens of how our society, our shared body of knowledge and understanding, and our species benefits as a whole.\"\n\nHis core argument: ideas are cheap (AI generates thousands), verification is the bottleneck, formal infrastructure (Lean) provides structural guarantees, but meaning and value remain irreducibly human. Replace \"mathematical theorems\" with \"claims about the world\" and \"Lean\" with \"Deliberus\" — the structural analogy is exact.\n\nTao also warns about **AI collapse**: AI trained on recursively generated AI outputs degrades over iterations. \"Without a sufficient amount of genuine content, AI becomes ungrounded from reality.\" A human-maintained argument graph — genuine reasoning, structured and persistent — is immune to this recursive degradation.\n\n---\n\n## Constitutional AI as Snapshot, Deliberus as Camera\n\nConstitutional AI aligns models with explicitly stated normative principles. But recent research reveals fundamental limits: constitutions embed cultural assumptions ([Pourdavood, March 2026](https://arxiv.org/html/2603.28123v1)) and compliance is inconsistent ([MATS 9.0, March 2026](https://www.alignmentforum.org/posts/Tk4SF8qFdMrzGJGGw/)).\n\nA constitution is a **snapshot** — a fixed list of principles at a moment in time. Deliberus is the **camera** — the deliberative *process* that generates, tests, decomposes, and evolves those principles through structured argumentation. \"Constitutional convention infrastructure\" for AI values.\n\n---\n\n## Philosophical Grounding: Peirce and Rorty\n\nThe convergence thesis is Peircean at its core: truth as the limit of inquiry, where all inquirers converge given sufficient evidence and decomposition. When two people disagree on a probability, the disagreement is about something specific and decomposable — different evidence, different weighting, different priors, different frameworks. Rorty's complementary contribution is irony — the recognition that the analytical framework itself (scheme taxonomy, QEM weights, four-type classification) is contingent, not a discovery about the fabric of reality. The irony keeps you from dogmatism; the convergence keeps you building. The system is designed so that convergence is testable, not presupposed — if decomposition at depth finds irreducible parallax rather than shared bedrock, that discovery is itself valuable and the graph will have made it visible. (Full treatment: [research/truth-graph-evidence-system.md §Philosophical Grounding](research/truth-graph-evidence-system.md))\n\n---\n\n## The Landing Page Closing\n\nAll of this is what the final sentence on deliberus.com gestures at:\n\n> *So we can finally think together about what matters most — and find, at the deepest level, how much we share.*\n\n## What Stops the Descent (Aug 2026)\n\nThe wager says decomposition keeps going until what remains is small and nameable. A separate question is how far it *can* go at all — and the answer turns out to be further than the goods-lists suggest, then stopped by something specific.\n\nValue decomposes continuously: from value words, through the empirically derived goods lists this doc already cites, through affective primitives, down to a **finite set of physically regulated bodily variables**, and then to a single generative principle — self-maintenance, of which the free-energy formulations are the formal version. No level in that descent is where a different kind of stuff enters.\n\nWhat stops it is not running out of ingredients. It is **two floors of different kinds**. The *phenomenal* floor asks why any of this is felt at all — the deep-descent experiment already stood on it. The *normative* floor is reached with the ingredient list complete: \"X is regulated\" does not entail \"X ought to be.\" They are independent, and neither dissolves the other.\n\nWhich is what makes hedonic monism structurally interesting rather than merely bold: it is exactly the claim that **the two floors are one floor**. That claim is already a live, challengeable node in the graph.\n\nFull treatment, including two vocabulary repairs the corpus owed — *\"no common scale\"* and *\"local and contextual\"* both replaced with checkable formulations — and a falsifier sharper than the gallery's: [research/decomposing-value.md](research/decomposing-value.md).\n\n---\n\n## As Alignment Target\n\nThe convergence thesis is closest to Yudkowsky's Coherent Extrapolated Volition (CEV) — \"what humanity would want if we knew more, thought faster, were more the people we wished we were\" — but with a concrete mechanism rather than a thought experiment. Structured argument decomposition with scheme classification and critical questioning *approximates* the CEV process in practice. The alignment target doesn't need to be metaphysically objective; it needs to be humanly shared. See [research/deliberus-as-alignment-infrastructure.md §1.3](research/deliberus-as-alignment-infrastructure.md).\n\n**The convergence thesis as \"Hail Mary\"** (Apr 25, 2026): Simon's scientific-method frame (\"The Scientific Method in Two Steps,\" learningmechanics.pub, Apr 2026) celebrates bold testable hypotheses even when disproven. The convergence thesis is the project's boldest Hail Mary and should be stated with explicit falsification conditions: in topic domain X, decompose claims from worldview groups A and B to depth N; measure whether the deepest shared premises converge or diverge. Without such conditions, the thesis remains \"beautiful and useful\" philosophy (Simon's Step A) rather than science (Steps A + B). See [research/scientific-method-two-steps-and-deliberus.md](research/scientific-method-two-steps-and-deliberus.md).\n\n**The wager, transformed: typed residues** (Jul 5, 2026): A six-agent research fan-out testing the thesis found it survives in a sharpened shape — decomposition converts vague weighting-disagreements into *typed, narrow, maximally legible residues* (fittingness claims, structural questions, axiom choices under proven impossibility, risk-function choices within a permissible band), which remain arguable in a different register and empirically account for a minority of real-world disagreement. Legibility is the guaranteed floor; convergence is the falsifiable wager built on it. Population ethics is the existence proof of the endgame: full decomposition into formally mapped axioms plus an impossibility theorem — bedrock-shaped, but maximally legible. **First data (Jul 2026)**: the live residue map's first two classified termini split 1–1 — one genuine fittingness residue (retribution), one value-dressed claim that dissolved into a checkable empirical question (assisted-dying consequences) — with the propose-only classifier discriminating correctly in both directions on blind reads. Two points are an existence proof of honest scoring, not a verdict. See [research/dogfood-run-2-orthogonal-experiments.md](research/dogfood-run-2-orthogonal-experiments.md). Full synthesis: [research/convergence-wager-typed-residues-synthesis.md](research/convergence-wager-typed-residues-synthesis.md); the weighting analysis: [research/value-weighting-decomposition.md](research/value-weighting-decomposition.md); the empirical record: [research/empirical-deliberation-convergence-evidence.md](research/empirical-deliberation-convergence-evidence.md); the metaethics survey: [research/moral-convergence-metaethics-and-psychology.md](research/moral-convergence-metaethics-and-psychology.md); the standing attacks: [research/convergence-wager-red-team.md](research/convergence-wager-red-team.md).\n\n**Residue, crux, double-crux — and the hinge score** (Jul 6, 2026, from the first residue experiment): three things the vocabulary must keep apart. A *residue* is where one branch bottoms out; a *crux* is a claim whose flip would flip a conclusion; a *double-crux* is a crux both sides acknowledge as their hinge — a social achievement the graph can invite but never decide. Crux-hood turns out to be computable: the hinge score (`/claims/{id}/hinge`) clamps each claim in a conclusion's decomposition subtree to fully-granted and fully-denied and measures how far the conclusion's QBAF strength swings — so \"which pebble does this debate actually stand on?\" gets a ranked, empirical answer. A typed residue with a high hinge is a load-bearing pebble. First experiment, live readings, and the QEM property that keeps scaffolding from burying a crux: [research/dogfood-run-1-friction-log.md §I2/§I4](research/dogfood-run-1-friction-log.md).\n\n**What a residue is, and what the map is for** (Jul 9, 2026): a residue is a *fallible fixed point*, not proven exhaustion. Proving that no further decomposition exists is impossible in principle (it would mean proving nobody will ever have a creative idea, Deutsch's point), so a terminus marks where a branch stopped splitting *under the current move-set applied seriously*, and the verdict is itself a challengeable claim that new moves can reopen. The map earns its keep even for residues that never resolve, four ways. (1) The compression is the product: two book-length hostile texts becoming one precisely stated claim transforms the disagreement socially even if the leaf never moves; everything above it becomes common ground or homework. (2) Decomposition is only one resolution operator: the typed leaf is where reframing, analogy transfer to already-settled residues, scope-narrowing, and the handoff from analysis to attunement get their exact coordinates. (3) The map is the wager's scoreboard: without it, \"values are irreconcilable\" and \"it's all confusion\" both remain vibes. (4) A well-stated open problem is one of the most productive artifacts a knowledge community can own (Hilbert's 23 problems organized a century of mathematics); the residue map is that list for normative disagreement. The population-ethics theorem, read this way, is not a wall the acid hit but *what completed decomposition looks like*: the one debate humanity has decomposed all the way down yielded not an answer but a proven map of the trade-space, a finite menu of exits each with a demonstrated price, and the theorem itself stays criticizable through its premises (transitivity, completeness, person-affecting restrictions are all live exit routes in the literature). Design implication, held lightly: terminus verdicts could carry *moves-tried* metadata so every residue is maximally reopenable.\n\n**Full technical exploration**: [research/ai-safety-and-tao-augmentation-research.md](research/ai-safety-and-tao-augmentation-research.md) — MGE, Constitutional AI limits, debate-based alignment, DCI paper, 80,000 Hours gap, Tao's full thesis. [research/deliberus-as-alignment-infrastructure.md](research/deliberus-as-alignment-infrastructure.md) — the deep case for Deliberus as alignment infrastructure. [research/non-zero-sum-economics-and-civilizational-cooperation.md](research/non-zero-sum-economics-and-civilizational-cooperation.md) — the economic and game-theoretic case for convergence as cooperation discovery.\n\n**The priors behind the conviction (2026-08-21)**: a cross-domain sweep found the strands where independent starting points converge — measured value-structure universality, cooperative-morality universals, autopoietic valence, reverse mathematics, universality classes — jointly predicting the *sharpened* wager (shared structure, contested weights) rather than the naive one: [research/fractal-priors-for-convergence.md](research/fractal-priors-for-convergence.md).\n\n**See also**: [Depth](depth.md) · [Bridging](bridging.md) · [Worldview Lenses](worldview-lenses.md) · [Civilizational Vision](civilizational-vision.md) · [Analysis ↔ Attunement](analysis-and-attunement.md) · [Lean Analogies](research/lean-deliberus-analogies.md)\n"}