{"path":"research/deliberus-as-alignment-infrastructure.md","content":"# Deliberus as Alignment Infrastructure\n\n**Date**: March 31, 2026\n**Purpose**: Deep research on how Deliberus could serve as AI safety infrastructure — not just relate to alignment, but BE an alignment tool.\n**Builds on**: [ai-safety-and-tao-augmentation-research.md](ai-safety-and-tao-augmentation-research.md) (initial survey)\n\n---\n\n> **July 2026 update**: LawZero's formal safety case for Scientist AI ([arXiv:2606.29657](https://arxiv.org/abs/2606.29657)) now documents, in its own text, the two boundaries this doc's oversight argument rests on: contested normative claims get no convergence guarantee (§2.2) and the safety threshold is \"a normative certification threshold fixed by the designer\" that the framework \"describes rather than attempts to resolve\" (Rem. 5.28). The machine-side architecture also independently converged on Deliberus's statement/provenance layer (\"epistemic contextualization\"). Full analysis: [bengio-safety-from-honesty-and-deliberus.md](bengio-safety-from-honesty-and-deliberus.md).\n\n## Executive Summary: The Case in One Page\n\nCurrent AI alignment approaches share a structural weakness: they treat human values as inputs to be extracted (RLHF), approximated (Constitutional AI), or inferred (inverse reinforcement learning) — but none provide persistent, evolving, publicly auditable infrastructure where the *reasoning behind* values is decomposed, debated, and refined.\n\nDeliberus could fill this gap. The argument:\n\n1. **The verification thesis** (Tao, 2026): Ideas are cheap; verification is the bottleneck. This applies to values as much as to mathematics. Anyone can assert a value. The hard problem is decomposing *why* that value matters, what it depends on, where it conflicts with other values, and what evidence supports it.\n\n2. **The convergence thesis** (founder's conviction, supported by empirical evidence from MGE and morality-as-cooperation research): Human values converge when decomposed to sufficient depth. This convergence — if real — is the natural alignment target. Not \"what most people voted for\" but \"what emerges when reasoning is made structurally transparent.\"\n\n3. **The Lean analogy taken seriously**: Lean provides type-checking for mathematical proofs. Deliberus provides scheme-checking for human reasoning — every argument classified by Walton scheme, every scheme's critical questions generated, every gap visible as a sorry marker. This is not metaphor; it is a structural parallel with concrete implementation.\n\n4. **The infrastructure gap**: Constitutional AI is a snapshot; RLHF is a noisy signal; debate-based alignment needs better judges; democratic input processes (Collective Constitutional AI, Polis) lack argument structure. Deliberus is the persistent, structured, community-maintained reasoning graph that each of these approaches needs but none provides.\n\n5. **The specific mechanism**: AI systems could verify their ethical reasoning against a Deliberus graph the way AlphaProof verifies mathematical reasoning against Lean — not as a hard constraint, but as a structured reference that makes the *basis* of any ethical claim transparent, questionable, and improvable.\n\nThe strongest version of this claim: Deliberus is not a tool for alignment research. It is alignment infrastructure — the epistemic substrate that makes alignment *possible* at civilizational scale.\n\nThe honest caveat: This vision requires solving several hard problems (value lock-in, extraction bias, gaming resistance, the is-ought gap) that may prove intractable. The research below explores both the promise and the limits.\n\n---\n\n## 1. Deliberus as Alignment Infrastructure — The Specific Mechanisms\n\n### 1.1 Value Learning: From Statistical Extraction to Structured Decomposition\n\n**The problem with current approaches:**\n\nRLHF (Reinforcement Learning from Human Feedback) learns values from preference rankings — \"humans preferred output A over output B.\" This is statistically powerful but structurally impoverished. It captures *what* humans prefer without decomposing *why*. The result is a reward model that approximates human values without understanding them — vulnerable to reward hacking, distribution shift, and cultural bias baked into the preference data.\n\nConstitutional AI improves on this by making principles explicit, but constitutions are static snapshots written by small teams. As Pourdavood (arXiv:2603.28123, March 2026) demonstrates, constitutions embed cultural assumptions. As MATS 9.0 research shows, compliance with constitutions is inconsistent even when they exist.\n\n**What structured value learning looks like:**\n\nA Deliberus graph provides something neither RLHF nor Constitutional AI can: the *reasoning structure* behind values. Consider the value \"privacy matters.\" In RLHF, this is a preference signal. In Constitutional AI, this is a principle statement. In Deliberus:\n\n- The claim \"privacy matters\" is classified as a **value premise**\n- It decomposes into supporting claims: \"autonomous individuals require informational self-determination\" (value premise), \"surveillance chills free expression\" (empirical claim with evidence), \"power asymmetries corrupt\" (value premise with historical evidence)\n- Each supporting claim has its own decomposition, critical questions, and connections to opposing arguments\n- The entire structure is navigable through different worldview filters — a libertarian lens weights autonomy differently than a communitarian lens, but both see the full reasoning structure\n\n**The interface for AI value learning:**\n\nAn AI system learning from a Deliberus graph would not learn \"privacy: good\" (the RLHF signal). It would learn:\n\n1. The argument *structure* supporting privacy (which claims, which schemes, which evidence)\n2. The conditions under which privacy competes with other values (security, transparency, public health)\n3. The critical questions that challenge privacy claims (and whether those questions have been answered)\n4. The worldview-dependent weights that different communities assign to the same reasoning\n\nThis is structurally richer than any current value learning approach. The closest precedent is **Moral Graph Elicitation** (MGE, Benkler et al., arXiv:2404.10636, 2024), which maps directional edges between values in context (\"in this situation, value A is wiser than value B\"). MGE found that participants *overwhelmingly converge* on these directional relationships — empirical support for the convergence thesis. Deliberus extends MGE from pairwise value comparisons to full argument graph topology with scheme classification and critical question decomposition.\n\n**Comparison to existing approaches:**\n\n| Approach | What it captures | What it misses |\n|---|---|---|\n| RLHF | Preference rankings | Why preferences exist; structure of reasoning |\n| Constitutional AI | Explicit principles | Reasoning behind principles; how principles interact |\n| Inverse RL (CIRL) | Behavior-inferred values | Explicit reasoning; value decomposition |\n| MGE (Benkler) | Pairwise value rankings in context | Argument structure; scheme classification; CQs |\n| **Deliberus** | Full argument topology with schemes, CQs, evidence, worldview weights | Requires explicit reasoning (cannot infer from behavior) |\n\n### 1.2 Reasoning Verification: The Lean Analogy Made Concrete\n\n**The structural parallel:**\n\nIn Lean, a proof is a sequence of steps. Each step must be justified by an axiom, a previously proven theorem, or a definition. The type-checker verifies that every step follows from what came before. Gaps are marked with `sorry` — visible, actionable invitations to fill in missing reasoning.\n\nIn Deliberus, an argument is a structure of claims. Each claim has a type (empirical, normative, definitional, value premise). Each relationship between claims is classified by Walton scheme (argument from expert opinion, argument from consequence, argument from analogy, etc.). Each scheme has critical questions that must be answered for the argument to hold. Unanswered critical questions are sorry markers.\n\n**What \"type-checking\" for ethical reasoning would look like:**\n\nAn AI system making an ethical judgment (\"it is permissible to break this promise because the consequences of keeping it are worse\") could be verified against a Deliberus graph:\n\n1. **Scheme identification**: This is an \"argument from consequences\" (Walton scheme). The critical questions include: \"Are the consequences accurately predicted?\" \"Are all relevant consequences considered?\" \"Is consequentialist reasoning appropriate here, or are there deontological constraints?\"\n2. **Premise verification**: The empirical claims about consequences can be checked against evidence nodes in the graph. The normative claim that consequences outweigh promise-keeping can be traced to its supporting argument structure.\n3. **Gap detection**: If the AI's reasoning skips a critical question — e.g., assumes consequentialism without addressing deontological objections — the Deliberus graph reveals this as a sorry marker, a structural gap in the reasoning.\n\nThis is not the same as mathematical proof verification. The is-ought gap means you cannot *prove* an ethical conclusion the way you prove a theorem (see Section 3.3). But you can make the *structure* of ethical reasoning transparent, identify gaps, and flag when an AI system is reasoning from unexamined assumptions.\n\n**Existing work on formal verification of ethical reasoning:**\n\n- **Deontic Temporal Logic** (Priya & Rao, arXiv:2501.05765, 2025): Proposes DTL for formal verification of AI ethical compliance. Uses temporal operators (always, eventually, until) combined with deontic operators (obligatory, permissible, forbidden) to specify and verify ethical constraints. This works for *codified* ethical rules (regulations, company policies) but cannot handle the open-ended reasoning Deliberus targets.\n\n- **The Ethical Compiler** (Bowman, UBC): Addresses the is-ought gap directly in compilation — formal methods can verify that a system *correctly implements* a given ethical specification, but cannot verify that the specification itself is ethically correct. This is the fundamental limit: Deliberus can verify reasoning *structure* but not reasoning *truth*.\n\n- **Computational Complexity of Ethics** (Stenseke, arXiv:2302.04218, 2024): Surveys the computational tractability of moral reasoning. Key finding: many ethical decision problems are NP-hard or worse — morally optimal action under uncertainty is computationally intractable in general. This suggests that decomposition (breaking hard moral problems into tractable sub-problems) is not just useful but *necessary* — exactly what argument scheme classification does.\n\n### 1.3 The Convergence Thesis as Alignment Target\n\n**The claim:** If human values converge when decomposed to sufficient depth, that convergence IS the alignment target.\n\nThis is a strong claim. It differs from other alignment targets:\n\n- **RLHF target**: What humans currently prefer (statistical, noisy, culturally biased)\n- **Constitutional AI target**: What a small team wrote down (snapshot, arbitrary, culturally embedded)\n- **CEV target**: What humanity *would* want if we \"knew more, thought faster, were more the people we wished we were\" (Yudkowsky, 2004)\n\nDeliberus's convergence thesis is closest to CEV but differs in a crucial way: CEV is a *thought experiment* about hypothetical ideal reasoners. Deliberus proposes a *concrete mechanism* — structured argument decomposition with scheme classification and critical questioning — that approximates the CEV process in practice. Not perfectly, but instrumentally.\n\n**Philosophical support for convergence:**\n\nDerek Parfit's magnum opus *On What Matters* (2011, three volumes) argues that the three major ethical traditions — consequentialism, Kantianism, and contractualism — converge on the same conclusions when properly understood. His \"Triple Theory\" claims these are \"climbing the same mountain on different sides.\" This is the philosophical version of the convergence thesis: seemingly irreconcilable moral frameworks produce the same judgments when decomposed to fundamental principles.\n\nCritics exist. Walen (\"Separate Peaks,\" 2024) argues Parfit's convergence requires gerrymandering each tradition beyond recognition. Baumann (\"In Search of the Trinity,\" 2021) identifies structural dilemmas in the conciliatory project. The philosophical jury is split.\n\n**Empirical support for convergence:**\n\n- **Morality-as-Cooperation** (Curry, Mullins & Whitehouse, 2019; Alfano, Cheong & Curry, Heliyon 2024): Machine-reading analysis of 256 societies finds seven moral values (family, group loyalty, reciprocity, bravery, respect, fairness, property rights) present across all studied cultures. These are not identical in weight or application, but they are universal in presence. This supports a weaker version of convergence: not that values are *identical* across cultures, but that they share a common structure.\n\n- **MGE empirical convergence** (Benkler et al., 2024): Participants overwhelmingly agree on directional value relationships when reasoning is structured. This is the strongest direct evidence: convergence *emerges from* structured decomposition, not from shared culture or demographics.\n\n- **Cross-cultural moral foundations** (Eriksson et al., 2021): Different populations agree on *which* moral arguments underlie *which* opinions, even when they disagree on the opinions themselves. The argument structure is shared even when conclusions differ.\n\n**The \"bedrock of consciousness\" claim:**\n\nThe founder's strongest version — that values converge on \"pure consciousness and love and shared humanity\" — connects to philosophy of mind. The claim that consciousness is morally relevant is indeed close to universal (Knobe et al., \"Consciousness and Morality,\" Oxford Handbook): across cultures, the capacity for subjective experience is the primary criterion for moral status. Whether this constitutes a \"bedrock\" or merely a shared starting point is debatable.\n\nThe philosophical literature on consciousness and moral status (Sauer, 2019; \"The argument from agreement: How universal values undermine moral realism\") notes an irony: the *universality* of certain moral intuitions might actually undermine moral realism. If we all agree that suffering matters because of shared evolutionary history, that's a *causal* explanation, not a *justificatory* one. The convergence could be an artifact of shared biology rather than evidence of objective moral facts.\n\n**For Deliberus, this distinction may not matter:** Whether convergence reflects objective moral facts or shared human nature, the practical result — a navigable structure of human reasoning that reveals common ground — serves alignment equally well. The alignment target doesn't need to be metaphysically objective; it needs to be *humanly shared*.\n\n### 1.4 Constitutional Convention for AI\n\n**The existing approach — Collective Constitutional AI (Anthropic + CIP, 2023):**\n\nAnthropic ran a public input process using Polis to generate a constitution for Claude. ~1,000 Americans voted on AI behavioral principles. The resulting \"publicly-sourced\" constitution produced a model that was slightly less biased and slightly more resistant to certain jailbreaks than Anthropic's default constitution.\n\n**Limitations:**\n- Polis captures opinion clusters, not argument structure. You know *what* people agree on but not *why*.\n- The process was a one-shot survey, not an evolving deliberation.\n- ~1,000 American participants is not humanity.\n- The resulting constitution is still a static snapshot.\n\n**Stanford's \"People Shaping AI\" (2026):**\n\nA more ambitious effort: Stanford's Deliberative Democracy Lab, in collaboration with Anthropic, Google, Meta, Microsoft, OpenAI, and xAI, invited public input on the future of AI agents through structured deliberation (announced February 2026). This is the most institutionally significant democratic AI governance initiative to date.\n\n**What Deliberus would add:**\n\nA Deliberus-powered constitutional convention would differ from both approaches:\n\n1. **Argument structure, not just opinion clusters**: Every principle would come with its decomposed reasoning — the claims it rests on, the evidence supporting those claims, the critical questions that challenge them, and the competing principles it trades off against.\n\n2. **Persistent and evolving**: Not a one-shot process but a continuously growing argument graph. As new evidence emerges, as AI capabilities change, as edge cases surface, the constitutional reasoning evolves.\n\n3. **Worldview transparency**: Through the worldview filter, different communities could see which principles their values support and which they contest — without forcing premature consensus.\n\n4. **Bridging arguments as the novel signal**: Instead of optimizing for majority agreement (which suppresses minority perspectives), Deliberus could surface *bridging arguments* — reasoning that crosses opinion divides. A constitutional principle supported by bridging arguments has deeper legitimacy than one supported by a 51% majority.\n\n5. **Sorry markers as democratic invitation**: Unresolved critical questions in the constitutional reasoning are visible gaps — explicit invitations for communities to contribute their perspective, not just their vote.\n\n### 1.5 Scalable Oversight via Argument Structure\n\n**The debate alignment literature's core problem:**\n\nIrving et al. (2018) propose AI debate as a scalable oversight mechanism: two AI systems debate, a human judges. If the debate protocol incentivizes honesty, even a weak human judge can oversee a superhuman AI.\n\nThe problem: this depends on judge quality. Brown-Cohen & Irving (AI Alignment Forum, 2025) address the \"obfuscated arguments\" problem — where a dishonest debater adversarially constructs arguments that sound plausible but are wrong. Their Prover-Estimator protocol provides theoretical guarantees, but the practical challenge remains: how do you help human judges evaluate complex, multi-step arguments?\n\n**Empirical evidence that argument structure helps judges:**\n\n- **\"AI Debate Aids Assessment of Controversial Claims\"** (Rahman et al., arXiv:2506.02175, 2025; presented at ICML workshop): Studies how AI-generated debates help humans assess controversial claims. Finding: structured debate — where AI systems present arguments and counterarguments — significantly improves human assessment accuracy compared to single-perspective summaries. This directly validates the \"judge enhancement\" thesis.\n\n- **Argument mapping improves reasoning** (van Gelder, 2015; Cullen et al., 2018, npj Science of Learning): Argument visualization produces large effect sizes (d = 0.72) on critical thinking measures — significantly larger than traditional instruction. The mechanism is exactly what Deliberus provides: making argument structure visible transforms implicit reasoning into explicit, evaluable structure.\n\n- **Intelligence analysis** (Kruger, 2022, Intelligence and National Security): Argument mapping improves clarity and rigour in written intelligence products. If it works for intelligence analysts evaluating national security claims, it should work for judges evaluating AI reasoning.\n\n- **\"Can LLMs Judge Debates?\"** (Sanayei et al., EMNLP 2025): Evaluates LLMs as debate judges using argumentation theory semantics. Finding: formal argumentation structure (attack/support relations, Dung's semantics) significantly improves judge accuracy over unstructured evaluation. This validates using argumentation-theoretic structure — exactly what Deliberus provides — as judge enhancement infrastructure.\n\n**The concrete mechanism:**\n\nA Deliberus-enhanced judge in an AI debate would see:\n\n1. Each AI's argument decomposed into atomic claims with scheme classifications\n2. Critical questions automatically generated for each argument scheme\n3. Cross-references to the existing argument graph — where does this reasoning connect to known arguments?\n4. Sorry markers highlighting structural gaps in each debater's reasoning\n5. Similarity links to related arguments from different contexts\n\nThis transforms judging from \"which argument sounds more convincing?\" to \"which argument has fewer structural gaps?\" — a much more tractable cognitive task.\n\n**Process-based oversight alignment:**\n\nOpenAI's research on process supervision (Lightman et al., 2023, \"Improving mathematical reasoning with process supervision\") found that rewarding correct *reasoning steps* — not just correct answers — dramatically improves model performance and reliability. This \"supervise process, not outcomes\" principle (Stuhlmuller & Jung, 2022, AI Alignment Forum) maps directly to Deliberus's approach: the platform evaluates the *structure* of reasoning (schemes, critical questions, evidence connections), not just the *conclusion*.\n\n---\n\n## 2. The Convergence Thesis in Philosophy\n\n### 2.1 Moral Realism and the Convergence Claim\n\nThe founder's conviction — that decomposing values far enough reveals shared bedrock — is a form of what philosophers call **procedural moral realism**: the view that correct moral reasoning, if pursued far enough, converges on determinate answers.\n\nThe philosophical landscape (2024 PhilPapers survey: 62% of professional philosophers identify as moral realists):\n\n**Strong realism (convergence as discovery):** Moral facts exist independent of human opinion. Deep decomposition *discovers* them. This is the strongest version of the convergence thesis — and the most philosophically contested. Sharon Street's \"Darwinian Dilemma\" (Philosophical Studies, 2006) argues that evolutionary pressures shaped our moral intuitions for survival, not truth. If our sense that \"suffering matters\" is an evolved response rather than a perception of moral reality, convergence proves nothing about objective moral facts.\n\n**Parfit's conciliatory project (convergence as reconciliation):** The three major ethical traditions agree at their best — \"climbing the same mountain on different sides.\" Parfit himself was a moral realist who believed his Triple Theory revealed objective moral truths. But even if the metaphysics is wrong, the *practical convergence* is significant for alignment: if consequentialism, Kantianism, and contractualism agree on most cases, the alignment target is well-defined for those cases.\n\n**Constructivism (convergence as construction):** Moral facts are constructed through rational procedures. This is actually the most natural fit for Deliberus: the platform doesn't claim to *discover* moral truths but to *construct* well-justified moral reasoning through structured deliberation. The convergence, if it occurs, emerges from the process — which is exactly what the MGE data shows.\n\n**Epistemological significance of convergence:** A 2025 paper directly addresses this: \"On the Epistemic Significance of Convergence in Ethical Theory\" (Ethical Theory and Moral Practice, 2025). The analysis suggests that convergence between independent ethical theories provides *some* evidence for the truth of shared conclusions — but the strength of this evidence depends on the degree of genuine independence between the converging theories. If Kantianism, consequentialism, and contractualism share historical influences (as they do), convergence is less surprising and less evidentially significant.\n\n### 2.2 Coherent Extrapolated Volition (CEV) — Is Deliberus a CEV Machine?\n\nYudkowsky's CEV (2004) proposes aligning AI with what humanity *would want* \"if we knew more, thought faster, were more the people we wished we were, had grown up farther together.\" The key features:\n\n1. **Extrapolation, not polling**: Not current preferences but *improved* preferences\n2. **Coherence**: The extrapolated volitions of different people should be compatible\n3. **Community**: \"Grown up farther together\" — a shared deliberative process\n4. **Dynamic**: Not a fixed target but an evolving one\n\n**Deliberus as CEV approximation:**\n\nThe structural parallels are striking:\n\n- **\"Knew more\"** → Deliberus's evidence nodes and auto-connection make relevant evidence visible. Each extraction enriches the graph with new information.\n- **\"Thought faster\"** → AI-assisted decomposition (8-pass extraction pipeline) does in minutes what would take humans hours — the argument structure emerges rapidly.\n- **\"Were more the people we wished we were\"** → The sorry model reveals gaps in our reasoning. Critical questions surface assumptions we haven't examined. The platform is explicitly designed to improve reasoning quality.\n- **\"Grown up farther together\"** → The accumulating graph is a shared epistemic commons. Each contribution builds on and connects to existing reasoning. This IS growing up farther together, epistemically.\n\n**Critical difference:** CEV is a thought experiment about a hypothetical process. Deliberus proposes a *concrete mechanism* — argument decomposition with scheme classification — that partially instantiates that process. The gap between \"partially instantiates\" and \"fully implements\" is where the honest assessment must happen (see Section 3).\n\n**Yudkowsky's own concerns:** CEV has been criticized within the alignment community as underspecified. The concept of \"extrapolation\" is doing enormous philosophical work — what counts as \"knowing more\"? Whose direction of growth counts? Yudkowsky himself noted (in the original CEV paper) that if human values are fundamentally incoherent — if there is no convergence point — then CEV cannot produce a well-defined target. Deliberus's extraction pipeline would *reveal* this incoherence rather than assume it away — which is itself valuable information for alignment.\n\n### 2.3 Moral Convergence Literature — What the Evidence Actually Shows\n\n**Hopster (2019), \"Explaining historical moral convergence\":** Observes a global historical trend toward liberal moral values (expanding circles of concern, rights-based frameworks). Argues this convergence is better explained by anti-realist mechanisms (power dynamics, economic development, information flow) than by realist ones (moral discovery). Key insight for Deliberus: convergence is real but its explanation matters. If convergence is driven by contingent historical forces rather than rational decomposition, Deliberus's argument structure adds genuine value — it transforms accidental convergence into *reasoned* convergence.\n\n**Hassan (2019), \"Moral Disagreement and Arational Convergence\":** Argues that some moral convergence is \"arational\" — driven by empathy, exposure, and emotional development rather than logical argument. This is actually the attunement pole of Deliberus's core dialectic: convergence through perspective-taking and emotional inhabitation, not just analytical decomposition. The platform needs both.\n\n**The morality-as-cooperation framework** (Curry et al., 2019, 2024): Provides the strongest empirical case for universal moral structure. Seven types of cooperative behavior (family, group, reciprocity, bravery, respect, fairness, property) are valued across all 60 (and later 256) societies studied. This doesn't mean identical values — the *weight* and *application* vary enormously — but the *structure* is shared. For Deliberus, this suggests the argument graph should reveal common structural patterns even when surface-level conclusions differ.\n\n### 2.4 The \"Bedrock of Consciousness\" — Philosophical Assessment\n\nThe claim that values converge on consciousness as foundational:\n\n**Support:** The near-universal attribution of moral status to conscious beings (Knobe et al., Oxford Handbook of Philosophy of Consciousness). The growing scientific consensus on animal consciousness (New York Declaration on Animal Consciousness, 2024, signed by dozens of consciousness researchers). The philosophical tradition from Bentham (\"Can they suffer?\") through Singer to contemporary sentientism.\n\n**Challenge:** \"Consciousness matters\" is genuinely cross-cultural, but the *conclusions* drawn from it diverge radically. Buddhist, Christian, utilitarian, and Kantian traditions all value consciousness but derive very different ethical systems from it. The bedrock may be real but underdeterminate — too thin a foundation to resolve most ethical disagreements.\n\n**For alignment:** Even a thin shared bedrock is valuable. If AI systems can be aligned to \"consciousness matters, suffering is bad\" as a starting point, and the Deliberus graph reveals how different traditions build from that shared foundation to different conclusions, the resulting alignment is much richer than any current approach provides.\n\n---\n\n## 3. Risks and Failure Modes — Honest Assessment\n\n### 3.1 Value Lock-in\n\n**The risk:** A mature Deliberus graph used for AI alignment could crystallize the values of early contributors. If the first 10,000 contributors are predominantly WEIRD (Western, Educated, Industrialized, Rich, Democratic), the graph's argument structure — what gets decomposed, which schemes are recognized, what counts as evidence — reflects their epistemic norms.\n\n**The lock-in mechanism:** It's subtle. The graph doesn't lock in *conclusions* (the worldview filter prevents that). It locks in *what counts as an argument*. Walton's 96 schemes were developed within Western analytic philosophy. Non-Western argumentation traditions (analogical reasoning in Chinese philosophy, narrative reasoning in indigenous traditions, dialectical reasoning in Buddhist logic) may not map cleanly onto Walton's taxonomy.\n\n**Research on value lock-in risk:** Qiu et al. (\"The Lock-in Hypothesis: Stagnation by Algorithm,\" OpenReview, 2025) demonstrate that LLM training creates feedback loops where AI-generated content reinforces existing values. Manifund has funded specific research on \"Operationalizing Value Lock-in from Frontier AI Systems\" and \"Moral Progress in AI to Prevent Premature Value Lock-in\" — indicating the alignment community takes this risk seriously.\n\n**Mitigations:**\n\n1. **Temporal versioning**: Deliberus should track how the graph evolves over time, allowing comparison between argument structures at different stages. If early contributions dominate the topology, this becomes visible.\n\n2. **Demographic metadata**: Track (anonymously) the demographic distribution of contributors. Alert when the graph's coverage is skewed.\n\n3. **Scheme extension**: Design the scheme classification to be extensible. Non-Western argumentation patterns should be first-class schemes, not forced into Western categories.\n\n4. **Deprecation over deletion** (from the Lean analogy): Never remove arguments from the graph. Mark them as superseded, questioned, or historically contingent. The full history of reasoning should be preserved.\n\n5. **Adversarial red-teaming**: Deliberately extract arguments from traditions that challenge the graph's current structure. Feed the pipeline Islamic jurisprudence, Confucian ethics, Ubuntu philosophy, indigenous knowledge systems — not to check boxes but to stress-test whether the graph's structure can accommodate genuinely different reasoning traditions.\n\n### 3.2 The \"Who Decomposes?\" Problem — Extraction Bias\n\n**The risk:** The 8-pass extraction pipeline uses Gemini 3 Flash. LLMs have documented political biases (Chen et al., arXiv:2601.08785, 2026; Yoo & Shin, EMNLP 2025). More fundamentally, \"source framing triggers systematic bias in large language models\" (Matschke et al., Science Advances, 2025) — LLMs evaluate the same text differently depending on its attributed source.\n\n**The specific danger:** When the pipeline classifies a claim as \"empirical\" vs \"normative\" vs \"value premise,\" that classification shapes how the graph treats it. If the LLM systematically misclassifies certain cultural perspectives — treating one tradition's value premises as empirical claims (requiring evidence) while treating another's as normative claims (requiring argument) — the resulting graph embeds a bias that's invisible in the graph structure itself.\n\n**This is worse than standard LLM bias** because it operates at the *structural* level. Users can see biased content and correct it. Users cannot easily see that the *classification system* is biased — that certain kinds of reasoning are systematically filed in categories that make them harder to defend.\n\n**Mitigations:**\n\n1. **Multi-model extraction**: Run the pipeline through multiple LLMs with different training data and compare classifications. Disagreements flag potential bias.\n\n2. **Human-in-the-loop for classifications**: The pipeline currently classifies automatically. For alignment-critical arguments, require human review of type classifications and scheme assignments.\n\n3. **Classification auditing**: Periodically analyze whether certain topics, traditions, or perspectives are systematically classified differently. Statistical analysis of classification patterns by source domain.\n\n4. **Extraction transparency**: Every classification should be traceable to the LLM's reasoning. Why was this claim classified as empirical rather than normative? The pipeline should log classification rationale, not just classification labels.\n\n### 3.3 The Is-Ought Gap in Verification\n\n**The fundamental limit:** David Hume's guillotine (1739): you cannot derive an \"ought\" from an \"is.\" No amount of factual decomposition logically entails a normative conclusion. This means Deliberus can verify the *structure* of ethical reasoning but not its *soundness* in the way Lean verifies mathematical proofs.\n\n**What this means concretely:**\n\n- Lean can verify: \"If axioms A hold, then theorem T follows.\" The axioms are assumed; the derivation is checked.\n- Deliberus can verify: \"If value premises V and empirical claims E hold, then normative conclusion N is supported via argument scheme S.\" The premises are tracked; the reasoning structure is checked.\n- But Deliberus CANNOT verify: \"Value premise V is true.\" Value premises are the axioms of ethical reasoning — they can be decomposed, questioned, compared, and contextualized, but not *proven*.\n\n**This is actually fine for alignment.** The is-ought gap doesn't invalidate Deliberus as alignment infrastructure; it clarifies its role. The platform provides:\n\n1. **Transparency**: What value premises does this ethical conclusion rest on?\n2. **Consistency**: Are these value premises consistent with each other and with the reasoner's other commitments?\n3. **Completeness**: Have all relevant critical questions been addressed?\n4. **Coverage**: What alternative argument structures reach different conclusions from similar premises?\n\nThese are enormous contributions to alignment even without the ability to prove value premises true. The alternative — current AI systems that reason from implicit, unexamined value assumptions — is strictly worse.\n\n**Bowman's \"Ethical Compiler\" framing** is exactly right: we can formally verify that a system *correctly implements* a given ethical specification (the derivation from premises to conclusion). We cannot formally verify that the specification itself is correct (that the premises are true). But making the specification *explicit and decomposed* is the prerequisite for any human evaluation of its correctness.\n\n> **External evidence delimiting this boundary (2026):** SciencePedia / LCoT \"Inverse Knowledge Search\" (arXiv:2510.26854) is a heavyweight independent convergence on this doc's core premise — radical compression of reasoning / verification-as-bottleneck — and an at-scale proof of reasoning-corpus + cross-model-consensus verification + inverse-search + emergent cross-disciplinary graph. **But its entire reliability mechanism depends on a mechanically-checkable endpoint — i.e. it lives on the *is* side of exactly the §3.3 is-ought gap.** It therefore validates the *premise* and the *verifiable slice*, while empirically delimiting Deliberus's non-overlapping niche: the normative / contested / no-verifiable-endpoint space such methods structurally cannot enter. Premise-validation + niche-delimiter, never \"they solved our problem.\"\n\n### 3.4 Gaming at Civilizational Stakes\n\n**The risk:** If Deliberus becomes alignment infrastructure, the incentives to manipulate it become enormous. Nation-states, corporations, and ideological movements would have strong motivation to shape the argument graph in their favor.\n\n**Attack vectors:**\n\n1. **Sybil attacks**: Creating many accounts to flood the graph with supporting arguments for a particular position. Traditional platform problem, but with alignment stakes.\n\n2. **Sophisticated argument injection**: Using AI to generate plausible-sounding arguments that contain subtle logical errors or misleading framings. Unlike spam, these arguments would pass quality checks.\n\n3. **Strategic sorry exploitation**: Identifying sorry markers (gaps in reasoning) and filling them with arguments that subtly redirect the reasoning chain.\n\n4. **Extraction poisoning**: If the extraction pipeline is used to process adversarially crafted texts, the resulting claims could be designed to connect to existing graph nodes in misleading ways.\n\n5. **Scheme gaming**: Constructing arguments that technically satisfy a scheme's critical questions while violating the spirit of the scheme — formal compliance without genuine reasoning.\n\n**Defenses:**\n\n1. **Argument quality is structural, not reputational**: Unlike social media (where influence = follower count), Deliberus evaluates argument *structure*. A well-decomposed argument with answered critical questions is strong regardless of who submitted it. Gaming requires actually producing good arguments — which is partially self-defeating.\n\n2. **Critical question generation is automatic**: The pipeline generates CQs from detected schemes. An injected argument must survive its own CQs, which are generated independently.\n\n3. **Cross-extraction auto-connect reveals inconsistencies**: Arguments from different sources that connect to the same graph region create a cross-referencing web. Manipulated arguments that contradict the existing graph become visible through auto-connect edge conflicts.\n\n4. **The worldview filter prevents consensus capture**: Even if an attacker floods the graph with arguments supporting their position, the worldview filter reveals this as one perspective among many, not as settled truth.\n\n5. **Temporal versioning as forensics**: Tracking when arguments were added, by whom, and how the graph structure changed enables forensic analysis of coordinated manipulation campaigns.\n\n**Honest assessment:** These defenses reduce but don't eliminate gaming risk. A sufficiently sophisticated adversary with nation-state resources could generate high-quality arguments that survive structural checks. The question is whether Deliberus is *more* resistant to manipulation than alternatives (RLHF preference data, Constitutional AI principles written by small teams, Twitter discourse). The answer is almost certainly yes — the structural evaluation provides a higher bar than any existing approach.\n\n### 3.5 The Scale Problem\n\n**The risk:** A civilizational reasoning graph that covers all human values would be incomprehensibly large. With 318 claims from 7 sources, the graph is already rich. At Wikipedia scale (60 million articles worth of reasoning), navigation becomes intractable.\n\n**Mitigation:** This is where the worldview filter, the sorry model, and the feed algorithm earn their keep. The graph doesn't need to be navigable in its entirety — it needs to be navigable *from any given perspective, for any given question*. The feed algorithm's 8 modes (including bridging detection) surface the most relevant arguments for any user in any context. The sorry model directs attention to the most important gaps. Gradual semantics (QEM, Potyka 2018) computes argument strength locally.\n\n---\n\n## 4. Adjacent Alignment Approaches — Detailed Comparison\n\n### 4.1 Recursive Reward Modeling (Leike et al., 2018)\n\n**Core idea:** Decompose complex tasks into simpler sub-tasks that humans can evaluate. A reward model for a complex task is built from reward models for its sub-tasks.\n\n**Connection to Deliberus:** RRM's decomposition is conceptually identical to Deliberus's argument decomposition — both break complex judgments into evaluable components. But RRM decomposes *tasks* while Deliberus decomposes *reasoning*. RRM asks \"can a human evaluate this sub-task?\" Deliberus asks \"can a human evaluate this sub-argument?\"\n\n**What Deliberus adds:** RRM assumes the decomposition is correct — if a task is wrongly decomposed, the reward model is wrong. Deliberus's scheme classification and critical question generation provide a *structured* decomposition validated against argumentation theory. The decomposition itself is checkable.\n\n### 4.2 Iterated Amplification (Christiano, 2018)\n\n**Core idea:** A weak human + AI assistant can solve problems that neither can solve alone. Iterate this: use amplified human+AI teams to train better AI assistants, which amplify humans further.\n\n**Connection to Deliberus:** Iterated amplification is exactly what the `@[simp]` flywheel does: each extraction adds knowledge to the graph, which makes subsequent extractions richer (auto-connect finds more connections), which makes the graph more useful, which enables better human judgment. The Deliberus graph IS the accumulated amplification.\n\n**What Deliberus adds:** Iterated amplification in its original formulation is opaque — the amplified capability is embedded in model weights, not visible to inspection. Deliberus's amplification is transparent: every connection, every scheme classification, every critical question is explicitly represented in the graph. If the amplification goes wrong, you can trace where.\n\n### 4.3 Cooperative Inverse Reinforcement Learning (Hadfield-Menell et al., 2016)\n\n**Core idea:** The AI and human are playing a cooperative game where the AI doesn't know the human's reward function and must learn it through interaction. The AI's uncertainty about human values makes it deferential — it asks rather than assumes.\n\n**Connection to Deliberus:** CIRL's key insight — that the AI should be *uncertain* about human values and should *actively learn* them through interaction — maps to Deliberus's sorry model. Every sorry marker is an explicit uncertainty: \"this reasoning has a gap.\" The sorry model makes value uncertainty structural and visible.\n\n**What Deliberus adds:** CIRL infers values from behavior (what the human does). Deliberus learns values from *reasoning* (what arguments the human finds compelling and why). Behavioral inference is vulnerable to revealed preference problems — people don't always act according to their values. Reasoning structures capture stated values and their justifications, which is richer even if imperfect.\n\n### 4.4 Moral Parliament (Newberry & Ord, FHI, 2021)\n\n**Core idea:** Under moral uncertainty, run multiple ethical theories in parallel (like a parliament of moral theories). Each theory \"votes\" on actions proportional to the credence assigned to it. The result is a weighted aggregate that respects uncertainty.\n\n**Connection to Deliberus:** The worldview filter IS a moral parliament. Different value weightings — utilitarian, deontological, virtue-ethical, communitarian — navigate the same argument graph and produce different judgments. The graph structure persists; the evaluative lens varies.\n\n**What Deliberus adds:** The Moral Parliament treats ethical theories as black boxes that output votes. Deliberus decomposes the *reasoning within* each theory: what premises does consequentialism rely on? Where does Kantianism's reasoning depend on empirical assumptions? The worldview filter shows not just *that* theories disagree but *where and why* — which specific premises or value weights produce the divergence. This is much more informative than a vote count.\n\n### 4.5 Process-Based Oversight\n\n**Core idea:** Evaluate the *process* of AI reasoning, not just the *outcome*. A process that is transparent, well-structured, and follows valid reasoning steps is more trustworthy than one that produces the right answer by opaque means.\n\n**Connection to Deliberus:** This IS Deliberus's core contribution. The entire platform is about making the *process* of reasoning visible: what claims, what evidence, what argument schemes, what critical questions. An AI system whose reasoning passes through a Deliberus-like structure is process-supervised by design.\n\n**What Deliberus adds:** Process-based oversight research (Lightman et al., 2023; Stuhlmuller & Jung, 2022) has focused on mathematical and logical reasoning where process correctness can be verified step by step. Ethical reasoning adds the complication that value premises cannot be verified — they can only be decomposed and questioned. Deliberus extends process-based oversight to the normative domain by providing structure (schemes, CQs) without claiming verification in the mathematical sense.\n\n---\n\n## 5. The \"Judge Infrastructure\" Angle\n\n### 5.1 What Makes a Good Judge in AI Debate?\n\nThe debate alignment literature identifies several requirements for effective judges:\n\n1. **Ability to follow multi-step arguments** without losing track of premises\n2. **Resistance to rhetorical manipulation** — evaluating argument structure rather than persuasiveness\n3. **Domain knowledge sufficient** to evaluate key claims (or tools to compensate for missing knowledge)\n4. **Awareness of common logical fallacies** and biased reasoning patterns\n5. **Capacity for perspective-taking** — understanding arguments from worldviews different from the judge's own\n\n### 5.2 How Argument Decomposition Improves Judgment\n\nEvery one of these requirements is enhanced by argument decomposition:\n\n1. **Multi-step argument tracking**: Scheme classification breaks complex arguments into identified reasoning patterns. Instead of following a 500-word argument, the judge sees: \"This is an argument from expert opinion (Walton scheme 1) supporting an argument from consequences (scheme 62), which in turn supports a normative conclusion.\" Each scheme has known critical questions.\n\n2. **Rhetorical resistance**: The decomposition strips rhetoric from structure. A beautifully written argument from analogy and a clumsily written one receive the same scheme classification and face the same critical questions. Structure evaluation is more resistant to persuasion than holistic evaluation.\n\n3. **Domain knowledge augmentation**: Auto-connect links the current argument to related arguments in the graph. A judge evaluating a climate argument sees connections to previously decomposed economic, ethical, and scientific arguments on related topics.\n\n4. **Fallacy detection**: Many logical fallacies are structural — they violate the critical questions of the scheme they purport to use. An argument from authority where the authority lacks relevant expertise fails CQ1 (\"Is the source a genuine expert in the relevant domain?\"). The CQ system makes fallacy detection mechanical rather than requiring expertise.\n\n5. **Perspective-taking**: The worldview filter shows how the same argument structure is evaluated from different value frameworks. A judge can see that an argument succeeds under utilitarian evaluation but fails under deontological evaluation — and understand specifically where the divergence occurs.\n\n### 5.3 Deliberus as Judge Enhancement Infrastructure — The Concrete Proposal\n\nAn AI safety debate could be conducted \"through\" a Deliberus instance:\n\n1. **Pre-debate**: The debate topic's existing argument graph is loaded, showing what is already known about the relevant claims, evidence, and reasoning structures.\n\n2. **Debate phase**: Each AI debater's arguments are fed through the extraction pipeline in real-time. The judge sees not raw text but decomposed argument structure with scheme classifications and auto-generated critical questions.\n\n3. **Cross-referencing**: The judge can see where each debater's arguments connect to (or contradict) the existing graph. Novel claims are highlighted; claims that depend on contested premises are flagged.\n\n4. **Gap analysis**: Sorry markers show where each debater's reasoning has structural gaps — unanswered critical questions, unsupported premises, missing evidence.\n\n5. **Post-debate**: The debate's arguments are integrated into the persistent graph, enriching it for future debates on related topics.\n\nThis transforms the judge's task from \"evaluate two persuasive texts\" to \"compare two argument structures against a shared knowledge base.\" The latter is a much more tractable cognitive task, especially for complex ethical reasoning.\n\n---\n\n## 6. What Would Need to Be True\n\nFor Deliberus to actually serve as alignment infrastructure, the following conditions must hold:\n\n### 6.1 The Convergence Thesis Must Be At Least Partially Correct\n\nIf human values are fundamentally incommensurable — if deep decomposition reveals irreducible disagreement all the way down — then there is no alignment target to converge on. The graph would be valuable as a map of disagreement but not as alignment infrastructure.\n\n**Current evidence:** Partially supportive. MGE shows convergence on value *relationships*. Morality-as-cooperation shows universal moral *types*. Parfit argues for theoretical convergence across ethical traditions. But none of these proves convergence on all questions — and the most important alignment questions (how to handle existential risk, how to distribute AI benefits, whether to create sentient AI) may be precisely the ones where convergence fails.\n\n**What would strengthen this:** Running the extraction pipeline on genuinely diverse moral traditions (not just Western analytic philosophy and its usual interlocutors) and empirically measuring convergence patterns. Does the graph reveal shared structure across Buddhist ethics, Islamic jurisprudence, Ubuntu philosophy, and Western liberalism? This is a testable question.\n\n### 6.2 The Extraction Pipeline Must Be Robust Against Systematic Bias\n\nIf the LLM-based extraction consistently frames one tradition's reasoning as more rigorous than another's, the graph doesn't reveal genuine convergence — it imposes a particular tradition's structure on all others.\n\n**Testable:** Run the same arguments through multiple LLMs. Measure classification agreement. Conduct adversarial testing with arguments deliberately written to challenge the classification system.\n\n### 6.3 The Graph Must Scale Without Losing Structural Integrity\n\nAt 318 claims, the graph is navigable. At 318,000 claims, it must still be navigable — and the auto-connect system must still find genuine connections rather than noise.\n\n**Current trajectory:** The embedding-based auto-connect (Qwen3-Embedding-4B, cosine threshold 0.80) works well at current scale. Whether it works at 1000x scale is an open empirical question. The embedding tension (documented in the project) is acute: embeddings flatten contested concepts toward the average. At scale, false connections may overwhelm genuine ones.\n\n### 6.4 The Sorry Model Must Incentivize Genuine Contribution\n\nIf sorry markers become bureaucratic checkboxes rather than genuine invitations to improve reasoning, the correction mechanism fails. The platform needs contribution incentives that reward quality (answering critical questions with evidence and reasoning) over quantity (filling gaps with plausible-sounding but shallow responses).\n\n### 6.5 The Platform Must Resist Capture\n\nIf any single entity — a corporation, a government, an ideological movement — gains disproportionate control over the graph's structure, the platform becomes a tool of that entity rather than alignment infrastructure for humanity.\n\n**Structural defense:** The graph's structure is transparent and auditable. Capture requires producing better arguments, not just more arguments. But at civilizational stakes, \"producing better arguments\" is well within the capacity of well-funded adversaries.\n\n### 6.6 AI Systems Must Be Able to Interface With the Graph Meaningfully\n\nFor Deliberus to serve as alignment infrastructure, AI systems must be able to:\n- Query the graph for relevant argument structures\n- Verify their reasoning against the graph's schemes and CQs\n- Identify when their conclusions depend on premises that are contested in the graph\n- Update their reasoning when the graph evolves\n\nThis requires a well-defined API for machine reasoning against argument graphs — a \"Deliberus type-checker\" that AI systems can call. This does not yet exist but is technically feasible given the current architecture (FalkorDB graph + embedding-based search + scheme classification).\n\n---\n\n## 7. Concrete Next Steps / Research Directions\n\n### 7.1 Immediate (Can Be Done Now)\n\n1. **Convergence experiment**: Extract arguments from 5 genuinely diverse traditions (e.g., Buddhist ethics, Islamic jurisprudence, Ubuntu philosophy, libertarian philosophy, Confucian ethics) on the same topic. Measure auto-connect patterns. Do cross-tradition connections cluster around shared premises? This is a direct empirical test of the convergence thesis.\n\n2. **Extraction bias audit**: Run the same arguments through Gemini, GPT, and Claude extraction pipelines. Compare classifications. Where do they disagree? Is disagreement systematic or random?\n\n3. **Judge enhancement experiment**: Take a set of complex ethical arguments. Present them to judges in raw text form and in Deliberus-decomposed form (with scheme classifications and CQs). Measure judgment quality. This directly tests the \"judge infrastructure\" hypothesis.\n\n### 7.2 Near-Term (Requires Design Work)\n\n4. **Reasoning verification API**: Design the interface through which an AI system could check its ethical reasoning against the Deliberus graph. What does a \"verification query\" look like? What does a \"verification result\" contain?\n\n5. **Scheme extension framework**: Research non-Western argumentation traditions and design a framework for extending Walton's 96 schemes to accommodate them. Not as a token gesture but as a structural requirement for genuine universality.\n\n6. **Worldview filter formalization**: The current worldview concept is intuitive but underspecified. How exactly do different worldviews weight different value premises? Is this learnable from user behavior? Can it be made explicit?\n\n### 7.3 Longer-Term (Requires Community + Scale)\n\n7. **Democratic AI constitution pilot**: Partner with an AI lab to run a structured deliberation (through Deliberus) generating a model constitution. Compare the result to Collective Constitutional AI (Polis-based) on coverage, depth, and downstream model behavior.\n\n8. **Alignment benchmark**: Create a benchmark measuring how well AI systems can verify ethical reasoning against a Deliberus graph. Analogous to MATH or GSM8K but for structured ethical reasoning.\n\n9. **Cross-platform interoperability**: If Deliberus-style argument graphs become alignment infrastructure, they need interoperability standards. Design an \"Argument Interchange Format for Alignment\" (extending AIF) that other platforms can implement.\n\n10. **Temporal dynamics study**: Track how the graph evolves over months/years. Does it converge, diverge, or oscillate? Does the convergence thesis hold empirically at scale and over time?\n\n---\n\n## Sources\n\n### Alignment Approaches\n- Yudkowsky, E. (2004). \"Coherent Extrapolated Volition.\" Singularity Institute. — https://intelligence.org/files/CEV.pdf\n- Leike, J. et al. (2018). \"Scalable agent alignment via reward modeling.\" DeepMind. — https://arxiv.org/pdf/1811.07871\n- Christiano, P. (2018). \"Iterated Amplification.\" AI Alignment Forum. — https://www.alignmentforum.org/s/EmDuGeRw749sD3GKd\n- Hadfield-Menell, D. et al. (2016). \"Cooperative Inverse Reinforcement Learning.\" NeurIPS. — https://arxiv.org/abs/1606.03137\n- Newberry, T. & Ord, T. (2021). \"The Parliamentary Approach to Moral Uncertainty.\" FHI Technical Report. — https://www.fhi.ox.ac.uk/wp-content/uploads/2021/06/Parliamentary-Approach-to-Moral-Uncertainty.pdf\n- Stuhlmuller, A. & Jung, J. (2022). \"Supervise Process, not Outcomes.\" AI Alignment Forum. — https://www.alignmentforum.org/posts/pYcFPMBtQveAjcSfH/supervise-process-not-outcomes\n- Lightman, H. et al. (2023). \"Improving mathematical reasoning with process supervision.\" OpenAI. — https://openai.com/index/improving-mathematical-reasoning-with-process-supervision/\n- Sorensen, T. et al. (2024). \"Position: A Roadmap to Pluralistic Alignment.\" ICML 2024. — https://proceedings.mlr.press/v235/sorensen24a.html\n- Pluralistic Alignment Workshop (2026). ICML 2026. — https://pluralistic-alignment.github.io/\n- Lloyd, H. (2025). \"Disagreement, AI alignment, and bargaining.\" PhilArchive. — https://philarchive.org/archive/LLODAA-2\n\n### Democratic AI Governance\n- Anthropic (2023). \"Collective Constitutional AI: Aligning a Language Model with Public Input.\" — https://www.anthropic.com/research/collective-constitutional-ai-aligning-a-language-model-with-public-input\n- Huang, S. et al. (2024). \"Collective Constitutional AI.\" FAccT 2024. — https://facctconference.org/static/papers24/facct24-94.pdf\n- Small, C.T. et al. (2023). \"Opportunities and Risks of LLMs for Scalable Deliberation with Polis.\" — https://arxiv.org/pdf/2306.11932\n- Stanford FSI (2026). \"People Shaping AI: Industry-Wide Forum.\" — https://fsi.stanford.edu/news/people-shaping-ai-groundbreaking-industry-wide-forum-invites-public-input-future-ai-agents\n- Nature Scientific Reports (2026). \"Democratic governance through DAO-based deliberation.\" — https://www.nature.com/articles/s41598-026-40180-8\n\n### Moral Philosophy and Convergence\n- Parfit, D. (2011). *On What Matters*, Vols. 1-3. Oxford University Press. Summary: https://www.goodthoughts.blog/p/parfits-triple-theory\n- Street, S. (2006). \"A Darwinian Dilemma for Realist Theories of Value.\" *Philosophical Studies* 127(1): 109-166. — https://philpapers.org/rec/STRADD\n- Hopster, J. (2019). \"Explaining historical moral convergence.\" *Synthese*. — https://link.springer.com/content/pdf/10.1007/s11098-019-01251-x.pdf\n- Hassan, P. (2019). \"Moral Disagreement and Arational Convergence.\" *The Journal of Ethics* 23: 145-161. — https://link.springer.com/article/10.1007/s10892-019-09284-4\n- Sauer (2019). \"The argument from agreement: How universal values undermine moral realism.\" *Ratio*. — https://onlinelibrary.wiley.com/doi/full/10.1111/rati.12233\n- Walen, A. (2024). \"Separate Peaks: Reasons to Reject Derek Parfit's Views about Theoretical Moral Convergence.\" Academia.edu. — https://www.academia.edu/86251254/\n- Baumann, M. (2021). \"In Search of the Trinity.\" *Ethical Theory and Moral Practice*. — https://link.springer.com/article/10.1007/s10677-021-10161-z\n- \"On the Epistemic Significance of Convergence in Ethical Theory\" (2025). *Ethical Theory and Moral Practice*. — https://link.springer.com/article/10.1007/s10677-025-10524-w\n- Carini, J. (2026). \"Most Philosophers Believe in Objective Morality.\" Substack. — https://joelcarini.substack.com/p/most-philosophers-believe-in-objective\n\n### Empirical Moral Convergence\n- Curry, O.S. et al. (2019). \"Is It Good to Cooperate? Testing the Theory of Morality-as-Cooperation in 60 Societies.\" *Current Anthropology*. — https://static1.squarespace.com/static/5e25e21ecd9ab432cd5c5043/t/5e2ef1439cbdd23a7d7ae248/1580134733449/curry.hraf.2019.pdf\n- Alfano, M., Cheong, M. & Curry, O.S. (2024). \"Moral universals: A machine-reading analysis of 256 societies.\" *Heliyon*. — https://www.sciencedirect.com/science/article/pii/S2405844024019716\n- Eriksson, K. et al. (2021). \"Different Populations Agree on Which Moral Arguments Underlie Which Opinions.\" *Frontiers in Psychology*. — https://www.frontiersin.org/articles/10.3389/fpsyg.2021.648405/pdf\n- Bentahila, L. et al. (2021). \"Universality and Cultural Diversity in Moral Reasoning and Judgment.\" *PMC*. — https://ncbi.nlm.nih.gov/pmc/articles/PMC8710723/\n- Knobe, J. et al. \"Consciousness and Morality.\" *Oxford Handbook of Philosophy of Consciousness*. — https://www.ncbi.nlm.nih.gov/books/NBK563591/\n- Benkler, Y. et al. (2024). \"What Are Human Values, and How Do We Align AI to Them? (MGE)\" — https://arxiv.org/abs/2404.10636\n\n### Formal Ethics and Computational Limits\n- Priya T.V. & Rao, S. (2025). \"Deontic Temporal Logic for Formal Verification of AI Ethics.\" — https://arxiv.org/abs/2501.05765\n- Bowman, W.J. \"The Ethical Compiler: Addressing the Is-Ought Gap in Compilation.\" UBC. — https://williamjbowman.com/resources/wjb2024-ethical-compiler.pdf\n- Stenseke, J. (2024). \"On the computational complexity of ethics: moral tractability for minds and machines.\" *Artificial Intelligence Review*. — https://link.springer.com/article/10.1007/s10462-024-10732-3\n- Herzog, C. (2021). \"On formal ethics versus inclusive moral deliberation.\" *AI and Ethics* 1: 313-329. — https://link.springer.com/content/pdf/10.1007/s43681-021-00045-4.pdf\n\n### Argument Mapping and Judgment Quality\n- van Gelder, T. (2015). \"Using Argument Mapping to Improve Critical Thinking Skills.\" *Palgrave Handbook of Critical Thinking in Higher Education*. — https://www.reasoninglab.com/wp-content/uploads/2013/10/TvG-Using-argument-mapping-to-improve-critical-thinking-skills-2015.pdf\n- Cullen, S. et al. (2018). \"Improving analytical reasoning and argument understanding.\" *npj Science of Learning* 3:21. — https://www.nature.com/articles/s41539-018-0038-5\n- Kruger, A. (2022). \"Using argument mapping to improve clarity and rigour in written intelligence products.\" *Intelligence and National Security* 37(5). — https://www.tandfonline.com/doi/abs/10.1080/02684527.2022.2026584\n- Crudele, F. & Raffaghelli, J.E. (2023). \"Promoting Critical Thinking Through Argument Mapping.\" *JITE* 22. — https://www.jite.org/documents/Vol22/JITE-Rv22p497-525Crudele9486.pdf\n- Rahman, S. et al. (2025). \"AI Debate Aids Assessment of Controversial Claims.\" — https://arxiv.org/abs/2506.02175\n- Sanayei, R. et al. (2025). \"Can LLMs Judge Debates? Evaluating Non-Linear Reasoning via Argumentation Theory Semantics.\" EMNLP Findings. — https://aclanthology.org/2025.findings-emnlp.1159.pdf\n\n### LLM Bias in Extraction\n- Matschke et al. (2025). \"Source framing triggers systematic bias in large language models.\" *Science Advances*. — https://www.science.org/doi/10.1126/sciadv.adz2924\n- Chen, J. et al. (2026). \"Uncovering Political Bias in LLMs using Parliamentary Voting Records.\" — https://arxiv.org/html/2601.08785\n- Yoo, J. & Shin, Y. (2025). \"Fair or Framed? Political Bias in News Articles Generated by LLMs.\" EMNLP. — https://aclanthology.org/2025.emnlp-main.856.pdf\n\n### Value Lock-in Risk\n- Qiu, T.A. et al. (2025). \"The Lock-in Hypothesis: Stagnation by Algorithm.\" OpenReview. — https://openreview.net/pdf/3cd76acea8528520e435223cbe94054f58584e8f.pdf\n- Manifund. \"Operationalizing Value Lock-in From Frontier AI Systems.\" — https://manifund.org/projects/operationalizing-value-lock-in-from-frontier-ai-systems\n- Manifund. \"Moral Progress in AI to Prevent Premature Value Lock-in.\" — https://manifund.org/projects/moral-progress-in-ai-to-prevent-premature-value-lock-in\n\n### Argumentation Theory\n- Macagno, F. (2021). \"Argumentation schemes in AI: A literature review.\" *Argument & Computation*. — https://philarchive.org/archive/MACASI-11\n- Yu, S. & Zenker, F. (2020). \"Schemes, Critical Questions, and Complete Argument Evaluation.\" *Argumentation*. — https://d-nb.info/121098167X/34\n\n### Scalable Oversight\n- Pallavi Sudhir, A. et al. (2025). \"A Benchmark for Scalable Oversight Mechanisms.\" ICLR BiAlign Workshop. — https://arxiv.org/abs/2504.03731\n- Nay, J.J. (2025). \"Aligning AI Agents with Humans through Law as Information.\" Stanford CodeX. — https://law.stanford.edu/wp-content/uploads/2025/10/Aligning-AI-Agents-with-Humans-through-Law-as-Information.pdf\n\n### AI Governance and Democratic Risk\n- Jin, Z. et al. (2026). \"AI Poses Risks to Democratic and Social Systems.\" — https://zhijing-jin.com/d/2026-ai-risk.pdf\n- Chatham House (2026). \"Breaking the deadlock on AI governance.\" — https://www.chathamhouse.org/2026/03/breaking-deadlock-ai-governance\n- One Project (2026). \"How to Make AI Serve the Public.\" — https://oneproject.org/how-to-make-ai-serve-the-public/\n"}