{"path":"research/walton-argument-schemes.md","content":"# Walton's Argumentation Schemes: Deep Research for Deliberus\n\n*Research compiled March 27, 2026. Sources cited throughout.*\n\n---\n\n## 1. The Full Taxonomy\n\nDouglas Walton (1942-2020), together with Christopher Reed and Fabrizio Macagno, catalogued **96 argumentation schemes** in their definitive 2008 book *Argumentation Schemes* (Cambridge University Press). An earlier 1996 work identified 29 core schemes; the 2008 compendium expanded this substantially, and a 2015 classification paper by Walton and Macagno proposed a systematic taxonomy.\n\n### How the 96 Schemes Are Organized\n\nThe Walton-Macagno classification uses several orthogonal dimensions:\n\n**By epistemic vs. practical conclusion:**\n- **Epistemic schemes**: Conclusion is that a proposition is known to be true or false (e.g., argument from expert opinion, argument from sign)\n- **Practical/deliberative schemes**: Conclusion is that an action should or should not be carried out (e.g., practical reasoning, argument from consequences)\n\n**By source of inferential warrant:**\n- **Source-based arguments**: Transfer credibility from a source to a claim (expert opinion, position to know, witness testimony, popular opinion)\n- **Rule-based arguments**: Apply a general rule to a specific case (argument from established rule, verbal classification, precedent)\n- **Causal arguments**: Infer from cause to effect or effect to cause (cause to effect, correlation to cause, sign)\n- **Analogy/similarity arguments**: Transfer properties between similar cases (analogy, example, precedent)\n- **Ad hominem / character-based**: Attack or support based on the arguer's character or commitments\n\n**By the 2008 book's chapter organization (the practical grouping):**\n\n| Category | Example Schemes |\n|----------|----------------|\n| Analogy, Classification, Precedent | Argument from analogy, verbal classification, precedent, example |\n| Knowledge-Related | Position to know, expert opinion, ignorance, evidence to hypothesis |\n| Practical Reasoning | Practical reasoning, consequences (positive/negative), values, waste/sunk costs |\n| Popular Opinion & Commitment | Popular opinion, popular practice, commitment, inconsistent commitment |\n| Character & Ad Hominem | Ethotic ad hominem, circumstantial ad hominem, bias, pragmatic inconsistency |\n| Causal | Cause to effect, effect to cause, correlation to cause, sign |\n| Threat & Fear | Argument from threat, fear appeal, danger appeal |\n| Gradualism & Slippery Slope | Gradualism, causal slippery slope, precedent slippery slope, full slippery slope |\n| Composition & Division | Argument from composition, argument from division |\n| Other | Need for help, distress, alternatives, opposites, abductive reasoning |\n\n### The 29 Core Schemes (1996 Original Set)\n\nThese form the most commonly referenced and annotated subset:\n\n1. Argument from Position to Know\n2. Argument from Expert Opinion\n3. Argument from Analogy\n4. Argument from Verbal Classification\n5. Argument from Established Rule\n6. Argument from Example\n7. Argument from Exception to a Rule\n8. Argument from Precedent\n9. Practical Reasoning\n10. Arguments from Ignorance / Lack of Knowledge\n11. Argument from Positive Consequences\n12. Argument from Negative Consequences\n13. Fear Appeal\n14. Danger Appeal\n15. Arguments from Alternatives\n16. Arguments from Opposites\n17. Plea for Help / Excuse\n18. Argument from Composition\n19. Argument from Division\n20. Slippery Slope Arguments (causal, precedent, full)\n21. Argument from Popular Opinion\n22. Argument from Popular Practice\n23. Argument from Commitment\n24. Arguments from Inconsistency\n25. Ethotic Ad Hominem\n26. Circumstantial Ad Hominem\n27. Argument from Bias\n28. Argument from Cause to Effect / Effect to Cause / Correlation to Cause\n29. Argument from Evidence to a Hypothesis / Abductive Reasoning\n30. Argument from Waste (Sunk Costs)\n\nSources: [Walton, Reed & Macagno (2008)](https://www.cambridge.org/core/books/argumentation-schemes/9AE7E4E6ABDE690565442B2BD516A8B6), [Walton & Macagno (2015)](https://journals.sagepub.com/doi/10.1080/19462166.2015.1123772), [Wikipedia: Argumentation Scheme](https://en.wikipedia.org/wiki/Argumentation_scheme)\n\n---\n\n## 2. Critical Questions for the Most Common Schemes\n\nEach scheme comes with a set of **critical questions** (CQs) -- structured challenges that test whether the argument holds. If a CQ cannot be satisfactorily answered, the argument is undercut. This is the core mechanism that maps to Deliberus's \"VALID badge\" concept.\n\n### Argument from Expert Opinion\n*E is an expert in domain S; E asserts A; therefore A is (presumably) true.*\n\n| # | Critical Question |\n|---|-------------------|\n| CQ1 | **Expertise**: How credible is E as an expert source? |\n| CQ2 | **Field**: Is E an expert in the specific field that A belongs to? |\n| CQ3 | **Opinion**: What exactly did E assert that implies A? |\n| CQ4 | **Trustworthiness**: Is E personally reliable as a source? |\n| CQ5 | **Consistency**: Is A consistent with what other experts assert? |\n| CQ6 | **Evidence**: Is E's assertion based on evidence? |\n\n### Argument from Position to Know\n*Source a is in a position to know about domain S; a asserts A; therefore A is true.*\n\n| # | Critical Question |\n|---|-------------------|\n| CQ1 | Is a in a position to know whether A is true (false)? |\n| CQ2 | Is a an honest, trustworthy, reliable source? |\n| CQ3 | Did a actually assert that A is true (false)? |\n\n### Practical Reasoning (Argument from Goal to Action)\n*I have goal G; action A is a means to realize G; therefore I ought to carry out A.*\n\n| # | Critical Question |\n|---|-------------------|\n| CQ1 | What other goals might conflict with G? |\n| CQ2 | What alternative actions could also bring about G? |\n| CQ3 | Which action (A or alternatives) is most efficient? |\n| CQ4 | Is it practically possible to bring about A? |\n| CQ5 | What consequences of bringing about A should be considered? |\n\n### Argument from Analogy\n*Case C1 is similar to case C2; A is true in C1; therefore A is true in C2.*\n\n| # | Critical Question |\n|---|-------------------|\n| CQ1 | Are there relevant differences between C1 and C2 that undermine the analogy? |\n| CQ2 | Is A actually true in C1? |\n| CQ3 | Is there a stronger counter-analogy (a case similar to C2 where A is false)? |\n\n### Argument from Consequences (Positive)\n*If A is brought about, good consequences will plausibly occur; therefore A should be brought about.*\n\n| # | Critical Question |\n|---|-------------------|\n| CQ1 | How strong is the causal link between A and the consequences? |\n| CQ2 | Are there negative consequences of A that outweigh the positive ones? |\n| CQ3 | Are there alternative actions that produce the same good consequences? |\n\n### Argument from Consequences (Negative)\n*If A is brought about, bad consequences will occur; therefore A should not be brought about.*\n\n| # | Critical Question |\n|---|-------------------|\n| CQ1 | How strong is the evidence that A actually leads to those consequences? |\n| CQ2 | Are there positive consequences of A that outweigh the negative ones? |\n| CQ3 | Can the bad consequences be mitigated or avoided while still doing A? |\n\n### Argument from Popular Opinion\n*A large majority believes A; therefore A is (presumably) true.*\n\n| # | Critical Question |\n|---|-------------------|\n| CQ1 | Is it actually true that a large majority believes A? (Evidence for the claim of popularity?) |\n| CQ2 | Could the popular belief be based on incomplete information or bias? |\n| CQ3 | Even if popular, has A been evaluated on its own merits with evidence? |\n\n### Argument from Sign\n*B is true; B is a sign of A; therefore A is true.*\n\n| # | Critical Question |\n|---|-------------------|\n| CQ1 | Is the correlation between B and A strong enough to be reliable? |\n| CQ2 | Could other factors explain B besides A? |\n\n### Argument from Commitment\n*Agent a has committed to proposition A; therefore a should act consistently with A.*\n\n| # | Critical Question |\n|---|-------------------|\n| CQ1 | Did a actually commit to A? |\n| CQ2 | Has a retracted commitment to A? |\n| CQ3 | Are there grounds justifying retraction of the commitment? |\n\n### Ad Hominem (Circumstantial)\n*Person a advocates A but a's circumstances are inconsistent with A; therefore A is doubtful.*\n\n| # | Critical Question |\n|---|-------------------|\n| CQ1 | Is a's circumstance actually inconsistent with A? |\n| CQ2 | Does a's inconsistency have any bearing on the truth of A? |\n| CQ3 | Is there an explanation that resolves the apparent inconsistency? |\n\n### Slippery Slope Argument\n*A0 seems acceptable; but A0 leads to A1, A1 to A2, ..., An is terrible; therefore A0 should not be done.*\n\n| # | Critical Question |\n|---|-------------------|\n| CQ1 | Is each step in the causal chain plausible? |\n| CQ2 | Can the chain be broken at any point? (Are there firewalls?) |\n| CQ3 | Is the final outcome really as bad as claimed? |\n\n### Argument from Waste (Sunk Costs)\n*If a stops now, all previous efforts are wasted; wasting effort is bad; therefore a should continue.*\n\n| # | Critical Question |\n|---|-------------------|\n| CQ1 | Is the goal still achievable? |\n| CQ2 | Should a reassessment of cost/benefit from this point forward be made, ignoring sunk costs? |\n\n### Argument from Ignorance\n*If A were true, A would be known; A is not known; therefore A is false.*\n\n| # | Critical Question |\n|---|-------------------|\n| CQ1 | Has a thorough enough search been conducted? |\n| CQ2 | Is the burden of proof correctly allocated? |\n| CQ3 | What strength of proof is required? |\n\n### Argument from Cause to Effect\n*Cause C generally produces effect E; C is present; therefore E will occur.*\n\n| # | Critical Question |\n|---|-------------------|\n| CQ1 | How strong is the causal generalization? |\n| CQ2 | Is the evidence for C being present adequate? |\n| CQ3 | Are there other factors that could prevent E despite C? |\n\n### Argument from Evidence to a Hypothesis (Abductive)\n*Hypothesis H explains evidence E better than alternatives; therefore H is (provisionally) true.*\n\n| # | Critical Question |\n|---|-------------------|\n| CQ1 | How satisfactorily does H explain E? |\n| CQ2 | How well do competing hypotheses explain E? |\n| CQ3 | How thorough was the search for alternative hypotheses? |\n\nSources: [Walton et al. (2008)](https://www.cambridge.org/core/books/argumentation-schemes/9AE7E4E6ABDE690565442B2BD516A8B6), [Studies in Critical Thinking](https://ecampusontario.pressbooks.pub/criticalthinking1234/chapter/__unknown__-29/), [ReasoningLab](https://www.reasoninglab.com/patterns-of-argument/argumentation-schemes/waltons-argumentation-schemes/)\n\n---\n\n## 3. LLM Classification of Argument Schemes (2024-2026)\n\n### State of the Art\n\nThis is an **active and rapidly evolving** research area. Key findings:\n\n**\"Can Large Language Models Understand Argument Schemes?\" (Bezou-Vrakatseli, Cocarascu & Modgil, ACL Findings 2025)**\n- The most directly relevant paper. Systematic evaluation of **7 LLMs** for classifying argument schemes based on Walton's taxonomy.\n- Tested zero-shot, few-shot, and chain-of-thought prompting.\n- Two enhancement strategies: formal definitions from Walton and LLM-generated descriptions.\n- Finding: Larger models show \"satisfactory performance\" in identifying schemes, but the task remains challenging -- especially for enthymemes (arguments with implicit premises).\n- Evaluated on both manually annotated and automatically generated arguments.\n- Source: [ACL Anthology](https://aclanthology.org/2025.findings-acl.702/)\n\n**\"A Comprehensive Study of LLM-Based Argument Classification\" (2026)**\n- Evaluated GPT-5.2, Llama 4, and DeepSeek on Args.me and UKP corpora.\n- Best result: **GPT-5.2 at 91.9% (Args.me) and 78.0% (UKP)** for argument component classification (claims, premises, stance).\n- Chain-of-thought prompting added 2-8 percentage points.\n- Systematic failure modes: instability with prompt formulation, difficulty detecting implicit criticism, challenges with complex argument structures.\n- Source: [arXiv:2603.19253](https://arxiv.org/abs/2603.19253)\n\n**\"Large Language Models in Argument Mining: A Survey\" (2025)**\n- Comprehensive survey showing LLMs significantly outperform traditional supervised approaches.\n- GPT-4 matches or exceeds fine-tuned RoBERTa on argument quality assessment.\n- RAG-GPT-4 adds further improvement.\n- **Critical gap identified**: Argumentation scheme classification specifically remains under-studied compared to component detection and quality assessment.\n- Source: [arXiv:2506.16383](https://arxiv.org/abs/2506.16383)\n\n**US2016 Corpus Annotation Study (Visser et al.)**\n- 505 inferential relations in Clinton-Trump debate transcripts annotated with Walton's schemes.\n- **97% of relations** (491/505) could be classified into one of Walton's scheme types.\n- Most common scheme: \"Argument from example\" (81 instances).\n- Inter-annotator agreement: Cohen's kappa = **0.723** (substantial agreement) for Walton scheme annotation.\n- Key challenge: Distinguishing similar schemes (e.g., practical reasoning vs. argument from values).\n- Source: [Springer: Annotating Argument Schemes](https://link.springer.com/article/10.1007/s10503-020-09519-x)\n\n### Practical Assessment for Deliberus\n\n| Task | LLM Capability (2026) | Readiness |\n|------|----------------------|-----------|\n| Detect if text is argumentative | High (>90% F1) | Production-ready |\n| Classify claim vs. premise | High (78-92%) | Production-ready |\n| Identify specific Walton scheme | Moderate (~70-80% for common schemes) | Viable with human review |\n| Handle enthymemes (implicit premises) | Low-Moderate | Needs augmentation |\n| Multi-scheme arguments | Low | Research stage |\n\n---\n\n## 4. Automatic Critical Question Generation\n\n### The CQs-Gen 2025 Shared Task (ACL ArgMining Workshop)\n\nA landmark event: the **first shared task specifically on generating critical questions for arguments**, hosted at the 12th Workshop on Argument Mining, co-located with ACL 2025 in Vienna.\n\n**Task format**: Given an argumentative text (intervention from a real debate), annotated with argumentation schemes, generate exactly 3 useful critical questions.\n\n**Results**: 13 teams participated. Best score: **67.6** (human-evaluated usefulness). Three of the four top-performing teams incorporated argumentation scheme annotations.\n\n**Key systems**:\n\n1. **ELLIS Alicante (1st place, score 67.6)**: Two-stage pipeline with a Questioner (Llama 3.1 8B) generating 8 candidate questions and a Judge (Gemma 2 9B) selecting the best 3. **81% of selected questions** were generated with argumentation scheme information in the prompt. Both were small open-source models running locally. Source: [arXiv](https://arxiv.org/html/2506.14371)\n\n2. **DayDreamer (4th place)**: Three-stage pipeline -- extract structured arguments using Walton's taxonomy, generate critical questions using scheme-specific CQ templates, then rank. GPT-4o-mini outperformed Llama-3.1-8B. Achieved 60 helpful questions, 25 unhelpful, 17 invalid. Source: [arXiv](https://arxiv.org/html/2505.15554)\n\n3. **TriLLaMa (6th place)**: Two-stage generation+classification using LLaMA 3.1 (8B/70B/405B). Found larger models better for generation, medium for classification. Source: [ACL Anthology](https://aclanthology.org/2025.argmining-1.34/)\n\n4. **CUET_SR34**: Few-shot with Llama-3-8B integrating NER and argument schemes. Source: [ACL Anthology](https://aclanthology.org/2025.argmining-1.28/)\n\n### Implications for Deliberus\n\nThe CQs-Gen results demonstrate that **automatic critical question generation is viable today** with relatively small open-source models. The winning system used two 8-9B models, not massive proprietary ones. Argumentation scheme information materially improves question quality (81% of best questions used scheme context).\n\nHowever, the best score of 67.6 out of 100 shows there is still significant room for improvement, particularly around generating questions that are both useful AND non-obvious.\n\n---\n\n## 5. How Schemes Map to Formal Argumentation\n\n### ASPIC+ Integration\n\nASPIC+ is the dominant framework for structured argumentation in AI. The connection to Walton's schemes is theoretically clean:\n\n**Argumentation schemes = defeasible inference rules in ASPIC+.** An argument scheme provides a pattern where the premises create a *presumption* in favor of the conclusion (not a guarantee). This maps directly to ASPIC+'s distinction between:\n- **Strict rules**: Premises guarantee conclusion (deductive logic)\n- **Defeasible rules**: Premises create a presumption (argumentation schemes)\n\n**Critical questions = pointers to counterarguments.** Each CQ identifies a specific way the defeasible inference can be attacked:\n- **Undercutters**: Challenge the inference itself (e.g., \"Is E really an expert in this field?\")\n- **Rebutters**: Provide reasons to believe the conclusion is false\n- **Undermining attacks**: Challenge the truth of a premise\n\n**Example formalization**: Argument from Expert Opinion becomes:\n```\nDefeasible rule: expert(E, S) AND asserts(E, A) AND in_domain(A, S) =>d A\nUndercutter from CQ2: NOT in_domain(A, expertise_of(E))\nUndercutter from CQ5: disagrees(E2, A) AND expert(E2, S)\n```\n\n### QBAF (Quantitative Bipolar Argumentation Framework) Integration\n\nQBAFs extend abstract argumentation with numerical strengths. Scheme information can inform:\n- **Base strength**: Arguments using well-answered scheme CQs get higher base weight\n- **Attack/support relations**: CQ failures map to attack edges; CQ satisfactions map to support edges\n- **Gradual semantics**: The number of answered vs. unanswered CQs provides a natural gradient for argument strength\n\nRecent work (KR 2025) on gradual semantics for Assumption-Based Argumentation shows these approaches can satisfy desirable properties like balance and monotonicity.\n\n**Gap**: While ASPIC+ theoretically supports schemes as defeasible rules, practical computational implementations have not fully realized this. The theoretical framework is well-established; tooling lags behind.\n\nSources: [ASPIC+ Tutorial (Modgil & Prakken, 2014)](https://journals.sagepub.com/doi/10.1080/19462166.2013.869766), [Walton & Reed on Defeasible Inferences](https://cgi.csc.liv.ac.uk/~floriana/CMNA/WaltonReed.pdf), [KR 2025: Gradual Semantics](https://proceedings.kr.org/2025/50/kr2025-0050-rapberger-et-al.pdf)\n\n---\n\n## 6. The \"VALID Badge\" Concept\n\n### Has Anyone Implemented Scheme-Based Validation?\n\n**No platform has implemented a user-facing \"VALID badge\" based on argumentation scheme analysis.** This would be novel to Deliberus.\n\nExisting platforms:\n- **Kialo**: 1M+ users, argument trees with pro/con structure, but **no scheme awareness**. Arguments are just nested claims without formal scheme classification. Source: [Wikipedia: Kialo](https://en.wikipedia.org/wiki/Kialo)\n- **Araucaria/OVA+**: Academic tools that support scheme annotation, but for researchers, not general users. No validation badges.\n- **Carneades**: Walton's own computational system (with Thomas Gordon) supports scheme-based argument construction and evaluation, but is a research tool, not a social platform.\n- **The Argument Web**: An ecosystem of interoperable tools for argumentation analysis (OVA+, Arvina, AIFdb), but oriented toward researchers. Source: [Philosophy & Technology](https://link.springer.com/article/10.1007/s13347-017-0260-8)\n\n### Proposed Deliberus Implementation\n\nThe pipeline would work as follows:\n\n```\nUser writes argument\n    |\n    v\nLLM classifies argumentation scheme (e.g., \"Argument from Expert Opinion\")\n    |\n    v\nSystem surfaces the scheme's critical questions as structured prompts:\n    - \"Is the cited expert credible in this specific domain?\"\n    - \"Do other experts agree?\"\n    - \"Is the assertion based on evidence?\"\n    |\n    v\nCommunity members answer critical questions (with evidence)\n    |\n    v\nUnanswered CQs    ->  Argument marked INCOMPLETE (amber)\nAll CQs answered   ->  Argument marked for review\nAll CQs satisfied  ->  Argument earns VALID badge (green)\nAny CQ failed      ->  Argument marked CHALLENGED (red)\n```\n\n**Key design considerations**:\n- Not binary: An argument can be partially validated (3/6 CQs answered)\n- CQ answers are themselves arguments, recursively subject to the same process\n- Multiple schemes may apply to one argument (multi-scheme classification)\n- The badge should show *which* CQs are answered and which are open\n- VALID does not mean TRUE -- it means \"this argument has survived structured scrutiny under its identified reasoning pattern\"\n\n---\n\n## 7. Walton's Work with AI\n\n### Walton's Own Computational Contributions\n\nDouglas Walton (1942-2020) actively bridged argumentation theory and AI throughout his career. He produced nearly 60 books and hundreds of papers.\n\n**Araucaria** (with Reed, Macagno, Rowe): A free software tool for argument diagramming. Users load text containing an argument and create diagrams by dragging lines between propositions. Supports Walton's schemes natively -- users can tag inferences with specific scheme types. Used in critical thinking courses at universities and in legal evidence analysis. Source: [Springer](https://link.springer.com/chapter/10.1007/978-1-84800-149-7_8)\n\n**Carneades Argumentation System** (with Thomas Gordon): A computational model for argument construction, evaluation, and visualization. Carneades can analyze arguments, evaluate them against schemes, make argument diagrams, and construct arguments from a knowledge base. Walton and Gordon compared Carneades with IBM Watson Debater as leading systems for argument invention. Source: [Springer](https://link.springer.com/article/10.1007/s10503-017-9439-5)\n\n**The Argument Web** (Reed et al.): Walton contributed to the broader Argument Web ecosystem -- an online infrastructure hosting the largest publicly accessible corpora of argumentation and interoperable analysis tools. Source: [PMC](https://pmc.ncbi.nlm.nih.gov/articles/PMC6979725/)\n\n**Dialogue Types and Agent Communication**: Walton's typology of dialogues (persuasion, inquiry, negotiation, information-seeking, deliberation, eristic) was adopted in multi-agent systems for governing how software agents argue with each other. This influenced the design of inter-agent communication protocols. Source: [Springer: In Memoriam](https://link.springer.com/article/10.1007/s10506-020-09272-2)\n\n**AI and Law**: Walton's work on argumentation schemes was particularly influential in legal AI, where defeasible reasoning is central to case-based reasoning, burden of proof, and evidence evaluation. Source: [IOS Press](https://content.iospress.com/articles/argument-and-computation/aac200543)\n\n### Key Collaborators\n\n- **Christopher Reed** (University of Dundee): Co-author of the 2008 book, developer of Araucaria and the Argument Web infrastructure\n- **Fabrizio Macagno** (Universidade Nova de Lisboa): Co-author, focused on linguistic and pragmatic aspects of schemes\n- **Thomas F. Gordon**: Co-developer of Carneades\n- **Henry Prakken** (Utrecht University): ASPIC+ framework that formalized schemes as defeasible rules\n- **Sanjay Modgil** (King's College London): Extended ASPIC+ with preferences, now supervising LLM scheme research\n\n---\n\n## 8. Practical Implications for Deliberus\n\n### The Full User-Facing Pipeline\n\n```\n1. COMPOSE:  User writes a claim/argument in natural language\n\n2. CLASSIFY: LLM identifies the argumentation scheme(s)\n             - Show user: \"This looks like an Argument from Expert Opinion\"\n             - Allow user to confirm, correct, or add additional schemes\n             - Handle multi-scheme arguments (common in practice)\n\n3. SURFACE:  System presents the scheme's critical questions\n             - Prioritize unanswered CQs\n             - Show which CQs have been addressed by others\n             - Allow community to add new CQs beyond the template\n\n4. RESPOND:  Community members answer critical questions\n             - Each answer is itself an argument (recursive scheme analysis)\n             - Evidence attachments (links, citations, data)\n             - Voting/rating on CQ answers\n\n5. ASSESS:   Argument gains/loses status based on CQ coverage\n             - INCOMPLETE: CQs unanswered (amber)\n             - VALID: All CQs answered satisfactorily (green)\n             - CHALLENGED: CQ answer undermines the argument (red)\n             - CONTESTED: Conflicting CQ answers, unresolved (yellow)\n\n6. EVOLVE:   Status updates as new evidence/answers arrive\n             - Temporal tracking: argument was VALID until date X\n             - Notification when status changes\n```\n\n### Design Tensions to Resolve\n\n**Friction vs. rigor**: Surfacing 6 critical questions for every argument is rigorous but creates friction. Possible mitigations:\n- Progressive disclosure: Show the most important 1-2 CQs first\n- AI pre-filling: LLM drafts initial CQ answers that the community validates\n- Gamification: Answering CQs earns reputation/badges\n\n**Scheme ambiguity**: Real arguments often fit multiple schemes or none perfectly. The US2016 annotation study found 97% of arguments classifiable, but annotators struggled with similar schemes (practical reasoning vs. argument from values). For a user-facing platform, this means:\n- Allow multiple scheme tags\n- Show scheme suggestions as probabilistic (top 3 with confidence)\n- Don't require scheme classification for participation\n\n**Enthymemes**: Most natural arguments have implicit premises. \"We should fund this project because Dr. Smith says it's promising\" has the implicit premise \"Dr. Smith is an expert.\" LLMs struggle with enthymeme reconstruction. Consider:\n- LLM reconstructs implicit premises and shows them to the user for confirmation\n- This makes the full argument visible and scheme-classifiable\n\n**Recursive depth**: If CQ answers are themselves arguments needing CQ analysis, you get infinite recursion. Practical solutions:\n- Limit recursion depth (e.g., 2-3 levels)\n- Lower-level CQ answers use simpler validation (votes, evidence links)\n- AI-assisted assessment at deeper levels\n\n### Integration with Deliberus's Existing Concepts\n\n| Deliberus Concept | Walton Scheme Connection |\n|-------------------|--------------------------|\n| Contention diagram | Schemes structure the edges (support/attack types become scheme-specific) |\n| Argument bundles | A bundle is an argument + its scheme + its CQ responses |\n| Scoring axes | CQ satisfaction provides one scoring axis; scheme strength is another |\n| VALID badge | Fully answered CQs = VALID; this is the concrete implementation |\n| Feed algorithm | Prioritize arguments with unanswered CQs (\"this argument needs scrutiny\") |\n\n### Technical Architecture Considerations\n\n**Scheme classification model**: Start with GPT-4/Claude API calls for classification, validated against the US2016 annotated corpus. Fine-tune a smaller model (Llama 8B) as volume grows. The ELLIS Alicante CQs-Gen winner used 8-9B parameter models -- this is achievable on Darwin's GTX 1650 with quantization.\n\n**AIF (Argument Interchange Format)**: The standard data format for computational argumentation. Schemes map naturally to AIF's scheme nodes. Using AIF from the start enables interoperability with the Argument Web ecosystem (AIFdb, OVA+, Arvina). This aligns with the academic foundation analysis recommending AIF.\n\n**Graph database fit**: Arguments, schemes, CQs, and CQ responses form a natural graph. FalkorDB/Graphiti can represent:\n- Argument nodes with scheme labels\n- CQ edges connecting arguments to their challenges\n- CQ-response edges with satisfaction status\n- Temporal edges tracking when validity changed\n\n---\n\n## 9. Limitations of Argumentation Schemes\n\n### Coverage: What Doesn't Fit?\n\nThe US2016 annotation study found **97% of arguments in political debate** could be classified into Walton's schemes (491/505 relations). The 3% remainder were tagged \"default inference.\" This is encouraging but comes with caveats:\n\n- **Political debate is adversarial and structured** -- it may over-represent common schemes. Casual conversation, scientific discourse, or artistic critique may have different distributions.\n- **The schemes list is not closed**: Reviewers have noted that schemes could multiply indefinitely. Leo Groarke (Notre Dame Philosophical Reviews) proposed an \"argument from fallibility\" not in Walton's compendium and argued that \"the list of schemes is large but not comprehensive.\" Source: [NDPR Review](https://ndpr.nd.edu/reviews/argumentation-schemes/)\n- **Distinguishing similar schemes is hard**: Even trained annotators struggle with practical reasoning vs. argument from values, or argument from sign vs. argument from evidence to hypothesis (kappa 0.723 is \"substantial\" but not \"perfect\").\n- **Complex multi-step arguments** may use multiple schemes chained together -- no single scheme captures the whole.\n\n### Cultural Limitations\n\n**Walton's schemes emerge from Western philosophical tradition**, specifically from Aristotle's *Topics* through medieval disputation to modern informal logic. Research on cross-cultural argumentation reveals significant gaps:\n\n- **Different cultures privilege different forms of reasoning**: Japanese argumentation favors indirectness and consensus-building; Chinese philosophical tradition has rich argumentation but emphasizes different patterns (harmony-seeking vs. adversarial). Source: [Brill: Cross-Cultural Differences](https://brill.com/view/journals/jocc/18/3-4/article-p358_7.xml)\n\n- **Argument quality norms vary**: What counts as a \"good argument\" differs across cultures. In some traditions, appeal to authority carries more weight; in others, experiential knowledge or communal wisdom take precedence. Source: [Springer: Argument Quality and Cultural Difference](https://link.springer.com/article/10.1023/A:1026466310894)\n\n- **However, recent research challenges oversimplification**: Hugo Mercier's work shows that humans across cultures are surprisingly good at argumentation when motivated, and that \"the ancient Chinese were in fact skilled arguers.\" The principle of non-contradiction appears universal; differences are more about *norms of discourse* than *logical capacity*. Source: [Cognition and Culture Institute](https://cognitionandculture.net/blogs/hugo-mercier/cross-cultural-differences-in-argumentation/)\n\n- **Practical implication for Deliberus**: Don't hard-code Western schemes as the only valid reasoning patterns. Allow community-defined schemes. Track which schemes are most used by different communities. Consider supplementing Walton's taxonomy with non-Western reasoning patterns (e.g., reasoning from harmony, reasoning from elder wisdom, reasoning from lived experience).\n\n### Other Limitations\n\n- **Emotional and narrative arguments**: Walton's schemes focus on logical/inferential patterns. Arguments that persuade through narrative, metaphor, or emotional resonance don't fit neatly (though \"fear appeal\" and \"danger appeal\" begin to address this).\n- **Value disagreements**: Schemes assume shared evaluative standards. When people disagree about fundamental values (e.g., liberty vs. equality), scheme analysis can identify the structure but cannot resolve the disagreement.\n- **Scheme proliferation**: The jump from 29 to 96 schemes raises the question: where does systematization end? Too many schemes reduces their utility as cognitive tools.\n- **Static taxonomy**: The scheme list is fixed. New forms of reasoning (e.g., \"argument from machine learning model output\") are not covered. A living taxonomy would better serve a platform like Deliberus.\n\n---\n\n## Key Takeaways for Deliberus\n\n1. **Walton's schemes are the most comprehensive and well-studied framework** for classifying everyday arguments. 96 schemes with structured critical questions provide a rich foundation.\n\n2. **LLMs can identify schemes with moderate accuracy today** (~70-80% for common schemes) and this is improving rapidly. Production use is viable with human-in-the-loop validation.\n\n3. **Automatic critical question generation is a solved-enough problem**: The CQs-Gen 2025 shared task shows small open-source models can generate useful critical questions, especially when given scheme information. The winning system (ELLIS Alicante) used two 8-9B models.\n\n4. **The \"VALID badge\" concept is novel** -- no existing platform implements scheme-based argument validation as a user-facing feature. This is a genuine differentiator for Deliberus.\n\n5. **ASPIC+ provides the formal backbone**: Schemes as defeasible rules, CQs as attack pointers, QBAFs for numerical strength. The theory is mature; tooling needs building.\n\n6. **Cultural limitations are real but manageable**: Start with Walton's Western-origin taxonomy (which covers 97% of political debate), but design the system to accommodate community-defined schemes and non-Western reasoning patterns.\n\n7. **Progressive disclosure is essential**: Don't show all 6 CQs upfront. Start with \"What's the evidence?\" (the most universal CQ) and reveal more as the argument develops.\n\n---\n\n## References\n\n### Primary Sources\n- Walton, D., Reed, C., & Macagno, F. (2008). *Argumentation Schemes*. Cambridge University Press. [Link](https://www.cambridge.org/core/books/argumentation-schemes/9AE7E4E6ABDE690565442B2BD516A8B6)\n- Walton, D. & Macagno, F. (2015). A classification system for argumentation schemes. *Argument & Computation*, 6(3). [Link](https://journals.sagepub.com/doi/10.1080/19462166.2015.1123772)\n- Walton, D. (1996). *Argumentation Schemes for Presumptive Reasoning*. Routledge. [Link](https://www.routledge.com/Argumentation-Schemes-for-Presumptive-Reasoning/Walton/p/book/9780805820720)\n\n### LLM and Argument Mining\n- Bezou-Vrakatseli, E., Cocarascu, O., & Modgil, S. (2025). Can Large Language Models Understand Argument Schemes? *Findings of ACL 2025*. [Link](https://aclanthology.org/2025.findings-acl.702/)\n- Comprehensive Study of LLM-Based Argument Classification (2026). [arXiv:2603.19253](https://arxiv.org/abs/2603.19253)\n- Large Language Models in Argument Mining: A Survey (2025). [arXiv:2506.16383](https://arxiv.org/abs/2506.16383)\n\n### CQs-Gen 2025 Shared Task\n- ELLIS Alicante (1st place). [arXiv](https://arxiv.org/html/2506.14371)\n- DayDreamer (4th place). [arXiv](https://arxiv.org/html/2505.15554)\n- TriLLaMa (6th place). [ACL Anthology](https://aclanthology.org/2025.argmining-1.34/)\n- Task Overview. [ACL Anthology](https://aclanthology.org/2025.argmining-1.23/)\n- Task Website. [Link](https://hitz-zentroa.github.io/shared-task-critical-questions-generation/)\n\n### Annotation and Corpora\n- Visser, J. et al. (2020). Annotating Argument Schemes. *Argumentation*. [Link](https://link.springer.com/article/10.1007/s10503-020-09519-x)\n- US2016 Corpus. [Link](https://link.springer.com/article/10.1007/s10579-019-09446-8)\n\n### Formal Frameworks\n- Modgil, S. & Prakken, H. (2014). The ASPIC+ framework for structured argumentation: a tutorial. *Argument & Computation*. [Link](https://journals.sagepub.com/doi/10.1080/19462166.2013.869766)\n- Walton, D. & Reed, C. Argumentation Schemes and Defeasible Inferences. [PDF](https://cgi.csc.liv.ac.uk/~floriana/CMNA/WaltonReed.pdf)\n\n### Computational Tools\n- Araucaria Project. [Link](https://link.springer.com/chapter/10.1007/978-1-84800-149-7_8)\n- The Argument Web. [Link](https://link.springer.com/article/10.1007/s13347-017-0260-8)\n- Walton & Gordon on Carneades and argument invention. [Link](https://link.springer.com/article/10.1007/s10503-017-9439-5)\n\n### Cultural Perspectives\n- Mercier, H. Cross-cultural differences in argumentation. [Link](https://cognitionandculture.net/blogs/hugo-mercier/cross-cultural-differences-in-argumentation/)\n- Nisbett, R. et al. Cross-Cultural Differences in Informal Argumentation. *Journal of Cognition and Culture*. [Link](https://brill.com/view/journals/jocc/18/3-4/article-p358_7.xml)\n\n### Walton's Legacy\n- Baroni, P. & Verheij, B. (2020). Douglas Walton (1942-2020). *Argument & Computation*. [Link](https://journals.sagepub.com/doi/10.3233/AAC-190482)\n- In Memoriam: Influence on AI and Law. [Link](https://link.springer.com/article/10.1007/s10506-020-09272-2)\n"}