{"path":"research/scheme-set-exhaustiveness.md","content":"# Is Walton's Scheme Set Exhaustive? The Count Question, the Periodic Table, and the Generative Basis\n\n**Status**: Literature findings (August 2026 web research pass). Nothing here is decided — the founder decides what, if anything, changes in the ontology. The question investigated: is Walton's argument-scheme set exhaustive, and what is it converging toward? Two candidate readings were on the table going in: (a) the **periodic-table reading** — an open empirical taxonomy with predictable gaps; (b) the **generative reading** — schemes decompose into a small closed basis of primitive inference moves plus the critical-question mechanism, the way 64 codons are generated from 4 bases. The literature turns out to contain a direct, named instance of reading (b) — Wagemans' Periodic Table of Arguments — and, more surprisingly, a partial endorsement of it from inside Walton's own school. Honesty notes on what was *not* verified are inline and in §6.\n\n## §0 Plain-language summary\n\nWalton's scheme set is not exhaustive, and Walton's own school never seriously claimed it was. The same 2008 book is described as containing 60, 65, 96, or 104 schemes depending on who is counting, and Walton's earlier lists had 25 (1996) and 14 (the 2006 textbook list). When Walton himself field-tested a scheme list against 256 real election arguments, over a third didn't fit; he responded by minting two brand-new schemes and institutionalizing a \"none of the above\" category. The set is a species catalogue built by induction from examples — a field guide, not a periodic table, because nothing in it predicts where the gaps are.\n\nThe generative reading is not a fantasy, though. Three separate research programs have tried to build exactly that: Wagemans' Periodic Table of Arguments literally derives a closed factorial space (2 × 2 × 9 = 36 basic types in its 2016 version) from first principles; Katzav & Reed reduce scheme-completeness to enumerating \"relations of conveyance\" (an ontology problem); and Macagno & Walton themselves analyze every scheme as a *combination* of a semantic relation, a type of reasoning, and a logical rule — one causal relation yields five different argument moves depending on which inference rule rides on it. The catch, consistent across all three: the closed part of each system is real but coarse, and the scheme-level richness that people actually recognize and name re-enters through an open inventory (Wagemans' \"levers\", Katzav & Reed's \"Etc.\"-riddled relation tree, Walton's semantic relations). Empirically, on the same corpus, the a priori exhaustive system classified *less* (83%) than the empirical open list (97%).\n\nFor Deliberus the practical upshot is threefold: the scheme enum is a curated open set and needs a first-class does-not-fit value (the field's own annotation practice — the \"default inference\" label — is precisely this confession channel, in production in AIFdb); a two-layer design (small closed form layer + open curated scheme layer) is buildable and is roughly what both Wagemans' and Walton's mature systems already are; and critical-question generation does not die when no scheme fits, because a provably general meta-level CQ list exists for the bare defeasible schema (Yu & Zenker 2020).\n\n## §1 The count question: evidence that Walton's set is open\n\n**The counts do not agree — not even about the same book.** Walton, Reed and Macagno's *Argumentation Schemes* (Cambridge UP, 2008) is described by its own publisher as \"a compendium of 96 schemes\" ([Cambridge](https://www.cambridge.org/core/books/argumentation-schemes/9AE7E4E6ABDE690565442B2BD516A8B6)); the annotation literature that operationalizes it consistently counts \"60 main scheme types\" in Chapter 9, \"disregarding the many listed variants\" (Visser et al. 2020); Lumer counts \"60 main argument schemes and a further 44 sub-schemes\", i.e. 104 total ([Lumer, *Walton's Argumentation Schemes*](https://usiena-air.unisi.it/retrieve/e0feeaa8-a949-44d2-e053-6605fe0a8db0/A111.2_Lumer_Walton%e2%80%99sArgumentationSchemes_Print.pdf)); and the Notre Dame Philosophical Reviews review describes \"a large compendium of sixty-five schemes\" ([NDPR](https://ndpr.nd.edu/reviews/argumentation-schemes/)). Four counts for one book. Across Walton's own works the trajectory is: 25 schemes in *Argumentation Schemes for Presumptive Reasoning* (1996), 14 in the *Fundamentals of Critical Argumentation* textbook list (2006), 60+variants in 2008 (counts per Lumer and per Visser et al.). A set whose cardinality depends on the counter's individuation choices is not behaving like a closed inventory; it is behaving like a curated collection with fuzzy species boundaries.\n\n**Walton's method guarantees openness.** The Walton school's own account of how a scheme earns its place (Macagno, Walton & Reed 2017, [*Argumentation Schemes: History, Classifications, and Computational Applications*](http://arg.tech/people/chris/publications/2017/flap2017.pdf)) is a four-step *inductive* justification: outline a candidate structure from the literature, analyze a significant mass of real examples with it, show the form matters in natural discourse, conclude it \"needs to be recognized as a basic scheme.\" Schemes are inducted from corpora. By construction, a new register can mint a new scheme — and it did:\n\n**The field test: Hansen & Walton 2013.** Collecting ~256 arguments from newspaper coverage of the 2011 Ontario provincial election and classifying them against the 14-scheme textbook list plus \"none of the above\": **over one-third (37.1%) of the arguments fit none of the 14 schemes** ([Walton & Hansen, *Arguments from Fairness and Misplaced Priorities*](https://ccsenet.org/journal/index.php/jpl/article/download/29977/17760); [Hansen & Walton 2013](https://www.jbe-platform.com/content/journals/10.1075/jaic.2.2.03han)). Supplementing with nine further known schemes (practical reasoning etc.) recovered most of these; the residue clustered into two argument kinds *not previously formalized anywhere* — **Appeal to Fairness** and **Argument from Misplaced Priorities** — which the authors then wrote up as new schemes. After all of that, 12 arguments (4.7%) remained unclassifiable. Their own summary: \"familiar lists of schemes … are not comprehensive enough.\" Note also the inverse signal: two schemes on the list (slippery slope, appeal to ignorance) had *zero* instances in that register. Scheme frequency and scheme existence are both register-dependent.\n\n**Critics say it more bluntly.** Lumer's assessment of the schemes approach as a whole: \"the lists of resulting schemes are long, often very long, never complete and always arbitrary in what they include and exclude\", and — of the 2008 compendium specifically — \"these lists are far from being complete — though Walton et al. (2008) for their compendium affirm something near to the opposite. … The compendium of Walton et al. (2008) probably does not even contain 1% of the total of argumentation schemes which could be generated in the same style\" (Lumer, op. cit.; the \"1%\" is Lumer's polemical estimate, not a measurement — he means the compendium omits most deductive forms, probabilistic arguments, interpretative arguments, arguments for/from definitions, historiographic arguments, scientific/statistical argumentation, and complex/molecular arguments generally). The NDPR reviewer makes the structural point: \"we can, in principle, multiply schemes indefinitely, turning any generalizable form of argument into a corresponding scheme\" ([NDPR](https://ndpr.nd.edu/reviews/argumentation-schemes/)). And Shecaira adds a diagnosis of *why* scheme counts never stabilize: rival schemes for the \"same\" argument type often serve different purposes (normative vs descriptive), so the individuation of schemes is purpose-relative all the way down ([Shecaira 2016, *How to Disagree About Argument Schemes*](https://informallogic.ca/index.php/informal_logic/article/view/4610/4007)).\n\n**Walton's own classification attempts stayed provisional.** Chapter 10 of the 2008 book (\"Refining the Classification of Schemes\") proposed three top categories — reasoning, source-based arguments, applying rules to cases. Walton & Macagno 2015 ([*A classification system for argumentation schemes*](https://journals.sagepub.com/doi/10.1080/19462166.2015.1123772)) revised this to four (discovery arguments, practical reasoning, source-based, rules-to-cases), calling the earlier 'reasoning' category defective because it \"did not offer any further criteria for a positive … classification\" — and describing the whole enterprise in the language of a young science: \"The literature on classification of argumentation schemes is still very new, and so it seems hard to know the best way to proceed.\" Their method combines a top-down conceptual pass with a bottom-up pass from hard cases, \"resulting in a classification system in which the higher levels have been developed a priori and the lower levels by empirical generalisations.\" That sentence is, in effect, the two-layer design of §5 stated by Walton himself. Not verified directly: the exact wording of the near-completeness claim Lumer attributes to the 2008 compendium (footnote 8 of his paper cites it; the primary text was not consulted).\n\n**Verdict on reading (a).** The set is open and empirically driven, yes — but the \"periodic table\" half of reading (a) is too generous. Mendeleev's table predicted gallium because an underlying generative variable (atomic number, eventually) ordered the space and made gaps *addressable*. Nothing in Walton's taxonomy predicts a missing scheme; Hansen & Walton discovered Appeal to Fairness by tripping over it in a corpus, not by noticing an empty cell. The Visser group's own analogy for their annotation procedure is the right one: it \"bears a striking resemblance to biological taxonomy, the identification of organisms\" ([Visser et al. 2020](https://link.springer.com/article/10.1007/s10503-020-09519-x)) — a field guide with an identification key, not a periodic table.\n\n## §2 Wagemans' Periodic Table of Arguments: the generative candidate, examined\n\nWagemans' PTA is the literature's explicit attempt at reading (b), and its founding paper opens with exactly the complaint that motivates it: \"The existing classifications of arguments are unsatisfying in a number of ways\", diagnosed as the \"absence or inconsistent application of an ordering principle\" ([Wagemans 2016, *Constructing a Periodic Table of Arguments*](https://pure.uva.nl/ws/files/26929021/Constructing_a_Periodic_Table_of_Arguments.pdf)).\n\n**What it claims.** Every argument is reduced to exactly two statements (one premise, one conclusion), each parsed as a categorical proposition with a subject and a predicate. Three factorial parameters then define the type:\n\n1. **Subject vs predicate argument.** Of the four possible ways premise and conclusion can share components, two are degenerate — no shared element (\"S is P because T is Q\") gives no justificatory force, and full identity (\"S is P because S is P\") is pragmatically inconsistent (begging the question). Exactly two viable forms remain: predicate arguments (*a is X because a is Y*) and subject arguments (*a is X because b is X*). This derivation-by-elimination is the most codon-like move in the whole literature: the space of forms is closed *by argument*, not by enumeration.\n2. **First-order vs second-order.** Second-order arguments (e.g. argument from authority) take the *acceptability of the conclusion as a whole* as their target: \"q [is acceptable] because q is Z\".\n3. **Statement-type combination.** Premise and conclusion are each fact (F), value (V), or policy (P) — nine combinations (FF, VF, PF, …).\n\nThe 2016 framework thus \"allows for 2 × 2 × 9 = 36 basic types of argument\", and it is explicitly billed as \"a factorial typology of argument … each basic type … described as a unique combination of independent features\" ([KRINO project description of the PTA](https://krino-project.com/theory/the-pta/)). Classic schemes get coordinates: argument from sign = first-order predicate FF; \"argument from criterion\" = first-order predicate VF; pragmatic argumentation = first-order predicate PF; argument from authority = second-order predicate. The current version (PTA 3.0, [periodic-table-of-arguments.org](https://periodic-table-of-arguments.org/)) reorganizes the same material as four *forms* named alpha (*a is X because a is Y*), beta (*a is X because b is X*), gamma (*a is X because b is Y* — comparing relations, e.g. a fortiori/opposition arguments, readmitted as a form rather than rejected as degenerate), and delta (*q is A because q is Z*), each crossed with a form-specific *substance* classification and a third parameter, the **lever** ([argument form](https://periodic-table-of-arguments.org/argument-form/); [basic terminology](https://periodic-table-of-arguments.org/basic-terminology/)).\n\n**Does it actually generate Walton's schemes?** Partially, and the partiality is instructive.\n\n- *What the closed part buys*: a PTA coordinate like \"1 pre FF\" genuinely covers argument from sign, argument from cause, argument from effect, argument from correlation, and argument from motive — but it covers them *jointly, without distinguishing them*. The cell is far coarser than a Walton scheme.\n- *Where the scheme-level richness comes back*: the **lever** — \"the underlying mechanism … why the premise should make the conclusion more acceptable.\" And the lever inventory is openly non-generated: for alpha-FF, \"our everyday knowledge suggests typical options: Y is a SIGN of X; Y is a CAUSE of X; Y is an EFFECT of X; Y is CORRELATED with X; Y is a MOTIVE for X\" ([argument lever](https://periodic-table-of-arguments.org/argument-lever/)). \"Everyday knowledge suggests typical options\" is the epistemology of Walton's compendium, not of a periodic law. Wagemans is candid that the lever \"isn't entirely open-ended\" because form and substance constrain it — but constrained-open is still open. In codon terms: form × substance are the four bases, and they really are closed; the levers are not the 64 generated codons but a curated species list *inside* each cell.\n- The dual annotation of US2016 (§4) allows an empirical mapping between the two systems — the published co-occurrence matrix maps PTA technical names (\"1 pre FF\") onto Walton colloquial names (\"argument from sign\") many-to-many, confirming that neither system's types are refinements of the other's ([Visser et al. 2018, ISSA](https://arg.tech/people/chris/publications/2018/issa2018.pdf)).\n\n**Criticisms on record.** Yun Xie's OSSA commentary ([Xie 2016](https://scholar.uwindsor.ca/ossaarchive/OSSA11/papersandcommentaries/21)) raises four: (1) the two-proposition assumption — real arguments are often linked, with multiple co-dependent premises (\"All men are mortal; Socrates is a man; so …\"), which the PTA must absorb by recursive reconstruction (Wagemans' reply reconstructs the major premise as support for the *justificatory force* of the minor — workable, but the \"exactly one premise\" atom is preserved by reconstruction fiat); (2) under a logical (quantifier-aware) reading of categorical propositions, forms like \"Some S are P, because all P are S\" fit neither subject- nor predicate-argument — Wagemans concedes two such cases \"surely pose a challenge: I will try to find some concrete examples and think about it\" ([Reply to commentary](https://pure.uva.nl/ws/files/26929006/Reply_to_commentary_on_Constructing_a_Periodic_Table_of_Arguments.pdf)); (3) arguments with no shared element (\"Abortion should be prohibited, because taking the life of an innocent person is wrong\") require a substitution premise the arguer never stated; (4) some traditional schemes appear to instantiate *more than one* proposition-combination (authority as VF and PF; pragmatic argumentation), threatening the \"unique place in the table\" requirement — Wagemans' reply resolves each case by further reconstruction, and admits the uniqueness question \"definitely need[s] more consideration.\"\n- The empirical criticism is the sharpest (§4 for numbers): on the same 505 inferences, PTA annotation left **17% as Default Inference against 3% for Walton's list** — the a priori exhaustive system *failed to classify more real arguments than the open empirical list*, because a classification failure in any one sub-task (a proposition too vague to be F/V/P; unclear order; propositions too incomplete to locate subject vs predicate) defaults the whole argument out ([Visser et al. 2020](https://pmc.ncbi.nlm.nih.gov/articles/PMC7888437/)). Exhaustiveness-in-principle bought classification-failure-in-practice: natural language does not arrive in categorical subject-predicate form, so every PTA classification includes an interpretive reconstruction step, and the reconstruction is where the arguments leak out.\n\n**Verdict on reading (b) via Wagemans.** The PTA proves the strong lesson and the cautionary one at once. Strong: a small closed basis for argument *form* is derivable, teachable, and factorial — 2×2×9 is real, the elimination argument for \"exactly two viable forms\" is genuinely generative, and the sub-task agreement numbers (§4) show the basis dimensions are more reliably codable than whole schemes. Cautionary: the closed basis underdetermines the scheme layer people actually recognize; recovering named schemes requires the lever, which is once again a curated open list; and the reconstruction burden means the a priori system covers *less* of a real corpus than the empirical catalogue it was meant to systematize.\n\n## §3 Other generative bases\n\n**Macagno & Walton's own decomposition — the in-house version of reading (b).** The 2017 history/classification paper analyzes schemes as \"stereotypical patterns of inference, *combining* semantic-ontological relations with types of reasoning and logical axioms\" and shows the factorization concretely: the single causal relation *fever causes fast breathing* yields five distinct argument moves depending on the logical rule applied — defeasible modus ponens, defeasible modus tollens, two abductive directions, and inductive generalization ([Macagno, Walton & Reed 2017](http://arg.tech/people/chris/publications/2017/flap2017.pdf)). Their conclusion: \"Schemes represent only the prototypical matching between semantic relations and logical rules … The material and the logical relations can combine in several different ways.\" This is Walton's school saying, in its own words, that the compendium's 60 entries are *frozen combinations* from a factor space — the generative reading, endorsed from inside. The open residue here is the inventory of semantic/material relations (causal, definitional, mereological, authority-based, …), which is an ontology question, not an argumentation question.\n\n**Katzav & Reed 2004 — completeness reduced to ontology.** Their \"natural classification\" identifies an argument's type with the **relation of conveyance** it represents (causation, class membership, genus-species, constitution, …): \"developing a natural classification of arguments requires an enumeration of possible relations of conveyance\", pursued by \"uncovering the presuppositions that various domains of natural discourse make about the types of entity there are\" ([Katzav & Reed 2004](http://arg.tech/people/chris/publications/2004/arg2004.pdf)). Their published relation tree (internal: specification, constitution, analyticity, identity; external: causal, non-causal) is explicit about its own openness — nearly every branch terminates in \"Etc.\" This is the clearest statement of *why* scheme-exhaustiveness is hard: the scheme set is exactly as complete as your ontology of relations in the world, so a closed scheme basis would require a closed metaphysics. (For a graph platform this cuts close to home: it is the same reason a closed claim-type enum breaks per register.)\n\n**Kienpointner 1992 — the semantic-relation typology.** *Alltagslogik* aims explicitly at \"a comprehensive typology of schemes of everyday argumentation\": ~60 schemes (Lumer counts 58 main + 15 sub) organized by warrant status into three classes — **rule-using** schemes (subdivided by semantic relation: classification schemes incl. definition, genus-species, part-whole; comparison; opposition; causal), **rule-establishing** schemes (inductive example argumentation), and schemes that are **neither** (illustrative example, analogy, authority) — on a modified Toulmin prototype, empirically grounded in a 300-passage corpus ([frommann-holzboog abstract](https://www.frommann-holzboog.de/reihen/71/710012620?lang=en-gb); [Kienpointner's table of contents](https://www.sgipt.org/wisms/sprache/BegrAna/Plausib/KienpointnerAlltagslogik_IV.pdf)). Katzav & Reed read him as \"attempt[ing] to compile an exhaustive list\" and locate the failure in the warrant notion's context-dependence. Two details matter here: Kienpointner *rejects* field-dependent classification precisely because \"the number of argument classifications would 'explode'\" — an early recognition that one axis of the space is unbounded; and Lumer observes that Kienpointner buys (quasi-)deductive validity for 47 of 58 schemes by strengthening premises until they are \"in many or most cases false\" — systematization pressure deforming the schemes, a warning for any tidy basis.\n\n**Perelman & Olbrechts-Tyteca 1958 — the loci.** The *New Rhetoric*'s techniques (quasi-logical arguments, arguments based on the structure of reality, arguments establishing the structure of reality, dissociation) are the modern scheme literature's acknowledged ancestor and are similar in scale (Lumer: \"similar dimensions\" to Walton's list) and similarly open. Not consulted directly for this pass; cited here as lineage.\n\n**The Argumentum Model of Topics (Rigotti & Greco) — the \"finite hierarchy\" claim.** The AMT advertises exactly what annotation projects want: \"a hierarchical and finite taxonomy of argument schemes as well as systematic, linguistically-informed criteria\", with maxims generated from loci (ontological relations: definitional, mereological, causal, opposition, analogy, authority, practical evaluation) ([Musi, Ghosh & Muresan 2016](https://aclanthology.org/W16-2810.pdf)). It is the strongest *claim* of a closed generative hierarchy in the field. Its empirical record is in §4 — reliability with real annotators was the binding constraint, not coverage.\n\n**Lumer's epistemological alternative.** Lumer's positive proposal derives scheme classes from epistemological principles: three branches — deductive, probabilistic, practical — each grounded in its epistemology (logic, probability theory, decision theory), with elementary vs \"molecular\" schemes (his \"comet\" model) ([Lumer, philarchive](https://philarchive.org/archive/LUMASE)). This is a genuine generative basis of a third kind (neither linguistic form nor semantic relation, but *justification source*), and it makes the honest concession the others make too: his own listed system \"is far from being complete.\"\n\n**Toulmin warrants.** The oldest generative story: every argument type corresponds to a warrant type, warrants are field-dependent, and fields are open-ended — so the generated set is unbounded by design. Kienpointner's typology and Katzav & Reed's critique both descend from wrestling with this (warrant-as-classifier is \"no more informative than claiming [arguments] should be classified in accordance with ways of arguing\" — Katzav & Reed).\n\n**The critical-question mechanism has a completeness result of its own.** Yu & Zenker 2020 ([*Schemes, Critical Questions, and Complete Argument Evaluation*](https://link.springer.com/article/10.1007/s10503-020-09512-4)) start from the observation that complete evaluation of a scheme instance means asking *all* relevant CQs, note the field \"currently lacks a method for providing a complete argument evaluation\", and construct — at the meta-level, over the fully general schema \"premise(s); if premise(s), then conclusion; so conclusion\" — a CQ list they argue is \"both complete and applicable\", intended to generate appropriate object-level CQs for specific schemes. Read against the two-layer question: **completeness may be available for the CQ-generator even though it is not available for the surface-scheme list.** The closed basis, if there is one, sits under the critical questions (premise acceptability / warrant applicability / exceptions-rebuttals), not under the scheme names.\n\n**Formal argumentation frameworks don't need a closed set — by design.** ASPIC+ models schemes as defeasible inference rules and is parametric over *whatever* rule set you supply (the basic scheme for defeasible rules in Bench-Capon & Prakken 2010 is just \"if P1…Pn apply, then Q may be inferred\"); Carneades and DefLog likewise accommodate schemes without fixing an inventory (Macagno et al. 2017). AIF, the interchange format, treats schemes as Forms that argument nodes *may* instantiate, and the AIFdb corpora carry a \"Default Inference\" scheme node for edges that instantiate none — the formal layer already treats the scheme set as open and confessable. No computational framework surveyed requires exhaustiveness; only classifiers and annotation guidelines do, and they all cope by adding an escape category or shrinking the set.\n\n## §4 Coverage and reliability numbers from annotation studies\n\nThe two questions — *what fraction of real arguments fits some scheme?* and *can two people agree on which?* — have different answers, and the second is the binding constraint.\n\n**Coverage (no-fit rates):**\n\n| Study | Corpus | Scheme set | No-fit rate |\n|---|---|---|---|\n| Hansen & Walton 2013 | 256 election arguments (Ontario 2011, newspapers) | 14 schemes (Walton 2006) + none-of-the-above | **37.1%** initially |\n| same, after augmenting | same | +9 further schemes incl. 2 newly minted | **4.7%** residual unclassifiable |\n| Visser et al. 2018/2020 | US2016G1tv, 505 inference relations (Clinton–Trump TV debate) | all 60 main schemes of Walton et al. 2008 ch. 9 + \"default inference\" | **2.8%** (14 of 505) |\n| Visser et al. 2018/2020 | same 505 inferences | Wagemans PTA (36 types via 3 sub-classifications) | **17%** (85 of 505) Default Inference |\n\nSources: [Walton & Hansen](https://ccsenet.org/journal/index.php/jpl/article/download/29977/17760), [Visser et al. 2020](https://pmc.ncbi.nlm.nih.gov/articles/PMC7888437/), [Visser et al. 2018 ISSA](https://arg.tech/people/chris/publications/2018/issa2018.pdf). Reading these together: with the *full* 60-scheme catalogue, raw no-fit on televised political debate is small (≈3%) — but that register had already shaped the catalogue (Walton's examples are heavily drawn from media/political argument), and the moment Hansen & Walton took a *shorter* list into a *new* register, a third of the material fell outside and two genuinely new schemes had to be minted. Coverage is a function of how much curation the register has already received. The PTA's 17% shows the complementary failure: an a priori-complete system leaks through its reconstruction requirements rather than through missing cells. Note also distribution skew: the most frequent scheme in US2016 is *argument from example* \"by some margin\"; in the Ontario corpus, *appeal to negative consequences*; two textbook schemes had zero Ontario instances. Any fixed prompt-side scheme list will be exercised very unevenly per register.\n\n**Reliability (inter-annotator agreement):**\n\n- Walton 60-scheme annotation, US2016: Cohen's κ **0.723** (substantial) — achieved by two argumentation-trained annotators using a heuristic *decision tree* (later formalized as the Argument Scheme Key), not the raw compendium ([Visser et al. 2020](https://link.springer.com/article/10.1007/s10503-020-09519-x)).\n- PTA annotation, same corpus: aggregated κ **0.689**; sub-tasks: fact/value/policy κ 0.778, predicate/subject κ **0.851**, first/second order κ 0.658 (imbalance-depressed; raw agreement 98%). The basis dimensions are individually *more* codable than whole schemes — the factorial decomposition helps reliability even though its aggregation hurts coverage.\n- AMT guidelines (Musi, Ghosh & Muresan 2016): Fleiss κ **0.1** with nine minimally trained annotators on 30 persuasive essays, improving to **0.307** (fair) after guideline refinement and training — the \"finite hierarchical taxonomy\" did not rescue naive annotators ([paper](https://aclanthology.org/W16-2810.pdf)). The follow-up on argumentative microtexts got κ 0.296 with trained annotators, and one annotator marked NONE 21 times where the other two always picked a scheme — scheme-vs-no-scheme is itself a low-agreement judgment ([Musi et al. 2018, LREC](https://aclanthology.org/L18-1258.pdf)).\n- Lindahl et al. 2019 (Swedish newspaper editorials, Walton schemes): \"low agreement between the annotators\" (as characterized by Visser et al. 2020; primary numbers not consulted).\n- Duschl 2007 (school science interviews): started with 9 Walton schemes, **collapsed to 4 generic labels mid-study** to get agreement — the taxonomy deformed under reliability pressure (via Visser et al. 2020).\n- Standard practice in the computational literature is to pre-select a small subset: Feng & Hirst 2011 classify only the 5 most frequent Araucaria schemes; Green, Song et al., Schneider et al. all restrict per-domain (per Musi et al. 2016's survey).\n\n**LLM-era classification (2024–2025):**\n\n- Bezou-Vrakatseli, Cocarascu & Modgil, ACL Findings 2025 ([*Can Large Language Models Understand Argument Schemes?*](https://aclanthology.org/2025.findings-acl.702/)): first systematic eval; seven LLMs on Walton-scheme classification (EthiX dataset + generated arguments incl. enthymemes). Larger models \"satisfactory\" in few-shot with scheme descriptions; Claude 3.5 Sonnet consistently best; small models 40–50 F1 points behind; the paper's framing explicitly notes the task \"poses challenges even for human annotators.\"\n- Ruiz-Dolz, Kikteva & Lawrence, ACL 2025 ([*Mining Complex Patterns of Argumentative Reasoning in Natural Language Dialogue*](https://aclanthology.org/2025.acl-long.368/)): QT-Schemes corpus — 441 dialogue arguments annotated with **24** schemes (again: a working subset, not 60); first state-of-the-art results for scheme mining in dialogue.\n- Heinrich, Al Khatib & Stein, ArgMining 2025 ([*Multi-Class versus Means-End*](https://aclanthology.org/2025.argmining-1.19/)): encoding Visser's expert decision tree as a means-end procedure for LLMs performs close to flat multi-class while localizing *which decision-tree step* fails per scheme — machine-side evidence that the identification-key structure (not the flat list) is the usable interface to the taxonomy.\n\nNo study was found that runs the full 60-scheme catalogue as a flat classification target for models — every computational effort restricts, clusters, or proceduralizes the set first. That is itself a finding about the set's shape.\n\n## §5 What this means for Deliberus\n\n**1. The scheme enum is a curated open set — treat it exactly like the closed-enum finding the corpus already made.** The project's own three-instances finding (every taxonomy induced from one register breaks on the next; `docs/research/taxonomy-gaps-and-the-closed-enum.md`) is not a local quirk: it is the documented history of the scheme literature itself. Hansen & Walton is the canonical instance — new register (election discourse), 37% initial no-fit, two new schemes minted, ~5% irreducible residue *even after* expansion. Deliberus's pipeline classifies schemes from a closed enum today; the literature says the honest floor of \"fits no scheme\" is roughly 3–5% in well-curated registers and dramatically higher in fresh ones, and that the residue is *informative* (it is where Appeal to Fairness came from). The field's own production practice is the confession channel this project keeps converging on: the `default inference` label, institutionalized in the AIFdb corpora precisely so that \"examples not fitting any of the 60 schemes … elude classification\" visibly rather than by nearest-fit. A scheme field that cannot say \"none\" will do what Duschl's annotators did — silently deform the taxonomy to make everything fit.\n\n**2. The does-not-fit value has independent confession value beyond honesty.** Two empirical facts sharpen it: scheme-vs-no-scheme is itself a low-agreement judgment (the AMT outlier annotator), so `no_scheme` proposals should carry the same propose-only humility as terminus verdicts; and no-fit clusters are *scheme nurseries* — a periodic review of default-inference edges grouped by similarity is exactly the procedure by which Hansen & Walton minted new schemes. That is a graph-daemon-shaped job (flag-tier: cluster unclassified edges, propose a candidate pattern, never auto-mint).\n\n**3. A two-layer design is buildable — and both mature systems already are one.** The strong version of reading (b) — a small closed basis that *generates* Walton's 60 — is not supported: every candidate basis (PTA form×substance, Katzav-Reed conveyance relations, Macagno-Walton semantic-relation×reasoning-type, AMT loci, Lumer's epistemic branches) is closed only at its top layer and reopens at the layer where named schemes live. But the weak version is well-supported and is how the field actually operates:\n\n- *Closed layer* (small, factorial, machine-friendly, high-κ): argument form (alpha/beta/gamma/delta — subject/predicate sharing pattern), statement types (F/V/P), first/second order, reasoning direction (deductive/abductive/inductive per Macagno-Walton). Sub-task κ up to 0.85; derivable by elimination arguments; stable across registers because it is about sentence structure, not world-ontology.\n- *Open layer* (curated, named, register-sensitive, carries the CQs): the Walton-style scheme, understood as a *prototypical frozen combination* of closed-layer coordinates plus a semantic relation (Wagemans' lever ≈ Katzav-Reed's relation of conveyance ≈ the warrant's material content). This layer grows by induction and needs the confession value.\n\nWalton & Macagno 2015 describe precisely this: \"higher levels … developed a priori and the lower levels by empirical generalisations.\" For Deliberus, the closed layer is cheap to add to extraction (three or four low-entropy classifications the disagreement-preservation and stance instruments could also read), and it gives every edge a coordinate *even when the scheme slot says none* — which is the difference between \"unclassified\" and \"unclassified, but a first-order predicate argument from fact to value, so the generic VF critical questions apply.\"\n\n**4. CQ generation does not die when no scheme fits.** This is the concrete engineering payoff of Yu & Zenker 2020: a meta-level CQ list over the bare defeasible schema (premise acceptability; warrant/conditional acceptability; exceptions) is arguably complete and is designed to generate object-level CQs. A `no_scheme` edge can therefore still open a descent — with generic CQs parameterized by the closed-layer coordinates rather than scheme-specific ones. The current pipeline behavior (schemes drive CQ generation on every relationship edge) can degrade gracefully instead of silently: no scheme → generic-schema CQs, labeled as such. Note also the strength-layer coupling from the Session 20 audit: badge computation reads only scheme-bearing edges, so today an unclassifiable-but-real inference is invisible to strength. A generic-schema fallback with its own CQ set would close that gap without pretending classification succeeded.\n\n**5. Retire the periodic-table metaphor for Walton's set; keep it, with care, for the form layer.** Chemistry's table earns \"predictable gaps\" from a generative law. Walton's catalogue has no such law — it is a field guide, and the right tooling for a field guide is an identification key (which is what Visser's Argument Scheme Key and decision tree are, and what the means-end LLM results say machines want too). Wagemans' table does have a generative law, but only for form; its gaps are cells like \"2 sub PV\" that may simply be rare-to-empty in real discourse (11 of 505 US2016 arguments were second-order *at all*). If the ontology ever exposes scheme metadata publicly, the honest register is: \"curated inventory, open by construction, with a residue class\" — the same register the residue map already uses.\n\n**6. Register-sensitivity predicts where Deliberus's enum will break next.** The corpus has already seen claim-type break on a speaker-attitude hedge, residue-type on is-bedrock, epistemic status on doctrinal inference. The scheme literature adds: legal-doctrinal argument needs rule/precedent schemes formulated differently from conversational ones (Walton & Macagno 2015 discuss the divergence explicitly); election discourse needed fairness/priorities; scientific argumentation is, per Lumer, nearly absent from the compendium (statistical arguments, theory-choice arguments). The registers Deliberus is ingesting (legal, meta-analytic, referee reports) are exactly the ones the compendium covers thinnest. Expect scheme no-fits to concentrate there, and treat their clustering as signal.\n\n## §6 Sources\n\nPrimary items consulted via web search and fetch, August 2026. Items marked ⚠ were read only through secondary quotation or abstracts; claims from them are attributed accordingly.\n\n- Walton, D., Reed, C. & Macagno, F. (2008). *Argumentation Schemes*. Cambridge UP. Publisher page (96-scheme count, front matter, ToC): https://www.cambridge.org/core/books/argumentation-schemes/9AE7E4E6ABDE690565442B2BD516A8B6 ; front matter PDF: https://assets.cambridge.org/97805217/23749/frontmatter/9780521723749_frontmatter.pdf ; Internet Archive record (full ToC incl. the 60-scheme compendium list): https://archive.org/details/argumentationsch0000walt . ⚠ Body text (ch. 9–10) not read directly this pass.\n- NDPR review of *Argumentation Schemes* (\"sixty-five schemes\"; \"multiply schemes indefinitely\"): https://ndpr.nd.edu/reviews/argumentation-schemes/\n- Lumer, C. \"Walton's Argumentation Schemes\" (counts 60+44; \"never complete and always arbitrary\"; \"not even 1%\"; Kienpointner counts and deductivism critique): https://usiena-air.unisi.it/retrieve/e0feeaa8-a949-44d2-e053-6605fe0a8db0/A111.2_Lumer_Walton%e2%80%99sArgumentationSchemes_Print.pdf\n- Lumer, C. \"An Epistemological Appraisal / epistemological theory of argument schemes\" (deductive/probabilistic/practical branches; comet model): https://philarchive.org/archive/LUMASE\n- Walton, D. & Macagno, F. (2015). \"A classification system for argumentation schemes.\" *Argument and Computation* 6(3): https://journals.sagepub.com/doi/10.1080/19462166.2015.1123772 ; open text: https://philarchive.org/archive/WALACS-9\n- Macagno, F., Walton, D. & Reed, C. (2017). \"Argumentation Schemes. History, Classifications, and Computational Applications.\" *IfCoLog Journal* 4(8) (semantic-relation × reasoning-type × axiom decomposition; fever example; four-step inductive justification; ASPIC+/DefLog/Carneades fit): http://arg.tech/people/chris/publications/2017/flap2017.pdf ; also https://philarchive.org/archive/MACASH-10\n- Hansen, H.V. & Walton, D. (2013). \"Argument kinds and argument roles in the Ontario provincial election, 2011.\" *J. of Argumentation in Context* 2(2): https://www.jbe-platform.com/content/journals/10.1075/jaic.2.2.03han ; ResearchGate copy with tables (37.1%, 62.9%, 4.7%): https://www.researchgate.net/publication/314114512\n- Walton, D. & Hansen, H.V. \"Arguments from Fairness and Misplaced Priorities in Political Argumentation\" (the two newly minted schemes; the 14+none-of-the-above list): https://ccsenet.org/journal/index.php/jpl/article/download/29977/17760\n- Wagemans, J.H.M. (2016). \"Constructing a Periodic Table of Arguments.\" OSSA 11: https://pure.uva.nl/ws/files/26929021/Constructing_a_Periodic_Table_of_Arguments.pdf ; SSRN: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=2769833\n- Xie, Y. (2016). \"Commentary on Constructing a Periodic Table of Arguments.\" OSSA 11: https://scholar.uwindsor.ca/ossaarchive/OSSA11/papersandcommentaries/21\n- Wagemans, J.H.M. (2016). \"Reply to commentary…\": https://pure.uva.nl/ws/files/26929006/Reply_to_commentary_on_Constructing_a_Periodic_Table_of_Arguments.pdf\n- Periodic Table of Arguments site (PTA 3.0: form/substance/lever; lever page with the \"everyday knowledge suggests typical options\" passage; ATIP v5 2025): https://periodic-table-of-arguments.org/ ; https://periodic-table-of-arguments.org/argument-form/ ; https://periodic-table-of-arguments.org/argument-lever/ ; https://periodic-table-of-arguments.org/basic-terminology/ ; ATIP v5 PDF: https://periodic-table-of-arguments.org/wp-content/uploads/2025/09/wagemans-2025-argument-type-identification-procedure-atip-version-5-corrected.pdf\n- KRINO project PTA summary (\"factorial typology\"; 2×2×9=36): https://krino-project.com/theory/the-pta/\n- Visser, J., Lawrence, J., Reed, C., Wagemans, J. & Walton, D. (2020). \"Annotating Argument Schemes.\" *Argumentation* 35: https://link.springer.com/article/10.1007/s10503-020-09519-x ; PMC full text: https://pmc.ncbi.nlm.nih.gov/articles/PMC7888437/\n- Visser, J. et al. (2018). \"An annotated corpus of argument schemes in US election debates\" (ISSA; PTA corpus, 17% default, sub-task κs): https://arg.tech/people/chris/publications/2018/issa2018.pdf\n- Visser, J. et al. (2018). \"Revisiting Computational Models of Argument Schemes\" (COMMA): http://arg.tech/people/chris/publications/2018/comma2018-schemes.pdf ; Dundee copy: https://discovery.dundee.ac.uk/ws/files/28682480/Final_Published_Version.pdf\n- Katzav, J. & Reed, C. (2004). \"On Argumentation Schemes and the Natural Classification of Arguments.\" *Argumentation* 18(2): http://arg.tech/people/chris/publications/2004/arg2004.pdf ; PhilPapers: https://philpapers.org/rec/JKAOAS\n- Kienpointner, M. (1992). *Alltagslogik*. ⚠ Consulted via publisher abstract (\"about 60 schemes\"; Toulmin prototype; comprehensiveness aim): https://www.frommann-holzboog.de/reihen/71/710012620?lang=en-gb ; ToC scan (three-class structure, rule-using/rule-establishing/neither): https://www.sgipt.org/wisms/sprache/BegrAna/Plausib/KienpointnerAlltagslogik_IV.pdf\n- Shecaira, F.P. (2016). \"How to Disagree About Argument Schemes.\" *Informal Logic* 36(4) (purpose-relativity of schemes; taxonomy of disagreements incl. exhaustiveness): https://informallogic.ca/index.php/informal_logic/article/view/4610/4007\n- Yu, S. & Zenker, F. (2020). \"Schemes, Critical Questions, and Complete Argument Evaluation.\" *Argumentation* 34:469–498 (meta-level complete CQ list): https://link.springer.com/article/10.1007/s10503-020-09512-4\n- Musi, E., Ghosh, D. & Muresan, S. (2016). \"Towards Feasible Guidelines for the Annotation of Argument Schemes.\" ArgMining@ACL (AMT; Fleiss κ 0.1→0.307; survey of subset-selection practice): https://aclanthology.org/W16-2810.pdf\n- Musi, E. et al. (2018). \"A Multi-layer Annotated Corpus of Argumentative Text\" (LREC; microtexts, κ 0.296, NONE-outlier annotator): https://aclanthology.org/L18-1258.pdf\n- Song, Y. et al. (2014). \"Applying Argumentation Schemes for Essay Scoring\" (per-prompt protocols; \"generic\" category): https://aclanthology.org/W14-2110.pdf\n- Bezou-Vrakatseli, E., Cocarascu, O. & Modgil, S. (2025). \"Can Large Language Models Understand Argument Schemes?\" ACL Findings: https://aclanthology.org/2025.findings-acl.702/\n- Ruiz-Dolz, R., Kikteva, Z. & Lawrence, J. (2025). \"Mining Complex Patterns of Argumentative Reasoning in Natural Language Dialogue.\" ACL (QT-Schemes, 24 schemes): https://aclanthology.org/2025.acl-long.368/\n- Heinrich, M., Al Khatib, K. & Stein, B. (2025). \"Multi-Class versus Means-End: Assessing Classification Approaches for Argument Patterns.\" ArgMining: https://aclanthology.org/2025.argmining-1.19/\n\nNot verified this pass (flagged where used): the body text of Walton et al. 2008 (all chapter-level claims are via the ToC, the annotation literature, Lumer, and NDPR); Perelman & Olbrechts-Tyteca primary text; Lindahl et al. 2019's actual κ values; Grennan 1997 (mentioned in Walton & Macagno's lineage list; his scheme count was not established); the full PTA 3.0 lever tables (the site page describes them but the table contents were not extracted).\n"}