{"path":"research/empirical-deliberation-convergence-evidence.md","content":"# The Empirical Record on Deliberation and Convergence\n\n**Date:** 2026-07-05\n**Status:** AI-conducted research synthesis. Primary studies and official reports cited inline with URLs. Confidence levels stated per claim.\n\n## Framing\n\nDeliberus rests on a wager: that most deep disagreement among people of good will is, at bottom, semantic confusion plus differently-interpreted feasibility sitting on a foundation of shared care — and that with enough high-quality decomposition and structured deliberation, positions substantially converge. This document assembles the honest empirical record on what actually happens to deep moral and political disagreement when real, high-quality deliberation is run at scale. It reports the strongest positive evidence and the strongest negative evidence with equal rigor, then attempts a \"residue profile\": when good-faith structured deliberation completes, roughly what fraction of disagreement dissolves, and what does the remainder look like?\n\nThe honest headline, stated up front: **the evidence is genuinely encouraging on some axes and genuinely sobering on others, and the two are not in contradiction — they apply to different kinds of disagreement.** Deliberation reliably moves affective hostility, factual understanding, and policy-preference extremity. It reliably does *not* dissolve disagreements whose roots are divergent worldview priors and value-weightings. The convergence wager is best read as strongly supported for a large class of disagreement and clearly bounded for another.\n\n---\n\n## (a) Citizens' Assemblies\n\n**Ireland — Citizens' Assembly on the Eighth Amendment (abortion), 2016–2017.** A 99-member body chosen by random selection to mirror the Irish population — including pro-life, pro-choice, and undecided members — deliberated over five weekends, heard 25 experts, and reviewed submissions. By the end, 87% of members agreed the constitutional provision was unfit for purpose; on the question of unrestricted access, 64% voted in favour of terminations without restriction. In the subsequent 2018 referendum, 66.4% of Irish voters chose to repeal the Eighth Amendment. The alignment between the assembly's 64% and the national 66.4% — a gap of roughly two percentage points — is the single most-cited datapoint in the \"deliberative wave\" literature ([Electoral Reform Society](https://electoral-reform.org.uk/the-irish-abortion-referendum-how-a-citizens-assembly-helped-to-break-years-of-political-deadlock/); [Involve](https://www.involve.org.uk/news-opinion/opinion/citizens-assembly-behind-irish-abortion-referendum); [Participedia](https://participedia.net/case/irish-citizens-assembly-the-eighth-amendment)).\n\nTwo honest caveats. First, the assembly did not start from a neutral baseline: members were \"in large part in favour of liberalisation\" from the outset, and the body moved further in that direction ([Wikipedia: Citizens' assembly](https://en.wikipedia.org/wiki/Citizens%27_assembly)). Second, the celebrated 64%↔66% alignment is a *coincidence of aggregates*, not a measured before/after opinion shift of the same individuals — the public referendum majority and the assembly majority landing near each other is powerful evidence of *representativeness*, but weaker evidence that deliberation itself *changed* minds. Research does suggest a downstream effect on the wider public: voters aware that a Citizens' Assembly preceded the vote were more willing to support the liberal position and to deviate from the status quo ([Participedia](https://participedia.net/case/irish-citizens-assembly-the-eighth-amendment)). The same body's earlier work on marriage equality similarly preceded the 2015 referendum that passed same-sex marriage. *Confidence: high on the vote figures; moderate on the causal \"deliberation changed minds\" interpretation.*\n\n**France — Citizens' Convention for Climate (CCC), 2019–2020.** 150 citizens selected by sortition deliberated across eight weekends and produced 149 measures, many passing with near-universal support; the lowest-support proposal (reducing the motorway speed limit to 110 km/h) still passed at 59.7% ([Wikipedia: Citizens Convention for Climate](https://en.wikipedia.org/wiki/Citizens_Convention_for_Climate); [IDDRI key messages](https://www.iddri.org/sites/default/files/PDF/Publications/Catalogue%20Iddri/Etude/ST0720-CCC%20EN.pdf)). This is a striking *internal* consensus rate on genuinely contested, high-stakes policy. But the CCC is also the field's cautionary tale on two fronts: (1) the deliberation ran in **silos** — five fixed thematic groups meant most citizens voted on measures they had never examined in their own deliberations, undercutting the \"collective intelligence of 150\" claim ([Verfassungsblog](https://verfassungsblog.de/lessons-from-the-french-citizens-climate-convention/)); and (2) the \"sans filtre\" political promise collapsed — of 149 measures, three were dropped by President Macron and many others were weakened in legislation, leaving members to give parliament failing grades at a reconvened evaluation session ([Deliberative Democracy Digest](https://www.publicdeliberation.net/the-promises-and-disappointments-of-the-french-citizens-convention-for-climate/)). The internal consensus was real; the downstream impact was not. *Confidence: high.*\n\n**The OECD \"deliberative wave\" report (2020).** *Innovative Citizen Participation and New Democratic Institutions* analysed 289 case studies (763 deliberative panels, 1986–2019). Its central claim: convening a broad cross-section for numerous days to learn and deliberate \"is an effective way of overcoming polarisation and finding consensus on the thorniest policy dilemmas… particularly true for issues that are values-driven, require weighing trade-offs, and involve long-term concerns\" ([OECD full report](https://www.oecd.org/en/publications/innovative-citizen-participation-and-new-democratic-institutions_339306da-en/full-report.html)). A follow-up database update identified close to 600 cases, 101 since 2019. Permanent institutions have emerged, notably the **Ostbelgien** (East Belgium) model — a standing Citizens' Council with agenda-setting power drawn by lottery ([Medium/Participo](https://medium.com/participo/available-now-oecd-catching-the-deliberative-wave-report-2a922f3144ed)).\n\nAn important interpretive flag: the OECD is an *advocate* institution here, and its framing (\"effective way of overcoming polarisation and finding consensus\") is a curated summary of practitioner-collected cases rather than a randomized-controlled evidence base. It documents that assemblies *reach* recommendations; it is less rigorous on counterfactual mind-change. *Confidence: high that assemblies reliably reach collective recommendations; moderate on the stronger \"overcoming polarisation\" causal claim, given selection and advocacy effects.*\n\n---\n\n## (b) Deliberative Polling — America in One Room (2019)\n\nJames Fishkin's Deliberative Polling is the most methodologically rigorous instrument in this literature because it uses a **randomly-sampled treatment group plus a control group** measured on the same items. *America in One Room* (A1R, September 2019) brought ~500+ registered voters to a hotel outside Dallas to deliberate over a long weekend on five issues (immigration, economy, health care, environment, foreign policy) ([Wikipedia: America in One Room](https://en.wikipedia.org/wiki/America_in_One_Room)).\n\nResults, published in the *American Political Science Review* (2021, Fishkin, Siu, Diamond, Bradburn): participants showed \"large, depolarizing changes in policy attitudes and large decreases in affective polarization.\" Across 26 policy proposals identified as extremely partisan, participants moved toward the center on **22 of 26**, with the shift statistically significant on **19** ([APSR/Cambridge](https://www.cambridge.org/core/journals/american-political-science-review/article/is-deliberation-an-antidote-to-extreme-partisan-polarization-reflections-on-america-in-one-room/5DEFB6F8D944ECDE77A5E80C3346D4DE)). The effect was strongest on *substantive complex policy matters* and weaker on *identity-saturated symbolic issues* — a pattern that recurs throughout this literature and is directly relevant to the Deliberus wager. Knowledge gains were also measured (reported in the online appendix, Table A6). *Confidence: high; this is a rigorous field experiment.*\n\n**Durability.** A one-year follow-up (2020, on the cusp of the presidential election) reached 463 of 523 original participants and found that \"long-term shifts in delegates' political views had occurred\" relative to the control group — i.e., the changes largely *persisted* rather than fully reverting ([Deliberative Democracy Lab](https://deliberation.stanford.edu/america-in-one-room); [NORC](https://www.norc.org/research/projects/america-in-one-room.html)). Fishkin also reports a behavioral durability marker: deliberators voted \"overwhelmingly according to their post-deliberation policy preferences\" up to a year later ([Persuasion/Fishkin](https://www.persuasion.community/p/deliberation-can-save-democracy)).\n\n**The honest counterweight on durability.** The broader Deliberative Polling literature is more mixed than A1R alone suggests. Reviews note that DP \"creates dramatic, statistically significant changes in views\" but that \"some of these changes are reversed over time,\" with documented backsliding between on-site measurement and midterm follow-ups ([Involve](https://www.involve.org.uk/resource/deliberative-polling); [Farrar et al., New Haven experiment](https://personal.lse.ac.uk/list/PDF-files/NewHaven.pdf)). Durability appears to depend on issue, pre-existing dispositions, and whether one measures *policy attitudes* (which can fade) versus *reason-based engagement and voting propensity* (which persists more robustly) ([Smets & Isernia, 2014](https://journals.sagepub.com/doi/10.1177/1465116514533016)). A1R may be a relatively strong case; the modal case shows more decay. *Confidence: high on the existence of both persistence and decay; moderate on which dominates in general.*\n\n---\n\n## (c) Depolarization Interventions\n\n**The Stanford Strengthening Democracy Challenge megastudy (Science, 2024).** This is the strongest single body of evidence on what moves partisan animosity. From 252 submissions, 25 interventions were tested head-to-head in a uniform randomized design with **>32,000 participants**, measuring three outcomes: partisan animosity, support for partisan violence, and anti-democratic attitudes ([Science](https://www.science.org/doi/10.1126/science.adh4764); [Northwestern IPR working paper](https://www.ipr.northwestern.edu/documents/working-papers/2022/wp-22-38.pdf)).\n\nKey findings:\n- **Partisan animosity is highly movable.** Contrary to expert forecasters' expectations, **23 of 25** interventions significantly reduced it. The two most effective *highlighted sympathetic, relatable individuals from the other side*: the largest effect was the \"Contact Project\" video (**Cohen's *d* = −0.53**), and the fourth-largest was \"Civity Storytelling\" (**Cohen's *d* = −0.45**) ([Stanford Impact Labs](https://impact.stanford.edu/article/megastudy-testing-25-treatments-reduce-antidemocratic-attitudes-and-partisan-animosity); [Cornell Chronicle](https://news.cornell.edu/stories/2024/10/how-get-democrats-republicans-strengthen-democracy-0)).\n- **Anti-democratic attitudes and support for violence are much harder.** There was \"little overlap\" between the interventions that moved animosity and those that moved support for undemocratic practices or violence — implying *separate causal mechanisms*. Warm feelings toward the other side do **not** automatically translate into democratic commitment.\n\nThe critical lesson for a convergence platform: **liking your opponents and agreeing with (or tolerating) their politics are distinct achievements with distinct levers.** *Confidence: very high; this is a preregistered megastudy with an exceptional N.*\n\n**Moral reframing (Feinberg & Willer).** Arguments recast in the *audience's* moral language are more persuasive across the aisle — e.g., conservatives support environmental policy more when it is framed around \"purity\" rather than \"harm\" ([Feinberg & Willer 2019 review](https://sdimakis.github.io/moral_psychology/readings/week_10/Feinberg_2019.pdf)). But two sobering facts: fewer than 10% of people spontaneously reframe to their opponents' values, even though 64–85% can *recognize* the reframed argument as more persuasive when shown it ([From Gulf to Bridge, Feinberg & Willer 2015](https://www.scicom-bellagio.com/wp-content/uploads/2017/10/Feinberg-and-Willer-2017_From-Gulf-to-Bridge-When-Do-Moral-Arguments-Facilitate-Political-Influence.pdf)). Note: this is a *narrative review*, not a pooled meta-analysis; a single aggregate effect size is not established, and later work finds the effect is conditional (e.g., stronger when targeting conservative moral foundations). *Confidence: moderate-to-high on directional effect; low on a precise pooled magnitude.*\n\n**Perception-gap corrections (More in Common, 2019).** Democrats and Republicans imagine roughly *twice* as many opponents hold \"extreme\" views as actually do; both sides overestimate opponents' immoderate views by ~20 percentage points or more ([The Perception Gap report](https://perceptiongap.us/media/zaslaroc/perception-gap-report-1-0-3.pdf); [More in Common](https://moreincommonus.com/case_study/americas-perception-gaps/)). Concretely: 15% of Democrats agree \"most police are bad people\" while Republicans guess 52%; 21% of Republicans deny racism exists while Democrats guess 49%. This is *direct evidence for the semantic/perceptual component of the Deliberus wager* — much apparent disagreement is a mirage of misperception. But the sting: the perception gap is *worse* among the most news-consuming and highly-educated, and greater information does not reliably fix it ([Heterodox Academy](https://heterodoxacademy.org/blog/social-science-partisan-perception-gap/)). *Confidence: high.*\n\n**Intergroup contact (Pettigrew & Tropp, 2006).** The canonical meta-analysis: **515 studies, 713 samples, >250,000 subjects, 38 nations**, mean effect **r = −.21** between contact and prejudice; 94% of studies show an inverse relationship, and the *more rigorous experimental* studies yield a *larger* mean effect (**r = −.33**) ([Pettigrew & Tropp PDF](https://ideas.wharton.upenn.edu/wp-content/uploads/2018/07/Pettigrew-Tropp.pdf)). Mechanisms: reduced anxiety and increased empathy/perspective-taking mediate more strongly than knowledge gain ([mediators paper](https://www.researchgate.net/publication/229779236_How_Does_Intergroup_Contact_Reduce_Prejudice_Meta-Analytic_Tests_of_Three_Mediators)). Structured deliberation is, among other things, a high-quality contact intervention — which is likely *why* the affective-depolarization findings above are so robust. *Confidence: very high.*\n\n---\n\n## (d) The Negative Results — With Equal Rigor\n\n**The Forecasting Research Institute's adversarial collaboration on AI existential risk (2024).** This is the cleanest negative case in the literature, and it deserves close attention because it is a *best-case* deliberation that still failed to converge. 11 \"AI skeptics\" (nine superforecasters, two domain experts) and 11 \"AI concerned\" domain experts, disagreeing sharply on P(existential catastrophe from AI by 2100), engaged for 8 weeks — skeptics investing a median of **80 hours**, the concerned group 31 — reading materials, forecasting, and holding moderated one-on-one crux-generation calls ([FRI report](https://forecastingresearch.org/ai-adversarial-collaboration); [EA Forum summary](https://forum.effectivealtruism.org/posts/orhjaZ3AJMHzDzckZ/results-from-an-adversarial-collaboration-on-ai-risk-fri)).\n\nThe result: **almost no convergence.** Skeptic median moved 0.10% → 0.12%; concerned median moved 25% → 20%. The ~20-point gap barely closed, and participants attributed much of what movement occurred to real-world events (the GPT-4 release) rather than to each other's arguments ([FRI report](https://forecastingresearch.org/ai-adversarial-collaboration)). Critically, this was **not a failure of mutual understanding**: both groups could accurately summarize each other's arguments, and most reported satisfaction with their counterparts.\n\nWhy did it fail? The researchers' diagnosis is exactly the distinction that should govern the Deliberus wager. The disagreement traced to **fundamental worldview differences, not resolvable facts**:\n- **Different anchoring priors.** Skeptics assumed \"the world usually changes slowly, making rapid extinction unlikely\"; the concerned anchored on \"the arrival of a higher-intelligence species… has often led to the extinction of lower-intelligence species.\"\n- **Different evidential standards.** The concerned \"was more willing to place weight on theoretical arguments with multiple steps of logic\"; skeptics \"tended to doubt the usefulness of such arguments for forecasting.\"\n- **Empirically empty cruxes.** The best available near-term crux (a METR/ARC-Evals autonomous-replication result by 2030) was expected to close the gap by only ~1.2 percentage points out of 22.7. Many commonly-discussed questions (economic growth rates, publication counts) had *zero* forecasted value-of-information — resolving them would move *neither* group.\n\nThe one genuinely hopeful note: disagreement *shrank as the horizon extended* — by the 1000-year frame both groups put P(extreme negative outcomes) at 30–40%. Deep-enough decomposition found shared ground, but only by moving to a question so abstract it had little decision-relevance ([Risk Analysis, Rosenberg et al. 2025](https://onlinelibrary.wiley.com/doi/abs/10.1111/risa.70023)). *Confidence: high; this is a well-documented structured study, and it is the strongest single caution against the strong form of the convergence wager.*\n\n**Group polarization (Sunstein, \"The Law of Group Polarization,\" 1999/2002).** Deliberation among the *like-minded* reliably moves the group *toward the extreme* in its pre-deliberation direction — the mechanism being persuasive-argument pooling plus social comparison ([Sunstein SSRN](https://papers.ssrn.com/sol3/papers.cfm?abstract_id=199668); [Chicago Unbound](https://chicagounbound.uchicago.edu/law_and_economics/542/)). This is the mirror image of the citizens'-assembly results and the reconciling factor is *composition*: deliberation converges toward the reasonable center when the room is *diverse and well-moderated*, and diverges toward the extreme when the room is *homogeneous*. Sunstein's own conclusion — \"pay far more attention to the circumstances and nature of deliberation, not merely to the fact that it is occurring\" — is a direct design constraint. *Confidence: high.*\n\n**Motivated numeracy (Kahan et al.).** The deepest challenge. On politically charged data, *higher* numeracy and *higher* cognitive reflection are associated with *greater* polarization, not less — people deploy quantitative skill selectively to reach identity-congenial conclusions ([Kahan et al., Motivated Numeracy](https://rcgd.isr.umich.edu/wp-content/uploads/2018/07/motivated_numeracy_and_enlightened_selfgovernment.pdf)). The unsettling implication: identity-protective cognition operates *through* effortful System-2 reasoning, not merely around it. Giving good-faith, high-capacity people more information and better tools can *widen* factual disagreement when identity is at stake. A Western-European replication broadly reproduced the pattern while noting that ordinary science-comprehension deficits and identity-protective cognition can *both* be in play ([Cambridge/Behavioural Public Policy](https://www.cambridge.org/core/journals/behavioural-public-policy/article/motivated-numeracy-and-active-reasoning-in-a-western-european-sample/5C462D05ED2D1FCAD2715F10287E6A0A)). The mechanism debate continues, but the core finding is robust enough to take seriously as a design threat. *Confidence: moderate-to-high on the effect; the precise mechanism remains contested.*\n\n**Polis / vTaiwan — reality vs. hype.** The genuinely positive datapoint: in the 2015 UberX deliberation, 31,115 votes surfaced a cross-cluster consensus on passenger safety that ~95% of participants agreed on, and over 80% of vTaiwan deliberations have led to decisive government action ([Democracy Technologies](https://democracy-technologies.org/participation/consensus-building-in-taiwan/); [arXiv 2502.05017](https://arxiv.org/html/2502.05017v1)). The honest caveat: Polis's \"group-informed consensus\" metric is engineered to *surface bridging statements and hide divisive ones* — it finds the agreement that exists rather than manufacturing agreement where none does, and the divisive statements do not disappear, they are de-emphasized. Polis is powerful evidence that *latent* common ground is often larger than it appears (the perception-gap finding, operationalized) — not evidence that deep value conflicts dissolve. *Confidence: high on the mechanism; the \"80% → government action\" figure is a practitioner-reported statistic and should be read as such.*\n\n---\n\n## (e) A Residue Profile\n\nCan we state, across the best studies, what fraction of disagreement dissolves under good-faith structured deliberation, and what the remainder looks like? **A single universal number is not supported by the literature, and any doc claiming one is overreaching.** What *is* supported is a structured decomposition of the disagreement into components that behave very differently:\n\n1. **Misperception / semantic mirage — largely dissolvable.** The perception-gap work says a substantial slice of apparent disagreement (~20 percentage points of imagined extremity) is simply *wrong beliefs about what the other side thinks* ([Perception Gap](https://perceptiongap.us/media/zaslaroc/perception-gap-report-1-0-3.pdf)). Contact and deliberation correct much of this. This is the component most favorable to the Deliberus wager, and it is not small.\n\n2. **Affective hostility — reliably reducible.** The megastudy's *d* = −0.45 to −0.53 for the best interventions, and contact's r = −.21 to −.33, show that animosity moves substantially. But it is a *feeling*, not a *judgment* — reduced hostility does not itself resolve the policy question.\n\n3. **Factual/feasibility disagreement — conditionally reducible, sometimes backfiring.** Deliberation improves knowledge and shifts policy extremity on *complex substantive* issues (A1R: 22/26 proposals). But where identity is fused to the fact, motivated numeracy shows information can *widen* the gap.\n\n4. **Worldview-prior and value-weighting disagreement — largely resistant.** This is the FRI residue: different base rates for how the world changes, different standards for what counts as evidence, different weightings of the same acknowledged considerations. Here even 80 hours of best-practice adversarial collaboration between mutually-respecting experts moved almost nothing.\n\nThe most defensible summary: **when good-faith structured deliberation completes, the components rooted in misperception and affect move a great deal; the components rooted in complex feasibility move some (and occasionally backfire); the components rooted in divergent worldview priors and value-weightings move very little.** The residue is disproportionately category 4. Notably, deliberation's *social* achievement — mutual respect and the capacity to state the other's view fairly — is real and durable *even when the substantive gap remains*. The scholarly term of art for the productive endpoint is Sunstein's **\"incompletely theorised agreement\"**: participants converge on *what to do* (e.g., \"put it to referendum,\" \"adopt this safety regulation\") while retaining different *reasons why* ([Trinity College Law Review](https://trinitycollegelawreview.org/citizens-assembly/)). Convergence on action without convergence on justification is the modal successful outcome — not full value-convergence.\n\n---\n\n## Implications for Deliberus\n\nAn honest empirical prior for the convergence wager, stated as three graded claims:\n\n**Strongly supported:** A large fraction of what *looks* like deep disagreement is misperception, affective hostility, and unexamined semantic ambiguity — and this fraction is substantial, movable, and exactly what a high-quality decomposition-and-deliberation system is built to surface. The perception-gap magnitudes, the megastudy effect sizes, the A1R depolarization on complex issues, and Polis's bridging results all vindicate the \"shared care beneath the surface, obscured by confusion\" half of the wager. Deliberus is building the right instrument for the largest and most tractable slice of the problem.\n\n**Bounded:** The wager's *strong* form — that positions \"substantially converge\" given *enough* decomposition — is not supported for the class of disagreement rooted in divergent worldview priors and value-weightings. The FRI adversarial collaboration is the decisive case: mutual understanding was achieved, cruxes were identified, and convergence still did not happen, because the disagreement was not a confusion to be dissolved but a genuine difference in how to weigh acknowledged considerations. Decomposition *far enough* tends to reach either shared abstractions (the 1000-year frame) or bedrock value-primitives — and the shared abstractions are often too abstract to be decision-relevant, while the value-primitives are simply different.\n\n**Design consequence:** The most valuable thing Deliberus can do is *not* to promise convergence but to *sort the disagreement* — to make legible, for any given contested claim, how much of the gap is (1) misperception, (2) affect, (3) contested feasibility, or (4) irreducible value-weighting. That sorting is itself the product. It is also the honest one: the composition finding (Sunstein) means a system that puts *diverse* participants in *well-moderated* structured contact will trend toward the reasonable center, whereas one that lets enclaves form will amplify extremity — so moderation quality and viewpoint diversity are not features but load-bearing walls. And \"incompletely theorised agreement\" — convergence on action while preserving plural justifications — is a more empirically-grounded target than convergence on truth.\n\nThe wager should therefore be reframed, not abandoned: **most good-faith disagreement contains far more dissolvable confusion than participants believe, and surfacing that is transformative — but a real residue of value and worldview difference will remain, and a mature platform earns trust by mapping that residue precisely rather than pretending to dissolve it.**\n\n---\n\n## Strongest Objection to My Own Conclusions\n\nThe strongest objection is that this synthesis may be **too pessimistic about the residue because the negative cases are drawn from an unusually adversarial and abstract domain.** The FRI collaboration concerned a single, extraordinarily hard question — the probability of human extinction from a technology that does not yet exist — where \"cruxes\" are inherently unresolvable within the study window and worldview priors do essentially all the work. Most disagreements a deliberation platform would actually host (local land use, school policy, healthcare trade-offs, even abortion within a shared legal frame) are *more* fact-laden and *less* about ur-priors on how the cosmos changes. On those, the Irish assembly's near-perfect aggregate alignment, A1R's durable depolarization, and the megastudy's breadth suggest convergence may be *more* achievable than the FRI case implies. If the residue is domain-dependent — large for speculative existential questions, small for concrete policy trade-offs — then Deliberus's real-world docket might sit mostly in the favorable regime, and the \"bounded\" claim above would be overstated for its actual use cases.\n\nThe counter-counter, held honestly: motivated numeracy and group polarization are *not* domain-exotic — they appear on gun control, climate, and vaccines, which are exactly the concrete-policy issues a platform would host — so the resistant residue is unlikely to vanish even off the existential-risk frontier. The truthful position is that the *size* of the irreducible residue is genuinely uncertain and probably varies by issue, and that this uncertainty is itself the strongest argument for building Deliberus as an instrument that *measures* the residue empirically, case by case, rather than one that assumes its magnitude in advance.\n"}