{"path":"research/cognitive-bias-codex-and-human-contribution.md","content":"# The Cognitive Bias Codex, and what humans can safely contribute\n\n*2026-08-27. Founder brought the Cognitive Bias Codex and asked what to make of it, how it bears on\nhuman contribution to Deliberus, and what risks it implies. Researched against the primary\nliterature rather than the poster. Companion to\n[what-human-judgment-is-for.md](what-human-judgment-is-for.md), whose four acts this either\nsupports or narrows, and to [the-residual-error-taxonomy.md](the-residual-error-taxonomy.md),\nwhose eight classes this gives a human twin.*\n\n---\n\n## 0. In one breath\n\nThe codex is organised by **what each bias is for**, not by what it gets wrong: four pressures\n(too much information, not enough meaning, must act now, cannot store everything) and roughly\ntwenty coping strategies, each with a price. Read that way it is not a list of human defects. It\nis **a map of what a mind does when it must decide under constraint** — and Deliberus's central\nstructural property is that **the graph never has to decide now**. That is the relation, and it is\na better claim than \"we debias people\", because it is about relieving a constraint rather than\nrepairing a person.\n\nFour results then bear directly on what humans should be asked to do here, and three of them\n*narrow* the ask rather than widening it:\n\n1. **Myside bias is independent of cognitive ability** — so there is no expert tier to gate on.\n2. **There is no general bias-proneness factor, and bias tasks measure it unreliably** — so\n   scoring a person's rationality is measuring noise. This kills reputation-weighting empirically,\n   not merely on values grounds.\n3. **The bias blind spot is not cured by awareness** and is, if anything, slightly *larger* in the\n   cognitively sophisticated — the literature under the principle *never ask a mind to audit its\n   own frame*.\n4. **Correlated errors do not cancel, they accumulate** — so recruiting more humans does not wash\n   out shared bias, it raises confidence in it. This is the human twin of same-hand bias.\n\nAnd one live risk the codex creates by existing: **a bias vocabulary is a weapon**. Naming a bias\nis the fastest way to dismiss an argument without engaging it. The graph already has the right\nanswer to this and did not know it (§5).\n\n---\n\n## 1. What the codex actually is\n\nBuster Benson's 2016 *Cognitive bias cheat sheet*, rendered as a poster by John Manoogian III,\nbuilt on Wikipedia's *List of cognitive biases* — roughly 175–188 entries, de-duplicated and\ngrouped into about twenty strategies under four problems\n([Benson](https://buster.medium.com/cognitive-bias-cheat-sheet-55a472476b18),\n[Daily Nous](https://dailynous.com/2016/09/14/cognitive-bias-codex/)). The file the founder\nbrought is the multilingual edition (English, Portuguese, Catalan, Basque).\n\nBenson's own framing is the part worth keeping: every bias is there **for a reason**, primarily to\nsave the brain time or energy. The four problems, in his words:\n\n| Pressure | Coping strategy | The price |\n|---|---|---|\n| **Too much information** | filter aggressively | we notice what is primed, repeated, changed, or confirms us — and never see what we filtered |\n| **Not enough meaning** | fill the gaps with story and pattern | we complete characteristics from stereotype and generality, and find patterns in noise |\n| **Need to act fast** | become confident, commit, finish | sunk cost, escalation, preference for simple-looking options over complex ones |\n| **What should we remember** | reduce, edit, generalise | events become lists, lists become gists, memories are revised after the fact |\n\n**Three cautions before using it for anything.** There is no widely accepted theory of where\nbiases come from, so a taxonomy sorted by hypothesised source may sort wrongly; the entries are\npresented as independent errors when they plainly are not; and Gigerenzer's whole programme\ndisputes the framing — under\n[ecological rationality](https://journals.sagepub.com/doi/abs/10.1111/j.1745-6916.2008.00058.x)\na \"biased\" mind that ignores information can be *more* accurate than an unbiased one that does\nnot, because rationality is correspondence with the world rather than coherence with logic. The\ncodex is a beautifully organised **inventory**, and the corpus's own discipline about borrowed\nauthority applies: cite it for its structure, never for its weight.\n\n---\n\n## 2. The generous reading: four pressures, four structural reliefs\n\nEach quadrant names a pressure that exists because a mind is finite and the moment is now. For\neach, Deliberus already has a structural answer — and in each case the answer is *the medium*,\nnot an instruction to think better.\n\n**Too much information → the filtered material has somewhere to go.** Filtering is the pressure;\nits cost is that what you filtered out leaves no trace, in you or in the record. This is exactly\n**coherent absence**, class 2 of the residual taxonomy, and it is the class the taxonomy names as\nleast defended. It is also the reason *nobody here has said X* is the first of the four human\nacts.\n\n**Not enough meaning → the gap-filling becomes writable.** A mind bridges a gap with a story and\ndoes not mark the join. Decomposition and implicit-premise extraction are that join, made\nvisible and therefore contestable. The unwritten crux finding — that the load-bearing premise is\nstated on one side and assumed on the other — is this quadrant measured in a real debate.\n\n**Need to act fast → the graph does not have to.** This is the sharpest of the four. Sunk cost,\nescalation of commitment, preference for complete-looking information, confidence as a\nprecondition for action: all of these are the price of *having to conclude*. A claim graph can\nhold `undecided`, can type a horizon rather than resolve it, and can leave a question open for\nyears without anything breaking. **The permissive zone and the first-class `undecided` verdict are\nliterally the anti-act-fast affordance**, and this is the strongest thing the codex says in\nDeliberus's favour.\n\n**What should we remember → persistence.** Memories are edited after the fact, in the direction of\nthe current self. A record is not. The temporal rung's first inhabitants — dates on every claim, a\nstaleness signal, ratified validity lapse — exist so that the record can age *honestly* instead of\nbeing silently revised.\n\nThe honest form of the claim: **Deliberus does not debias anyone. It builds a place where the\nconstraints that produce bias are absent, and lets a person put things there.** Whether that\nchanges what the person then thinks is an open empirical question and should be stated as one.\n\n---\n\n## 3. What this says about the four human acts\n\n[what-human-judgment-is-for.md](what-human-judgment-is-for.md) proposed four acts and argued\nratification is the worst available use of a human. The bias literature supports the analysis and\nnarrows it in one place.\n\n**Supported — no expert gate.** Stanovich, West & Toplak find myside bias has **very little\nrelation to intelligence**, in both naturalistic and within-subjects paradigms, and Stanovich's\n*dysrationalia* names the gap: rational thinking is a capacity distinct from the one IQ tests\nmeasure ([SAGE](https://journals.sagepub.com/doi/10.1177/0963721413480174)). The founder's own\nposition — that smart, wise and well-read people are victims too — is the literature's position.\nIt follows that gating contribution on expertise buys nothing on the axis that matters, which is\nindependent support for *an act's value is not its rung*.\n\n**Supported — never ask a mind to audit its own frame.** Pronin's bias blind spot is the\nrecognition of bias in others coupled with failure to see it in oneself, and the follow-up work\nfinds the blind spot **positively related to cognitive sophistication** — small in magnitude, and\ncrucially *not* mediated by actual susceptibility, meaning the sophisticated are not less biased,\nonly more confident they are not\n([Cambridge](https://www.cambridge.org/core/journals/judgment-and-decision-making/article/hypothesized-drivers-of-the-bias-blind-spotcognitive-sophistication-introspection-bias-and-conversational-processes/4DCC30BDA244D22EF1BA30AC75547364);\n[Pronin & Hazel 2023](https://journals.sagepub.com/doi/10.1177/09637214231178745)). Awareness is a\n*prerequisite* for correction and not a substitute for it. This is the empirical floor under the\nprinciple ratified today, and it also explains why **self-report about one's own reasoning is the\none testimony that deserves no elevated credence** — which is precisely the correction made to the\nthird act (a claim about one's own past intention is a self-interpretation; only the present\nrefusal to endorse is incorrigible, because it is a speech act rather than a report).\n\n**Narrowed — you cannot score a person.** The measurement literature is blunter than expected:\ncorrelations between bias measures are low, **suggesting the absence of any general factor of\nsusceptibility**, composite scores are unreliable, and test–retest reliabilities across seven\nclassic tasks ranged from 0 to .82, with framing effects and sunk cost repeatedly at the bottom\n([Frontiers review](https://www.frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2021.630177/full);\n[reliability paradox](https://pubmed.ncbi.nlm.nih.gov/28726177/)). Two consequences. Any future\nproposal to weight contributions by a contributor's demonstrated rationality is measuring noise\ndressed as a trait — **and this is an empirical kill, which is more durable than the values\nargument**, because a values argument invites a counter-values argument and this does not. And it\nretroactively supports demoting votes to instrument readings rather than promoting them to\nverdicts: an aggregate over people whose individual reliability is unmeasurable is not a\nmeasurement.\n\n---\n\n## 4. Four risks, ranked\n\n### 4.1 A bias vocabulary is a weapon — and the ontology already answers it\n\nNaming a bias is the cheapest available way to dismiss an argument without engaging it. Walton\nclassifies exactly this as the **bias subtype of ad hominem** — one of five, alongside direct,\ncircumstantial, poisoning the well and tu quoque — which \"identifies extra-logical motives for why\nP defends A's conclusion\"\n([Walton, *Ad Hominem Arguments*](https://www.uapress.ua.edu/9780817391140/ad-hominem-arguments/)).\n\n**Checked against the code: `bias` is already a shipped scheme** (`deliberus/extraction/schemes.py:73`)\nwith critical questions `bias_exists` (\"{person} has a conflict of interest or bias regarding\n{claim_topic}\") and `bias_influences` (\"{person}'s bias plausibly influenced their reasoning\").\nSo a bias accusation entering this graph is **not a verdict and not a conversation-ender** — it is\na defeasible inference that must answer its own two questions, and its answers are claims like any\nother. The system's answer to weaponised bias language is the answer it already gives to every\nother rhetorical move: *state it as a claim and let it be attacked*.\n\nTwo things this does not solve. The accusation is a claim **about a person**, and the corpus's\ncommitment is to map reasoning rather than people — so there is a live ontology question about\nwhether person-claims should be first-class here at all. And the move is self-referentially\ntrapped: belief bias means an opponent's ad hominem looks fallacious while one's own looks\nrelevant, so the people most confident that a bias accusation is warranted are exactly the ones\nwhose judgement of that is least trustworthy.\n\n**Design consequence: never build a bias-tagging affordance.** A dropdown of 188 biases attached\nto other people's claims would be an ad-hominem machine with an academic finish. If bias\nattribution enters, it enters the long way — as a claim, with its scheme, answering its questions.\n\n### 4.2 Correlated error: more contributors does not mean less bias\n\nThe wisdom of crowds requires errors to be **approximately independent**, each person erring in\ntheir own idiosyncratic direction so that errors cancel. Correlated errors do not cancel; they\naccumulate, and aggregation then **amplifies systematic error rather than cancelling random\nerror**. Crowd members sharing information sources or perceptual habits is sufficient to correlate\nthem.\n\nA bias, by definition, is the correlated kind. That is what distinguishes it from noise. So the\nintuition that a big enough crowd washes out bias is exactly backwards for the errors that matter,\nand it is **the human twin of same-hand bias** (taxonomy class 1, where generator and checker share\na model family). It also gives the daemon co-stimulation rule its human form: *two people from the\nsame reading community are one detector sampled twice.*\n\nPractical consequence for the workshop and any future contributor round: **register diversity is\nnot a nice-to-have, it is the condition under which aggregation means anything at all.**\n\n### 4.3 The live session is a deliberating group, and those have four documented failure modes\n\nSunstein's four failures of deliberating groups: predeliberation errors get **amplified rather\nthan merely propagated**; **cascades** form as later speakers follow earlier ones and withhold what\nthey know; **group polarization** moves the group further in its predeliberation direction, found\nin hundreds of studies across a dozen countries; and **hidden profiles**, where shared information\ncrowds out unshared, so the group never learns what its members knew\n([four failures](https://hls.harvard.edu/bibliography/four-failures-of-deliberating-groups)).\n\nThe hidden-profile failure is already in the corpus from the machine side — agent populations\nscored 17–36% against near-100% solo ceilings on hidden-profile tasks. It is the same failure. The\nco-present two-person modality inherits all four, and the design already has partial answers:\na persistent minted claim is exactly an anti-hidden-profile device (private information becomes an\nobject rather than something you must find the moment to say), and *nobody here has said X* is an\nanti-cascade device. **Polarization has no answer in the design and should be measured rather than\nassumed away** — pre/post attitude on the mapped question is a two-minute instrument and the\nworkshop should carry it.\n\n### 4.4 Making reasoning explicit may entrench it — contested, and testable here\n\nThe sharpest finding in the batch, and the one to hold most loosely. Fernbach, Rogers, Fox &\nSloman (2013) found that asking people for a **mechanistic explanation** of a policy reduced both\nthe illusion of explanatory depth and attitude extremity — but the effect **did not occur when\npeople were asked instead to enumerate reasons** for their preference, and asking people to\n*justify* their position makes beliefs more extreme\n([Psychological Science](https://journals.sagepub.com/doi/abs/10.1177/0956797612464058)).\n\n**Replication is genuinely mixed**: three preregistered close replications failed, later ones\nsucceeded, and the debate is open. So this is a candidate mechanism, not a fact to build on.\n\nBut note what it would mean if it holds, because Deliberus does **both** things and can tell them\napart. *Add a supporting claim* is enumerate-reasons — the entrenching condition. *Decompose this\ninto how it actually works, step by step* is mechanistic explanation — the moderating condition.\nIf the effect is real, decomposition is the debiasing act and endorsement is the entrenching one,\nand a product that makes endorsement easier than descent would be an extremity machine with a\ngraph on the front.\n\n**⚠ SUPERSEDED the same day — read [does-descending-change-the-descender.md](does-descending-change-the-descender.md).**\nA proper read found the prediction wrong in three ways and the testability claim half false. Short\nversion: the literature has **three** states, not two (Crawford & Ruscio's three preregistered\nclose replications nulled the attitude effect while the understanding-collapse *did* replicate;\nWalker et al. at N=5,139 found the **reverse**); the moderator is **what the claim bottoms out in**,\nwhich the terminus enum already encodes; and **decomposing a value claim into premises is\nstructurally the reason-enumeration condition** — the arm that does *not* moderate — so the naive\nmapping inverts for much of this graph. And the testability line was wrong: descent and endorsement\n*are* attributable per user, but `VoteRequest.vote` is binary `agree`/`disagree`, there is no\nattitude scale anywhere, and **a binary vote cannot show moderation**.\n\n---\n\n### 4.5 Belief bias, and the assumption this corpus already calls its weakest\n\n*Added 2026-08-27 on founder challenge — \"do you mean to tell me it would not be beneficial to take\ninto account the nature of any of the biases that have been studied?\" The answer is no, that is not\nwhat was meant, and this section is the demonstration: one specific entry, followed properly, lands\non the project's own weakest bet.*\n\n`assumption-ranking.md` ranks **two-axis separability** — that a person can rate *do I agree* and\n*is this well-argued* as independent judgments — as the **most novel and therefore least tested**\nclaim in the corpus, with *\"zero data on argument-specific two-axis voting.\"*\n\n**Belief bias is the named mechanism by which that assumption would fail**, and before today it\nappeared exactly once in this corpus, in an aside about ad hominem three sections above. Evans,\nBarston & Pollard (1983) is the canonical result: people endorse **believable conclusions as\nvalid and reject unbelievable ones as invalid**, and the effect is *stronger on invalid syllogisms*\n— i.e. the failure is accepting bad arguments whose conclusions you like, which is precisely the\nmotion a rigor axis is supposed to prevent.\n\n**But the refinement is the useful part, and it cuts in our favour.** ROC-based work reframes it as\na **response-bias effect**: believability shifts the *acceptance threshold* rather than destroying\nthe *ability to discriminate* valid from invalid, and a hierarchical Bayesian meta-analysis finds\nbelievability does not influence discriminability unconditionally, with individual differences\nmediating. So people **can** still tell a good argument from a bad one while they disagree with it.\nWhat moves is where they set the bar.\n\nThat converts the risk from fatal to **calibratable**, and it is measurable in this graph without\nany new instrument: for each rater, compare their rigor ratings on claims they agreed with against\nthose they disagreed with. A systematic offset is belief bias in our own data. **A rater whose\nrigor scores are uncorrelated with their agreement is the two-axis assumption holding**, and the\ncorpus has been calling that untested for over a year.\n\n**And the persuasion literature has our problem, knows it, and names our answer.** Social-psychology\nwork routinely establishes \"argument quality\" by *pretest* — arguments that evoke favourable\nthoughts are labelled strong — which O'Keefe and Jackson attack as circular, arguing that an\n**independently-motivated account of argument quality** is required instead. Deliberus arguably has\none already: a scheme plus its critical questions is a normative criterion rather than a popularity\nreading. An argument is well-supported here because its critical questions are answered, not\nbecause raters liked it. That is worth stating explicitly the next time the rigor axis is designed.\n\n**Which is the point of this section.** The refusal in §5 is about making bias categories into\nschema cells that users fill. It is not, and must not be read as, a refusal of the findings. This\none changed a design question.\n\n## 5. What not to do with the codex\n\n**Do not put a bias taxonomy in the ontology — meaning the categories as schema cells, never the findings as design knowledge** (§4.5 is a worked case of a single entry doing real work). Everything in §1's cautions applies: no accepted\ntheory of source, contested classification, non-independent entries, and a rival programme that\ndenies the framing. Importing 188 categories would hand every user the weapon of §4.1, and it would\nbe **borrowing** rather than building — but note the failure mode is *not* the same as assembly\ntheory's, which the corpus also cautions against. AT **overclaims a mechanism** (a sharp,\nfalsifiable central claim under sustained attack from its own field); the codex has **no theory to\ncontest** and its arrangement borrows the appearance of one. The shared part is what borrowing\ncosts, and the discriminator that decides it — *a borrowed category earns its place by yielding a\nmove* — is worked out in\n[what-belongs-in-the-ontology.md](what-belongs-in-the-ontology.md) §6b, with the Walton scheme set\nas the controlled case of an import done right.\n\n**Do not claim Deliberus debiases people.** The debiasing literature's own summary is that a\nmajority of interventions are *at least partially* successful, that no meta-analysis was possible\nbecause of heterogeneity in how debiasing is conceptualised and measured, and — the one line worth\nkeeping — that **technological interventions are more likely to succeed than cognitive strategies**\n([Ludolph & Schulz](https://journals.sagepub.com/doi/full/10.1177/0272989X17716672)). Deliberus is\na technological intervention rather than a cognitive one, which is favourable, and that is the\nstrongest available version of the claim. Anything stronger overstates a literature that does not\nsupport it.\n\n**Do not let the LLM assess its own debiasing.** A 2026 study found an intervention that produced\nrobust effects among **LLM-simulated participants and none at all among human readers**, and in\nits second study the model's simulated effects were directionally right but significantly larger\nthan the real ones ([arXiv:2605.01006](https://arxiv.org/abs/2605.01006)). This is the DeepMind\nfacilitation warning arriving from a second direction: **a model asked whether its own mediation\nhelped will say yes.** Any evaluation of whether the graph changed a person's reasoning must be\nmeasured on the person.\n\n---\n\n## 6. What this changes\n\n**Nothing shipped, and one principle gains its literature.** *Never ask a mind to audit its own\nframe* was ratified today from the argument that a coherence-optimising system cannot see its own\nframe; the bias blind spot is the same result for humans, including the finding that being clever\ndoes not help and being told does not fix it.\n\n**Two things it narrows.** No contributor-rationality scoring, ever — not as a values position but\nbecause the trait it would score does not measure reliably enough to exist. And no bias-tagging\naffordance, because the scheme layer already handles the move correctly and a shortcut would\nconvert it into an ad hominem machine.\n\n**Two things it adds.** Register diversity is promoted from good practice to a precondition for\naggregation meaning anything. And an attitude-movement instrument (pre/post on the mapped\nquestion) belongs in the workshop design, because polarization is the one deliberating-group\nfailure the architecture has no structural answer to.\n\n**One prediction registered before its data exists**: if Fernbach's distinction holds, users who\ndescend will moderate and users who only attach support will harden — and this corpus records the\ntwo acts separately, so it can check.\n\n---\n\n## Sources\n\n- [Benson, *Cognitive bias cheat sheet*](https://buster.medium.com/cognitive-bias-cheat-sheet-55a472476b18) · [Daily Nous on the codex](https://dailynous.com/2016/09/14/cognitive-bias-codex/) · [Wikipedia, List of cognitive biases](https://en.wikipedia.org/wiki/List_of_cognitive_biases)\n- [Stanovich, West & Toplak, *Myside Bias, Rational Thinking, and Intelligence*](https://journals.sagepub.com/doi/10.1177/0963721413480174) · [Dysrationalia](https://en.wikipedia.org/wiki/Dysrationalia)\n- [Pronin & Hazel, *Humans' Bias Blind Spot and Its Societal Significance* (2023)](https://journals.sagepub.com/doi/10.1177/09637214231178745) · [Hypothesized drivers of the bias blind spot](https://www.cambridge.org/core/journals/judgment-and-decision-making/article/hypothesized-drivers-of-the-bias-blind-spotcognitive-sophistication-introspection-bias-and-conversational-processes/4DCC30BDA244D22EF1BA30AC75547364)\n- [The Measurement of Individual Differences in Cognitive Biases](https://www.frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2021.630177/full) · [The reliability paradox](https://pubmed.ncbi.nlm.nih.gov/28726177/)\n- [Gigerenzer, *Why Heuristics Work*](https://journals.sagepub.com/doi/abs/10.1111/j.1745-6916.2008.00058.x) · [Homo Heuristicus](https://constable.blog/wp-content/uploads/2021/12/2009-gigerenzer-brighton-homo-heuristicus.pdf)\n- [Walton, *Ad Hominem Arguments*](https://www.uapress.ua.edu/9780817391140/ad-hominem-arguments/) · [Ad Hominem Fallacies, Bias, and Testimony](https://www.researchgate.net/publication/257520212_Ad_Hominem_Fallacies_Bias_and_Testimony)\n- [Sunstein, four failures of deliberating groups](https://hls.harvard.edu/bibliography/four-failures-of-deliberating-groups) · [The Law of Group Polarization](https://www.researchgate.net/publication/279547233_The_Law_of_Group_Polarization)\n- [Fernbach, Rogers, Fox & Sloman (2013)](https://journals.sagepub.com/doi/abs/10.1177/0956797612464058) · [Three preregistered failures to replicate](https://www.researchgate.net/publication/349839502) · [Illusion of explanatory depth](https://en.wikipedia.org/wiki/Illusion_of_explanatory_depth)\n- [Ludolph & Schulz, *Debiasing Health-Related Judgments and Decision Making*](https://journals.sagepub.com/doi/full/10.1177/0272989X17716672) · [Retention and transfer of bias-mitigation interventions](https://www.frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2021.629354/full)\n- [Feroz & Kunst, *Can AI Debias the News?* (arXiv:2605.01006)](https://arxiv.org/abs/2605.01006)\n- [van Gelder on argument mapping and critical thinking](https://thinkeranalytix.org/wp-content/uploads/2018/09/TvG-Using-argument-mapping-to-improve-critical-thinking-skills-2015.pdf) — 0.85 SD in high-intensity courses, measured on critical-thinking tests rather than on bias in the wild, by the developer of the mapping software\n"}