{"path":"research/game-feel-and-the-first-screen.md","content":"# Game feel and the first screen — what the doorstep owes a stranger\n\n**Date**: 2026-08-28 · **Status**: research + design proposal, nothing ruled\n**Origin**: the founder's own diagnosis, and a sketch he drew more than a decade before the genre\nthat solved it existed.\n\n---\n\n## 0. The bind, in his words\n\n> *\"I have clung to the idea of getting fellow nerds to UNDERSTAND the foundation/convictions/\n> principles/design of Deliberus in the absence of it being deployed, fully functioning, heavily\n> populated… Since I imagine I'll need help to be able to continue building.\"*\n\nThis is an honest account of why the public surfaces over-explain, and it is a real bind rather than\na mistake: without a working, populated artifact, explanation is the only thing there is to offer.\nBut it is **condition (d) failure at first contact** ([fractal-priors-for-success.md](fractal-priors-for-success.md)):\nthe reader is asked to learn the code before receiving any value. Double-entry bookkeeping never did\nthat — merchants already tracked what they owed, and the notation gave their existing practice a\nshape. Esperanto did do it, and lost.\n\n**The resolution is not less explanation. It is payoff before understanding.**\n\n---\n\n## 1. The premise correction, which protects depth rather than limiting it\n\nThe founder's stated basis was that attention \"has been ground down over decades by social media and\nshort-form video.\" **The capacity half of that is myth and the behaviour half is solidly measured,\nand the two point at different designs.**\n\n**Myth**: the eight-second attention span. It traces to a 2015 Microsoft Canada report citing a\nmarketing firm, with no peer-reviewed research behind it, and the goldfish comparison is invented\ntwice over (fish attend far longer than nine seconds). A 2025 literature review found social-media\neffects on sustained attention to be *small, short-lived and highly variable*. **No major study shows\nbaseline capacity declining.**\n\n**Measured**: Gloria Mark's logging research, replicated by five independent studies 2014–2020\n(44 s, 47 s, 50 s), finds average time on a single screen fell **2.5 minutes (2004) → 75 seconds\n(2012) → 47 seconds (since ~2016)**.\n\n**So the phenomenon is faster triage, not reduced capacity** — in the reviewers' phrasing, people\nskim and switch faster *\"not because we can't focus, but because we've learned to decide more quickly\nwhat deserves our focus.\"* Which inverts the design consequence:\n\n| If the premise were… | The design would be… |\n|---|---|\n| capacity is degraded | permanent brevity; a ceiling on depth; simplify everything, forever |\n| **triage is faster** *(what the evidence supports)* | **the first screen must survive a several-second qualification; after it passes, people read deeply** |\n\n**This matters because a claim graph is inherently dense.** A permanent one-to-three-sentence rule\nwould cap the product below its own value. The rule that survives the evidence is narrower and\nbetter: **brevity buys the right to depth; it does not replace it.** The founder's instinct is right\nabout the first screen and would be wrong as a global law.\n\n---\n\n## 2. What is Deliberus's \"my books balance\"?\n\nCondition (d) is satisfied by a payoff a stranger already wanted, delivered before any vocabulary is\nlearned. Four candidates, ranked, with the reason each is or is not the demo:\n\n**(a) A fast, honest summary of a debate — AVAILABLE, but not the demo.** Two problems. It is the\ncommodity shape (a language model answers instantly and fluently), and — the sharper one — **it is\nthe exact intervention the DeepMind facilitation study measured as steering**: their *summarizing*\nfacilitator shifted real-money allocations by up to 5.5 percentage points while participants\nconsistently *preferred* it, and their own analysis attributes the steering to restatement that\nnormalises outliers. Leading with the summary leads with the mode that has a measured hazard.\n\n**(b) \"Is this argument any good?\" — the landing hook.** Paste a paragraph, an article, a post, and\nget back its structure and where it is weakest. A felt need with no good tool, no vocabulary required\nto receive the value, and — decisive — it is the **principles-based** mode (asking questions, probing\njustification), which is the arm of the same study that was *not* implicated in steering.\n\n**(c) \"What am I missing?\" — the return hook.** The corpus's own highest-value human act, and by\n*an act's value is not its rung* it needs no ontology fluency at all. A reader can answer it.\n\n**(d) The agent surface — the leverage play, and already designed.** The founder's MCP/CLI instinct\nhas a home: [agents-as-a-consumer-class.md](agents-as-a-consumer-class.md) settles the shape\n(**read-only**, dispute-shaped rather than verdict-returning, instruments as first-class tools) and\nstates the gate (**the strength layer must be sound first** — agent consumption launders provisional\nstructure into authority at machine speed). It also carries the strongest measurement argument in the\nproject: **the agent-side falsifier is the clean one**, because a machine client has no affordance to\nblame, so a null result is unambiguous.\n\n---\n\n## 3. What the games actually teach — five mechanics, each with a home here\n\nThe genre that solved \"navigate a web of claims and enjoy it\" is the deduction game, and its\ndesigners have written down how.\n\n### 3.1 A gap with grammar in it (*The Case of the Golden Idol*)\n\nTheir deduction screen is fill-in-the-blank, and the load-bearing detail is that the blanks are not\nblank. Klavins: *\"The scrolling text is not completely blank — it offers a lot of grammatical and\nsemantic context.\"* A slot that reads `[someone] argues [X] because [____]` is playable; *\"this claim\nhas an unsupported value premise\"* is not.\n\n**Home**: this is *no gap displayed without a move offered* given a concrete form, and it answers the\none standing violation of that principle — the completeness oracle's `unsupported_value_premises`,\nwhich reports what is unsupported and offers nothing.\n\n### 3.2 Curate to the solution plus interesting misdirection\n\nThey cut world-building because *\"loads of tangential information\"* overwhelmed players, adopting\n*\"a very minimalistic approach where we would only fill in the contents that were either necessary to\nthe solution or create interesting misdirections.\"* The graph-visualisation literature agrees from\nthe other side: the hairball comes from showing too much at once, and the standing recommendation is\nto **start from 20–50 relevant nodes** and expand on interaction (focus+context).\n\n**Home**: the default claim view must not be the neighbourhood. Deliberus already computes relevance\n(hinge sensitivity, `worth_asking`); it is not yet used to decide what to *hide*.\n\n### 3.3 Partial progress that does not resolve\n\nTheir fix for players who felt lost was the *\"two or fewer slots are incorrect\"* indicator, giving\n*\"a feeling that they are getting somewhere and were rewarded for figuring things out\"* without\nhanding over the answer.\n\n**Home**: *you have answered 3 of the 5 questions that would move this badge* — progress without a\nverdict. The hinge-ordered question list already computes exactly this and presents it as a list\nrather than as a signal.\n\n### 3.4 Teach through constrained action, with failure impossible at first\n\nPortal's opening room has no danger and no way to fail, standard controls, and one thing to do; each\nsubsequent room teaches exactly one mechanic or one new use of an old one.\n\n**Home**: this is the corpus's own **inversion principle** (*the machine holds the user's hand through\nthe descent — taps, not essays*) stated as level design. The first screen should have **one available\nmove**, and the second should have one more.\n\n### 3.5 Node types must differ at a glance\n\nThe standing criticism of *Slay the Spire*'s map is that encounters are *\"represented by a small\nsymbol with no distinct size or color differences,\"* forcing the player to visually comb each path.\n\n**Home**: Deliberus has four unused visual channels — claim type, badge state, terminus type,\nprovenance (asserted / minted / machine). A reader should be able to see *contested*, *bedrock* and\n*nobody has looked* without reading a word.\n\n---\n\n## 4. The strongest finding, and it is the founder's own sketch\n\nHis decades-old drawing has a conclusion at the top, three argument clusters hanging below — and, at\nthe bottom, **a scrubber**: `speed | dream | cobalt | sweet`, with a hand cursor on it.\n\n**That is not a graph control. It is a chapter selector**, and it is the structure of the two games\nthat solved this: *Return of the Obra Dinn*'s scene scrubber, and Golden Idol's vignettes, of which\nits designers say each *\"feel[s] like a self-contained puzzle with a natural pace\"* because each\ncarries **its own specific question set**.\n\n**Slicing — not zooming, not filtering, not a better layout — is the answer to the hairball.** And\nDeliberus already has the slices: **an argument is a vignette.** So the intuitive renderer is not a\nbetter force-directed graph. It is **one argument on screen at a time, with a scrubber across the\narguments of a source or a debate** — which is what he drew, before the genre that proves it existed.\n\n---\n\n## 5. The discipline that must ride along, or this becomes the thing we distrust\n\nGame feel optimises **engagement**, and this project's threat model contains the finding that\nparticipants *preferred* the facilitator that steered them. The corpus already states the\nconsequence: any friends-round result of the form *\"testers liked it\"* is near-worthless as evidence\nabout deliberative quality.\n\n**So: build for flow, and measure something else.** Time-on-app and return rate are the metrics a\ngame would optimise and are exactly the ones this project cannot trust. The measurable that survives\nis behavioural and already named — *do people point at claims* — plus the instrument suite's own\nreadings (disagreement preservation, ingestion balance) taken across sessions.\n\nSecond caution, held as a hypothesis rather than a finding: the **entrenchment hazard**\n([does-descending-change-the-descender.md](does-descending-change-the-descender.md)) says descent may\nharden the descender on value claims. Making descent *more* rewarding would amplify whichever\ndirection that turns out to run, and the entire instrument suite is blind to it by construction.\n\n---\n\n## 6. Read depth, honestly\n\nSearched and read this session: the attention-span literature (goldfish origin, the 2025 review,\nMark's replication set), the Golden Idol designer interview at Game Developer, Portal onboarding\nanalyses, graph-visualisation hairball and focus+context practice, Slay the Spire map criticism.\n\n**Not researched, and load-bearing if this proceeds**: self-determination theory as applied to games\n(competence / autonomy / relatedness — the return-motivation, which connects to\n[curiosity-as-growth-fuel.md](curiosity-as-growth-fuel.md) and is where the *why come back* answer\nlives); mobile-specific interaction patterns; accessibility of a graph-first interface; and any\nempirical work on whether game framing changes reasoning quality rather than reasoning *volume* —\nwhich is the question §5 says we would actually need to answer.\n\n**Sources**: [Gloria Mark, Attention Span](https://gloriamark.com/attention-span/) ·\n[Steelcase interview, the 47-second finding](https://www.steelcase.com/research/articles/our-47-second-attention-span-with-gloria-mark-s5-ep3-transcript/) ·\n[Northwell, the goldfish myth](https://thewell.northwell.edu/brain-nerve-health/attention-span-goldfish-myth) ·\n[The Myth of the Shrinking Attention Span](https://edspace.american.edu/thecfebeat/2025/01/01/the-myth-of-the-shrinking-attention-span-shed-siliman/) ·\n[Game Developer: pursuing the \"Aha!\" moment with The Case of the Golden Idol](https://www.gamedeveloper.com/design/case-of-the-golden-idol) ·\n[Thinky Games on the Golden Idol developers](https://thinkygames.com/features/how-the-case-of-the-golden-idol-developers-made-one-of-the-decades-best-detective-games-twice/) ·\n[Portal 2 and onboarding](https://medium.com/@mhkt/portal-2-taught-me-everything-i-know-about-onboarding-4e5abf0310c1) ·\n[Cambridge Intelligence: fixing data hairballs](https://cambridge-intelligence.com/blog/hairball-effect-in-graph-visualization/) ·\n[Cambridge Intelligence: graph visualization UX](https://cambridge-intelligence.com/blog/designing-intuitive-data-experiences-with-graph-visualizations/) ·\n[Slay the Spire UX analysis](https://medium.com/@n01578837/final-deliverable-632cfc09e673)\n\n**See also**: [ux-principles.md](../ux-principles.md) (the inversion principle, the engagement\ngradient, P20's structuring gradient) · [fractal-priors-for-success.md](fractal-priors-for-success.md)\n(condition (d)) · [agents-as-a-consumer-class.md](agents-as-a-consumer-class.md) (the MCP shape and\nits gate) · [engagement-gradient-priors.md](engagement-gradient-priors.md) · [sketches.md](../sketches.md)\n(the original drawing).\n\n---\n\n## 7. The correction: the constellation is telling too\n\n*Founder, 2026-08-29, correcting my read of his \"show, don't tell\": he was not describing the\ninvitation copy. He was describing **the whole landing strategy, the new star field included** — still\nexplaining the philosophy, the epistemology and the backdrop to fellow nerds, instead of letting them\nplay, learn by doing, and hit real curiosity or anxiety moments they navigate and contribute through.*\n\n**He is right, and it reaches the constellation.** Twenty-odd conviction tiles are assertions about\nwhat we believe. Compressed beautifully, stranger-tested, but still a wall of claims a visitor reads\nrather than a thing they do. The star field is the *most elegant possible* version of telling.\n\n**The resolution does not cancel the review; it relocates its output.** The constellation's real work\nis **internal**: forcing every conviction through the stranger test and the over-broad test is how\ncondition (d) gets held (§ *the doorstep* above), and that is worth doing whether or not a single\ntile is ever published. What was never established is that the tiles' **destination** is the landing\npage. Two different artifacts:\n\n| | The constellation | The landing |\n|---|---|---|\n| Job | prove each conviction can be said plainly to a stranger | let a stranger *do* something and feel what it is |\n| Audience | us, and eventually a reader who already wants the foundation | someone who arrived with an argument they cannot win |\n| Mode | telling, honestly and briefly | showing |\n| Where it belongs | a page you can reach, an about, a funder packet | the first screen |\n\nSo the star field earns a home — behind door two of the three-doors block, or as the *about* — and\nthe first screen becomes the thing working. **This is a proposal, not a ruling**; what is ruled is\nonly the founder's diagnosis that the current plan is still telling.\n\n**2026-09-19 — the founder placed it, for now.** Needing something to show at a workshop, he put the\nwhole star field on the landing itself, unruled stars included as marked drafts (*\"better to have\nthem present than to leave them out\"*). For a few hours it sat below the three worked examples.\nThen he made it the opening of the page (*\"I want the landing page to open on The constellation,\nhide the bits that come before\"*), with the hero and the examples hidden, never deleted. **That is\nthe reverse of the order this section argues for, and it is his call**: today the wall is what he\nwants a stranger to meet first. The argument above stands as the case for the day the first screen\ncan be the thing working; until then the telling is at least the distilled telling. State and upkeep rule: `specs/landing-redesign/constellation-map.md`\n§ DISPLAY STATE. **And the same evening he put a picture of a map above it** (*\"Could we put an example argument graph (like the animated/rendered one on the claim page, look it up) above the constellation? With just 1-2 elements of each (of the most central/distinctive) type of ontology detail, to illustrate what a quintessential Deliberus argument map/graph would look like (and just tiny tiny subtle hints of what the affordances are perhaps)?\"*), so the first screen is again the thing itself, shown, with the telling second: `web/src/lib/components/ExampleMap.svelte`. It is a hand-made picture with inert prompts, so it shows what a map looks like and is not yet the thing working; that step is still the one this section argues for.\n\n## 8. The cold start, and why it is the demo rather than the problem\n\nThe founder's two candidate answers to a young, sparsely populated graph are **live\nresearch/extraction while the user waits** and **a massive pre-backfill in the style of SciencePedia**\n(the Chinese project, arXiv:2510.26854).\n\n### 8.1 The backfill inherits a verification mechanism that does not exist in our domain\n\nTheir abstract states the filter in their own words: multiple independent solver models generate long\nchains, *\"rigorously filtered by prompt sanitization and cross-model answer consensus, **retaining\nonly those with verifiable endpoints**.\"* Three million first-principles questions over ~200 courses,\n~200,000 articles, and — their own scope note — the system *\"largely omits human-centric\ninformation.\"*\n\n**The reliability rests entirely on the checkable endpoint**, and Deliberus's differentiated domain is\ndefined by not having one. The corpus established this in May and it holds\n(`.private/sciencepedia_lcot_2510.26854_2026_05_16.md`,\n`.private/research_raw_2026_05_17_conceptual.md` § A): their method is *\"parasitic on the checkable\nendpoint — and the authors say so\"*, and the consensus filter has a sharper defect than mere\ninapplicability. Per **Correlated Errors (ICML 2025, 350+ models)**, models *\"agree on the wrong\nanswer far more than they would at random\"*, error correlation **rises with model accuracy** and with\nunder-the-hood convergence across providers — so cross-model consensus is *\"strongest exactly where\nit is least needed and weakest exactly where it matters.\"*\n\n**And a backfill at that scale collides with the pre-launch audit run the day before**\n([the-residual-error-taxonomy.md § 5](the-residual-error-taxonomy.md)): it would be same-hand bias\n(class 1) and confidence laundering (class 8) at three-million scale, and — the irreversible one —\n**frame lock-in (class 7) crystallized by machines before any human touched the vocabulary.** The\naudit's own finding is that frame lock-in is the single class where *waiting* makes things worse; a\nmass backfill makes it worse faster and permanently. Volume is the one input this graph cannot take\nbefore the frame is contestable.\n\n**What DOES transfer, and it is not the pipeline — it is the diagnosis.** Chen Kun's group state it\nalmost verbatim as ours: current sources *\"prioritize conclusions but omit the reasoning chains that\nproduce them\"*, and they call the missing layer **the dark matter of knowledge**. That is the\nmissing-layer thesis arriving independently from a well-funded lab with the Chinese Academy of\nSciences behind it, and their scope note (*omits human-centric information*) is the cleanest available\nstatement of the slot Deliberus claims. Positioning, not method.\n\n### 8.2 Live extraction is not a stopgap. Shown, it is the entire show-don't-tell answer\n\nMeasured on the paid tier: **9.4–15.6 words/s through the full eight-pass pipeline, ~40 claims per\n1,000 words, a fixed floor around 76 s** — so a 1,500-word source is roughly **100–165 seconds**.\n\nHidden behind a spinner that is dead time and a reason to leave. **Shown, it is the demo**, and it is\nthe one thing a language model cannot imitate: an answer arrives finished and opaque, while a map\n**assembles** — arguments appearing, edges forming between them, gaps marked as gaps, the questions\nthat would move each claim listed with what they would move. Watching structure being built out of\nyour own question is learning-by-doing at the doorstep, with no vocabulary demanded first.\n\nThree further reasons it beats the backfill beyond the safety argument: it is **demand-driven**, so\nthe graph grows where people actually contest things, which serves ingestion balance far better than a\nmachine sweep; it is **per-question**, so no frame crystallizes at scale; and the wait is where the\nsystem can honestly show what it does not know, which is the confession principle given a moving\npicture.\n\n### 8.3 On \"we'll just have to trust that the results speak for themselves\"\n\n**You do not have to trust it, and that is the cheaper position.** The claim — that a\nDeliberus-backed answer beats an unaided frontier answer while looking similar to most people — is\nmeasurable, and the measurement is already designed: the **substrate run** (raw sources / flat claims\n/ typed graph, judged by a machine reader that has no affordance to blame, so a null is unambiguous;\n[agents-as-a-consumer-class.md](agents-as-a-consumer-class.md)). It is blocked on quota, not on\nmethod.\n\nTwo reasons to want it early rather than late. If the advantage is real, an unfalsifiable *trust the\nresults* is strictly weaker than a number in a funding application. And if the advantage is **not**\ndistinguishable, that is the single most valuable thing to learn before the corpus grows large enough\nto make re-framing expensive. Note also that *indistinguishable to most people* cannot be resolved by\nasking people whether they liked it: the DeepMind facilitation result has participants preferring the\narm that steered them.\n\n### 8.4 The asymptote is bought by matching, not by volume\n\n*Founder's clarification: as the graph accumulates shared hidden and implicit premises, **less and\nless has to be created anew from extraction**. That is the vision, and it is the cost curve\n([lowering-the-cost.md](lowering-the-cost.md)) stated from the supply side.*\n\n**The vision is right and the machinery for it does not exist yet — which changes what volume is\nworth.** Measured 2026-08-29 (`uv run python scripts/graph_facts.py`): 4,852 claims of which 1,741\nsubstantive, **31 implicit premises across 31 sources**, and **56 `SIMILAR_TO` edges across the whole\nclaim population**. Concept-layer reuse is healthy (87% of concepts reused, measured earlier); claim\nreuse is roughly one percent.\n\n**⚠ Corrected 2026-08-29, and the correction matters.** The first version of this section said no\npath looks up an existing claim. **That was wrong**: `auto_connect` offers `decomposes_into` as one of\nfive classifications with explicit two-direction handling (*could A serve as a PREMISE that B\nlogically depends on?*), and the graph holds **84 cross-source DECOMPOSES_INTO edges** produced that\nway. Termination-in-known-substructure is **not impossible**; it happens *after the fact, at pair\nlevel*. The accurate gap is narrower: `auto_connect` is called only from the two extraction paths and\nis scoped to one source's claims against **other** sources, so a claim minted later by the correction\nUX never passes through it, `decompose_claim` itself mints without looking up, and same-source reuse\nis invisible.\n\n**And the measured surprise is bigger than the gap.** Per-source connection yield, measured\n2026-08-29, does **not** rise with corpus size — it tracks whether the new source has an **adversarial\ncounterpart already present**: the death-penalty pair 0.65 and 0.68 cross-source edges per claim, the\nassisted-dying pair 0.60 and 0.94, minimum wage 0.44, against 0.02 for the Wikipedia\nwisdom-of-crowds piece, 0.04 for Graeber's essay and 0.02 for a lone unified-theory article. So the\nflywheel is real and **paired**, not cumulative, which is direct evidence for the standing\ningest-by-debate-cluster rule and against the naive *it gets better as it grows* reading. (31 sources,\nso suggestive rather than settled — but the pattern is not subtle.)\n\n**Therefore adding volume does not bend the curve; it steepens the scattering problem.** Each new\nextraction mints fresh nodes for propositions the graph may already hold, so at N× the claims with an\nunchanged reuse rate you get N× the identity scattering (residual-error class 3) and the same\nnear-zero reuse. This is the second, independent reason the mass backfill is the wrong instrument for\nthis goal: §8.1 says it imports an inapplicable verification mechanism, and this says **it does not\neven serve the purpose it was proposed for.**\n\n**What actually buys the asymptote** is three pieces the corpus has already named, all designed and\nnone shipped:\n\n1. **A lookup inside decomposition** — before minting a subclaim, search for an existing one and link\n   instead. Nothing else makes a descent *terminate* in known structure.\n2. **Claim-sameness with careful auto-merge** — the founder's Session-25 ruling (anyone types\n   anything; careful merge afterwards; anyone can split). The binding constraint on every reuse claim,\n   and assembly theory's own hardest problem, so difficulty here is principled rather than local.\n3. **Cross-source implicit premises** — the shipped within-argument pass produces bridges *inside* one\n   argument, and its necessity filter deliberately excludes exactly the shared background the\n   asymptote is about. `cross_source_premises.py` produces the right object and is operator-gated and\n   unwired.\n\nOrdering follows from the same logic: **matching first, then volume.** Volume before matching is the\none sequence that makes the eventual matching harder.\n"}