{"path":"research/live-election-test-design.md","content":"# The Live Election Test: Two People, Real Topics, Pre-Registered Modalities\n\n**Date**: 2026-08-22 · **Status**: design, awaiting the quota decision and a participant pick — nothing here has run. · **Founder intent, verbatim**: *\"I wanna do a live test where two people debate current election topics and try new modalities of Deliberus also.\"* · **Governing priors**: the co-present ruling (participants use Deliberus live, from scratch — [islands-of-coherence.md](islands-of-coherence.md) §5c), the pre-registered modality set (expansion logged so attempts stay countable), the matching rule (real, live, sayable disagreements — never meta-community material), and the DeepMind steering result, sharpened below from the paper's full text.\n\n## Why now is structurally right\n\nThe riksdagsval is **September 13, 2026** — the material's own expiry date, not an artificial deadline. For the next three weeks election disagreements are maximally live and sayable; the corpus holds zero Swedish and zero electoral sources, so everything captured is new register; and election claims are naturally **event-bound** (\"partiet lovar…\" expires on election day), making this the first real material for the valid-while condition class the temporal rung's watcher design needs.\n\n## The pre-registered modality set\n\n| Modality | What happens | What it uniquely tests |\n|---|---|---|\n| **M-typed** (side-by-side, typed) | two laptops, both signed in; each writes their actual position through the landing input; then they explore *each other's* graphs and argue through the correction tools — decompose, add evidence, answer critical questions, vote, retarget an attack onto the part it means | the pointing signal (P21's primary observable), the contribution funnel from scratch, concurrent two-user writing |\n| **M-spoken** (diarized, post-hoc) | ten minutes of ordinary spoken debate recorded FIRST, before any screen; the Sarpetorp voice pipeline (KB-Whisper + pyannote + enrolled profiles) transcribes with speaker attribution; each person's segments enter as authored extractions | the graph's first claims attributed to live named participants; and **the same debate mapped twice** — typed versus spoken — a comparison nobody has data on |\n| **M-live** (turn-boundary, only if energy allows) | one exchange where structure appears on a shared screen mid-conversation | the facilitation zone — run last, under the restatement rule below |\n\nExpansion of this set during the session is allowed and **logged**, never silent — the terminus discipline applied to modalities.\n\n## Pre-registered signals (and the dead proxy)\n\n*Did they like it* is dead as evidence (the DeepMind preference-steering dissociation). What counts:\n\n1. **Pointing** — do they reference each other's claims as objects? (With the session-one-null caveat pre-attached: a null is under-determined across granularity / affordance / modality, per the pending Q2–Q4 rulings.)\n2. **The implicit crux** — run 3F's stable finding says each side's crux is stated by one and unwritten by the other; does live use surface it, and who says it first — a person or the descent?\n3. **The outcome arm** — does structure change what either person *does* next (a position revised, an evidence hunt started, a bet named)?\n4. **Instruments after**: disagreement preservation on the pair, stance conflicts, the weighing descent's state-deficit and satisfier-belief screens on their hottest weighting fight (their first non-synthetic use).\n\n## The DeepMind result, sharpened from the full text (2026-08-22 fetch)\n\nThe facilitator in arXiv:2605.14097 was a chat participant, intervening about once a minute. Study 2 compared **two strategies**: a *summarizing* facilitator (recaps, restating each person's proposals) and a *principles-based* one (active questioning targeting deliberation failure modes — premature consensus, weak justification). Consensus — Krippendorff's alpha over the members' allocation vectors, already near ceiling — improved under neither; participants preferred both; allocations still moved up to 5.5 percentage points. **The authors implicate summarization specifically**: restating participant-proposed splits *\"can turn outlier proposals into salient numeric anchors or increase their perceived legitimacy.\"* So restatement distorts in BOTH directions — it can flatten an outlier toward the center *or* amplify one into the anchor everyone negotiates around. The principles-based facilitator — the one that behaves like a critical-question descent — was their better arm and **not** the steering culprit.\n\n**The design rule this yields for M-live**: *the screen may show structure and ask questions; it never restates what either person said in its own words.* The one place restatement necessarily re-enters is extraction itself (a claim is a paraphrase of an utterance) — which is why disagreement preservation runs on the session's output, and why participants seeing their own claims get the ratification move the study's subjects never had: \"that's not what I said.\"\n\n## Blockers, with measured states (2026-08-22)\n\n| Blocker | Measured state | Action |\n|---|---|---|\n| **Gemini quota** | all key variants resolve to one blocked-account key at free tier (~20 calls/day/model shared); a session needs 10–25+ extraction calls | **the gating founder decision** — pay the bill / fresh billing account / accept free-tier brittleness. Note: the corpus holds two CONFLICTING key-identity measurements (TODO § key conflict); settle before deciding |\n| Rate limits | **softer than previously documented**: 3–10/minute on the LLM-invoking endpoints, not 1/minute | fine for two people; stale \"1/min\" corrected in the corpus |\n| Swedish weighing detection | the weighing/sanctity/reported lexicons are English-only — \"väger tyngre än\", \"viktigare än\" open nothing | half-day build + tests, pre-session |\n| Concurrent contribution | never exercised (two Users exist, no simultaneous writes ever) | two-browser dry-run by the operator before the session |\n| Extraction latency | 2–3 minutes per source | design around, never wait on screen: submit, keep talking, return to the map |\n| Badge boundary | fresh disagreements render as confident blue at 0.500 | the pending display-vocabulary decision — ideally ruled before participants see it |\n| Auth | Google sign-in per participant | acceptable; known at invite time |\n\n## Run of show (~90–120 min)\n\n1. **Baseline debate, spoken, recorded** (10 min, no screens) — the M-spoken material and the unstructured control sample.\n2. **M-typed** (40–50 min): positions in, cross-exploration, correction tools, votes. Observer notes pointing events with timestamps.\n3. **The weighing descent** on the sharpest weighting clash (10–15 min) — first live human run of the ten-question descent.\n4. **M-live, optional** (10 min): one exchange, turn-boundary, under the restatement rule, display log kept.\n5. **Debrief as authored input** — each participant's reflection goes in through the landing input, not a survey.\n6. **Post-session**: M-spoken extraction + diarization; instruments; typed-vs-spoken comparison; run doc (run 8) with the pre-registered signals scored.\n\nParticipant choice, topic choice within their real disagreement, and the quota decision are the founder's. Named-person planning stays in `.private/` per the born-private rule.\n\n## Cross-references\n\n[islands-of-coherence.md](islands-of-coherence.md) §5 (the workshop frame and the co-present ruling this instantiates) · [session21-single-surface-correction-and-the-modality-objection.md](session21-single-surface-correction-and-the-modality-objection.md) (the diarized modality's feasibility tiers and the pre-registered-set discipline) · [dogfood-run-2-orthogonal-experiments.md](dogfood-run-2-orthogonal-experiments.md) (the guided-path precedent) · [the-scrutiny-gap.md](the-scrutiny-gap.md) + the Standing Epistemic Threat Model (the steering stance M-live runs under) · [supersession-and-bitemporal-lifecycles.md](supersession-and-bitemporal-lifecycles.md) (event-bound validity — election claims as the watcher's first real class) · [live-test-simulation-vardforsakringen.md](live-test-simulation-vardforsakringen.md) (simulering två: gräddfilsfrågan, research-grundad)\n\n## Inbjudningstexten (svenska — utkast 2026-08-22, godkänt av founder, tandrad vässad på begäran)\n\n*Registret följer introduktionssidan; tandraden är landningsarbetets varningsregisterrad (\"Reasoning you don't structure gets summarized by someone else\"), här vässad snarare än mildrad, per teeth-and-confidence-linjen.*\n\n> Några veckor före valet sätter sig två personer ned som tycker olika i en fråga de båda bryr sig om — kärnkraften, skolan, tryggheten. Först pratar de helt vanligt, tio minuter, som vid vilket köksbord som helst. Sedan bjuder de in ett tredje sällskap: var och en lägger fram sin ståndpunkt i Deliberus, och samtalet fortsätter genom kartan som växer fram på skärmen. Varje påstående, varje skäl, varje belägg blir en synlig punkt som båda kan peka på, ifrågasätta och bygga vidare på.\n>\n> Varför då? För att politiska gräl går i cirklar av ett skäl som sällan syns: kärnan i oenigheten sägs nästan aldrig högt. Den ena tar något för givet som den andra aldrig har hört uttalas, och så pratar man förbi varandra i två timmar. På en karta blir just det synligt. Man slutar gräla om ytan och blir oense om rätt sak — och förvånansvärt ofta visar det sig att den verkliga oenigheten är mindre än grälet kändes. Några få vägval som går att namnge, i stället för en mur.\n>\n> Det som skiljer detta från alla goda samtal vi redan har är att ett samtal avdunstar. Här blir resonemanget kvar. Nästa person börjar inte om från noll utan fortsätter där ni slutade. De bästa skälen för och emot samlas på ett ställe, prövas bit för bit, och ingen — inte ens systemet självt — får hävda något utan att det går att syna. Resonemang du inte själv lägger fram blir återberättade av någon annan — i den version som passar dem, inte dig. Här ligger dina skäl kvar med dina egna ord, och den som vill angripa dem måste angripa vad du faktiskt sa.\n>\n> Den långa visionen är en gemensam karta över mänskligt tänkande — där det blir billigare för var och en av oss att förstå varför en annan människa tycker som hon gör. Inte för att alla ska tycka lika, utan för att den oenighet som återstår när missförstånden röjts undan är mindre, ärligare och går att leva med respektfullt. Det är vad en demokrati behöver mer av, inte minst nu. Och det är därför valet är rätt tillfälle att börja: frågorna är levande, och vi bryr oss på riktigt.\n\nTandradens vässning knyter medvetet an till misrepresentation-incitamentet (L1 i incitamentsanalysen): den kalla meningen namnger motpartens intresse, och boten blir en rustning — *den som vill angripa dina skäl måste angripa vad du faktiskt sa* — vilket är alltid-mynta-beslutets halmgubbe-detektor uttryckt som löfte till läsaren.\n"}