{"path":"research/session21-single-surface-correction-and-the-modality-objection.md","content":"# Session 21: A Correction That Lands in One Surface Is Not a Correction\n\n**Date**: 2026-08-17\n**Type**: docs-truth audit, three principles, one falsifier qualified. Cloud session, no LAN, no secrets, public HTTPS and the repo only.\n\n---\n\n## 1. The finding, which is a mechanism rather than a lapse\n\nThe session began as a recap and turned into an audit when one item exposed the pattern behind it. `TODO.md` was carrying **an item and its own answer as two separate open items**: *\"consume `heterogeneity` and `study_count` in edge strength\"* filed 2026-08-13, and *\"mint the evidence properties as claims\"* filed 2026-08-16, the second of which names the first and dissolves it — heterogeneity weighing nothing is not a missing weighting formula, it is a property that should be a claim.\n\nThat is not carelessness, and the mechanism is worth stating precisely because it recurs:\n\n> **A claim and its refutation live in different files, and nothing compares them.**\n\nEight instances surfaced in the corpus, and every one has the same shape — the correcting knowledge existed, in the repo, near the thing it falsified.\n\n| What one surface said | What another surface already knew |\n|---|---|\n| Consume heterogeneity in edge strength (TODO, Aug 13) | It is a property that should be a claim (TODO, Aug 16) |\n| Stance flattening: \"nothing covers it\" (TODO) | `stance.py` shipped two days later |\n| 1,175 tests (`CLAUDE.md`) | \"1,164 + 1, **not 1,175**\" (TODO) |\n| \"The graph already holds two live Swedish election disagreements\" | `/graph/stats`: 25 sources, none Swedish |\n| Flattening has four mechanisms (frontier) | Translation is a fifth, by policy (run 7) |\n| Always-mint's two safeguards, described as in force | Audit: both satisfied only by accident |\n| \"Two live forks\" (frontier) | Three forks then enumerated, one resolved |\n| Workshop designed on pre-crunched material | P20 forbids it in as many words |\n\nThe last one is the sharpest and became the fourth measured instance of [the uncollected idea](conviction-and-critique.md): **the workshop section cited P20's own lineage research elsewhere in the same file while contradicting P20 itself.** Three days between statement and violation, caught the same day it shipped — by the founder, not by anything in the corpus.\n\n**The commit that best shows the mechanism is `ad8a4fe`.** It corrected the Gemini billing reality — all key variants are one key, account blocked, free-tier quota — and its diff is `CLAUDE.md | 2 ++`. One file. The three TODO items resting on the premise it falsified were never re-trued, including one asking the founder to authorise a secrets-and-deploy swap onto a \"separate personal key\" the same measurement says does not exist as a separate key.\n\n### 1a. Five of the instances were mine, in this session\n\nRecording them because the pattern is the point and an audit that exempts the auditor is worthless.\n\n- **\"One-day interval\"**, asserted in five places for the P20 violation. Never checked. `git log`: P20 landed 2026-08-14, the workshop section 2026-08-17 — **three days**.\n- **\"Nine directives missing from the north-star doc.\"** Committed, then refuted by my own follow-up: **seven of the nine are present in that page's own words.** The grep searched `CLAUDE.md`'s phrasing — *\"One visible mental model per screen\"*, *\"Conservative relevance beats semantic reach\"* — and the page says the same things differently, so the search found nothing and I published nine absences.\n- **A link typed `structure-versive-scale.md` twice**, in two different files, hours apart.\n- **A stray CJK character** inside an English sentence in a new web-served doc.\n- **\"Post-hoc or incremental, never live\"**, published as the modality doc's verdict on the diarized session and challenged within the hour: *\"3x real-time but never live? How come? It's all a matter of how you build this. The sky is the limit. I know your imagination is greater than this.\"* The throughput figure was correct and the inference inverted it — **throughput above real time is the condition streaming needs.** The composed Mac figure is ≈1.3×, the A4000 ≈7×, and enrollment deletes the clustering stage that is the actual reason a cascaded diarizer favors batch.\n\n**Three of the five share one cause with the corpus instances: single-surface verification.** Concluding absence in document B from a grep written in document A's vocabulary is the same error as correcting file A and leaving file B, run backwards.\n\n**The liveness claim is a variant worth naming on its own, because it is subtler and I would not have caught it by re-reading my own text.** The disproving material sat in the file I had just cited, in the same session, as the source for the number I was reasoning from. That dotfiles research carries a section on *streaming* diarization and a diarization throughput figure that is also above unity — so the citation was used for the fact that supported the conclusion and not read for the facts that refute it. Call it **citing selectively within a source you have open**: it passes every check aimed at unsourced claims, because the claim is sourced. What it fails is the question of whether the source agrees with the conclusion drawn from it.\n\n### 1b. The buildable consequence, and it has a precedent in this repo\n\nFive of the eight corpus instances are **a number or a named artifact stated in two places**: 1,175 against 1,164, four mechanisms against five, two forks against three enumerated, 25 sources against \"already holds\", `count_safe_summary` has-no-caller. Those are mechanically checkable, and this repo already contains the pattern for making a prose promise bite — `tests/test_generate_frontier.py`, whose own comment records why it exists: *\"The prose promise became a mechanism, because this rule had been in CLAUDE.md for months and was violated the entire time.\"*\n\n**So the proposal is to extend that pattern to cross-surface numeric claims** rather than to try natural-language contradiction detection: a test that extracts stated counts and existence claims from `CLAUDE.md`, `TODO.md` and the frontier and compares them against each other and, where reachable, against the live source. It would have caught five of eight. It would not have caught the P20 violation, which is semantic — and saying so is part of the proposal, because a check advertised as catching everything is the failure mode it is meant to prevent.\n\nThe remaining three want the other half of the same idea: **an idea with a dependent breaks loudly.** P20 had none, which is why nothing noticed.\n\n---\n\n## 2. The founder dissolved two questions rather than answering them\n\nTwice in one session, options were offered inside a frame and the frame was rejected. Worth naming as a method observation, because both times the corpus turned out to already hold his answer.\n\n### 2a. The workshop format\n\nPresented with a false premise corrected and two options — ingest the Swedish material, or design on an adversarial pair already in the graph — he rejected what they shared:\n\n> *\"I dunno if we really need a pre-crunched material to start from in the case of Nils' workshops, why don't we just have participants USE Deliberus from scratch and submit stuff and interact with Deliberus live? Which is a process that really could use some UX love anyway. This will force it into existence, force us to work on that, to test it much more ourselves before the workshop, perhaps just with a friend, would be quite interesting to sit side by side with laptops and 'discuss through Deliberus'.\"*\n\nBoth options were pre-crunched. A session on material the software processed before anyone arrived tests the surfaces that display a graph and never the funnel that produces one. **P20 had said exactly this three days earlier** — *\"the guided path cannot be a read-only tour of an existing corpus\"* — and the load-bearing argument was also already written: pre-crunched material makes the project's own primary falsifier unobservable, since people only point at claims grown from their own words.\n\nThe phrase **\"discuss through Deliberus\"** named something the corpus had no principle for. Searched: `side-by-side`, `synchronous`, `co-present` returned nothing across the whole UX corpus. Everything documented was solo or asynchronous. That became **P21**.\n\n### 2b. The falsifier's fixed-modality assumption\n\nTold that *whether people point at claims* is the one signal that would redirect the build:\n\n> *\"Oooh this is very important. But it just inspires me to enable alternative UX modes/paths or whatever. We can and should adapt to whatever use Deliberus could conceivably transmogrify into. Considering dev velocity nowadays, this is the least of our worries… UI flows/views can be designed to adapt perfectly to different use cases and contexts, if we convince ourselves the effort is warranted.\"*\n\nThe objection holds. *\"With claim-level tools in hand\"* is not the same as *well afforded*, so a null is under-determined across three branches rather than one. **And it carries a danger worth having stated back**: an unbounded modality set means every null gets answered with *we had not tried the right interface yet*, which is precisely the defect that got an earlier reformulation of the architecture bet retracted. The resolution borrows his own terminus vocabulary — the verdict is *addressability did not pay under the modalities tried*, the set is pre-registered before a session, and expansion is logged so attempts stay countable. Full treatment: [interaction-modalities-and-the-pointing-test.md](interaction-modalities-and-the-pointing-test.md).\n\nHis strongest concrete idea in the same message — recording a two-person conversation and running it through the Sarpetorp diarization pipeline — **dissolves two of P21's three hardest blockers rather than mitigating them**, because nobody types. Verified rather than recalled: that pipeline runs KB-Whisper-Large, pyannote 3.1, a 256-dimension wespeaker embedder and eight enrolled profiles, with profile-aware assignment live since June. Deliberus has zero diarization in code or docs. **The gap is attribution, not transcription** — and a diarized session would give the corpus its first per-claim attribution to a live participant rather than a document byline.\n\n---\n\n## 3. How he asked the audit to be conducted, which is itself durable\n\n> *\"But vet and verify each one at a time as a proposal to me, and read up deeply for each. Search mannaminne, docs, git history etc.\"*\n\nand\n\n> *\"I don't want this to be halfassed so ground extremely deeply.\"*\n\nThis changed the work materially. The instinct on finding eight inconsistencies is to fix them in one sweep; **one-at-a-time-as-a-proposal is what surfaced that three of them were not repairs at all but founder decisions in disguise** — the workshop format, the badge boundary, the modality set. A batch fix would have silently decided all three.\n\nIt also has a boundary that had to be worked out rather than assumed, and it held for the rest of the session: **a factual lie gets fixed and reported; anything that changes an item's intent gets proposed.** The test count, the flattening count and the always-mint accuracy note were repairs. Retiring the Swedish-material item was a proposal.\n\nOne honest limit on the \"search mannaminne\" half: **`mannaminne` is unreachable from a cloud session** — it is a Mac-bound CLI against `darwin.home:5440` and there is no LAN. Said rather than skipped silently. The available arms were full git history, the whole corpus, the code, and the live public API, which was sufficient for this class because the strongest verifier is *does the named artifact exist*.\n\n---\n\n## 4. What shipped\n\nThree principles, on approval: **P21** co-present use, **P22** absence is a state not a score, **P23** every displayed claim is addressable. **P7a** promotes an Apr-8 note that held four principle-grade rules and was invisible to anyone scanning the heading list — the findability defect that let a later doc contradict a principle the page already held.\n\nP22 is the best-grounded item in the corpus: the founder's 2026 naming of *structural sycophancy*, plus his own 2009–2013 hand-drawing specifying *\"none addressed (orange)\"* as a distinct state, plus the measurement that 115 of 116 structured claims displayed confidence with nothing answered for five months.\n\nThe frontier was pruned rather than only appended to, on his instruction to cut what is no longer at the edge: the stale flattening count, the always-mint safeguards described as in force, and a resolved fork still listed as open whose removal also restored the section's own two-forks count. 2,891 of 3,000 words, contract green.\n\nAlso corrected: the Cloud-Session Notice claimed `.private/` is unreachable as verified fact. It **varies by launch** — one session the same day found it at a sibling path, this one found it absent at both — so an open handoff item asking to *drop* the standing-rule consequences would have been the same over-generalisation pointing the other way. The notice now prescribes `ls -d .private ../.private` at session start.\n\n---\n\n## 5. Open, and named rather than left implicit\n\nThe **three Gemini items** (`TODO.md`) still rest on the premise `ad8a4fe` falsified, and one of them requests a secrets-and-deploy action. This is now the first blocker on the live-funnel critical path rather than housekeeping, because a live co-use session is probably not runnable on free-tier quota at all.\n\nThe **pre-registration** of session one's modality set is this session's one dependent. If it does not happen before the first rehearsal, the modality doc has joined the four instances of the uncollected idea rather than fixing them. The liveness correction adds a second thing to pre-register with it: **which of the three tiers the session runs**, because tier 1 buys immediacy and takes on the steering risk while tier 3 avoids it and loses the immediacy. That is a trade to declare in advance, not a preference to discover afterwards.\n\nThe **cross-surface numeric check** in § 1b is proposed, not built.\n\n---\n\n## Cross-References\n\n[interaction-modalities-and-the-pointing-test.md](interaction-modalities-and-the-pointing-test.md) (the falsifier qualification) · [islands-of-coherence.md](islands-of-coherence.md) § 5c (the format decision) · [structure-versus-scale.md](structure-versus-scale.md) § The signal that would redirect the build (the third branch) · [../ux-principles.md](../ux-principles.md) P7a, P20–P23 · [conviction-and-critique.md](conviction-and-critique.md) (the fourth uncollected-idea instance) · [structuring-gradient-lineage.md](structuring-gradient-lineage.md) (P20's dependent, finally supplied) · [strength-layer-audit.md](strength-layer-audit.md) § 2b (P22's provenance) · [long-form-sources-and-meta-analysis-weighing.md](long-form-sources-and-meta-analysis-weighing.md) (the superseded heterogeneity recommendation)\n"}