{"path":"research/weighing-detection-semantic-tier.md","content":"# Should Weighing Detection Stay Deterministic? The Semantic Tier's Moment\n\n**Date**: 2026-08-23 · **Status**: research for a live founder question — NOT ratified. The question, near-verbatim: *\"Should the Swedish weighing lexicon even be deterministic? Should it not be LLM-based for much better judgment on that aspect?\"* Asked while planning the live election test, whose blocker table had scheduled a half-day deterministic Swedish lexicon build.\n\n## 0. Plain summary\n\nThe founder's instinct converges with what the corpus already concluded twice in its own words: the semantic (LLM) tier is \"the planned next step\" for weighing detection (dogfood run 2, G10), and the stance instrument's sibling limitation \"wants a semantic tier, not more strings.\" The recorded reasons for shipping deterministic v1 were sound and have since shifted: the silent-LLM-failure class it avoided is now a managed discipline rather than something only abstinence protects against, and the \"misses cost only an un-flagged invitation\" comfort was priced before the miss *rate* was measured — three of four English registers undetectable, zero of nine instrument firings on Swedish. A hand-built Swedish lexicon would repeat the sacredness brake's measured failure shape (introspection-built patterns, shipped inert against a corpus with zero examples to validate on). A live demonstration run for this document found **four of four weighings in a synthetic Swedish election paragraph, with dialect and mechanism** — including an idiom and a metaphor-carried implicit weighing that no lexicon can reach by construction; a hand-built lexicon would plausibly have caught one. What determinism still buys is real and named below, which is why the design space ends in a two-tier shape rather than a replacement — with the routing of LLM-detected weighings marked as a fresh founder decision, because the Jul-7 auto-open ratification was made under a deterministic detector.\n\n## 1. What the corpus already holds (read before answering, per house rule)\n\n**The v1 rationale, recorded in `deliberus/weighing.py`'s own docstring, verbatim:**\n\n> \"Detection is DELIBERATELY deterministic in v1 (lexical markers + claim type): no extraction-prompt changes, no new LLM pass — the Apr-1 outage class stays untouched, and misses cost only an un-flagged invitation (a human can open any claim explicitly).\"\n\n**The semantic tier was already the plan.** G10 (dogfood run 2), after ground-truthing zero detections against two rights-advocacy sources:\n\n> \"This is the empirical case for the planned semantic tier: when a text yields value premises but zero lexical weighings, an LLM pass should ask whether an implicit weighing structures the argument.\"\n\nAnd the stance instrument's logged limitation uses the same words about its own reported-speech signal: \"the axis wants a semantic tier, not more strings.\"\n\n**The register map** (measured, three runs): moral philosophy says \"outweigh\" (detected), policy analysis says \"net benefits\" (detected after the G1 lexicon extension), rights advocacy says neither — the value conflict carried entirely implicitly, \"undetectable lexically by construction.\" Run 3F added two further dialects (rank-ordering, superlative). Run 7 added the language axis: **0 of 9 instrument firings on Swedish against 9 of 9 on English.**\n\n**The determinism boundary is already porous.** Detection gates on `_WEIGHABLE_TYPES = {\"value_premise\", \"normative\"}` — and claim type is assigned by the LLM in Pass 2b. The \"deterministic\" detector has always consumed an upstream LLM judgment; the question is whether the *second* judgment (is this a weighing?) also deserves one.\n\n**The ratified action policy** (founder, Jul 7): provenance split — source texts auto-open the CQ1–8 descent, authored input gets an invitation, the sacredness brake downgrades any weighing to invitation. That ratification governs *what a detection does*; it was made when *what detects* was a regex.\n\n## 2. Why the recorded v1 rationale has shifted\n\n1. **\"The Apr-1 outage class stays untouched.\"** In July 2026 that class was avoided by not adding LLM passes. Since then it became a managed discipline: every pass routes through `llm_call` (no silent excepts, `logger.exception` mandatory), the Literal-string regression is pinned by tests, and a new pass inherits all of it. Abstinence is no longer the only protection — and under the Claude backend option the pass would not even compete for the starved quota.\n2. **\"Misses cost only an un-flagged invitation.\"** True per miss, and priced before the miss *rate* was measured. At one detectable register of four and zero of one language, the gentle-cost framing describes a tail event; the measured reality is an instrument dark in most of its domain. A cost argument calibrated on the tail does not govern the norm.\n3. **The Swedish-specific trap.** A Swedish lexicon would be introspection-built — the corpus holds zero Swedish weighing examples to validate against. That is exactly how the sacredness brake shipped: patterns from intuition, matched **zero of 1,678 claims**, inert not miscalibrated. Building lexicon-first for Swedish repeats a mistake this project has already paid to learn.\n4. **The quota context dissolved.** Whatever weight \"no new LLM pass\" carried as rationing discipline is removed by the subscription backend ([claude-code-as-extraction-engine.md](claude-code-as-extraction-engine.md)); batched, detection is one additional call per extraction.\n\n## 3. What determinism still buys (stated fairly — this is why the answer is a shape, not a swap)\n\n- **A decay-free, entropy-class detector.** A regex built once keeps working, costs nothing, cannot be prompt-injected, and produces identical firings on identical input — which keeps counts countable across runs, the property the pre-registration discipline leans on.\n- **Legible blindness, ex ante.** A lexicon's misses are enumerable before running it (no pattern for \"går före\" means \"går före\" will miss). An LLM tier's blindness cannot be enumerated in advance.\n- **But the confession asymmetry runs the other way ex post** — and this is the decisive point in house terms: a lexicon's *silence is ambiguous between absence and blindness*. G10 needed a human ground-truth pass to establish that the advocacy texts truly contained no lexical markers. An LLM tier can distinguish its own states — \"no weighing present,\" \"implicit weighing, low confidence,\" \"weighing in an unlisted dialect: ⟨named⟩\" — which is the confession principle applied to detection. A detector that cannot say \"I can't see here\" reports absence it has no license to report.\n\n## 4. The live demonstration (2026-08-23, cloud session, one call)\n\nA synthetic Swedish election-register paragraph — *\"Skolan är viktigare än skattesänkningar i årets val. Tryggheten måste gå före allt annat. Partiet lovar 5000 nya poliser till 2030, och det är värt varje krona. Vi har inte råd att vänta.\"* — was run through a Sonnet-class model in the structured-output harness the backend uses, asked for atomic claims plus weighings with span, dialect, and reasoning. Result, in full:\n\n| Span | Model's dialect label | Could a hand-built lexicon have caught it? |\n|---|---|---|\n| \"Skolan är viktigare än skattesänkningar\" | explicit comparative ranking (\"X är viktigare än Y\") | **Yes** — the one phrase every Swedish lexicon draft would contain |\n| \"Tryggheten måste gå före allt annat\" | absolute priority claim — noted as *stronger* than pairwise, \"excludes all competing considerations in advance\" | Maybe — \"gå före\" is guessable, the absolutist reading is not |\n| \"det är värt varje krona\" | idiomatic cost-benefit (\"worth every krona\") | **No** — a lexicalized idiom; enumeration would never end |\n| \"Vi har inte råd att vänta\" | urgency-over-deliberation, carried by economic metaphor | **No** — the implicit register G10 proved lexically undetectable by construction |\n\nAll five claims were extracted and correctly typed alongside (the promise correctly `empirical`, \"värt varje krona\" correctly split out as its own `value_premise`). Score against the founder's question: the LLM tier found 4/4 with mechanism explanations; a plausible hand-built lexicon catches 1, maybe 2 — and the two it misses include precisely the dialect classes (idiom, metaphor-borne implicit) that three dogfood runs established as the real distribution. One honest wrinkle logged: the model's Swedish *reasoning prose* twice conjugated \"vägar\" for \"väger\" — detection judgment sound, but any display-layer use of LLM-written Swedish still passes through the founder's language bar.\n\n## 5. The design space (founder decides; the lean is stated, not assumed)\n\n- **A. Extend the lexicon to Swedish.** Half a day, zero validation corpus, ceiling known in advance: repeats the brake's inert-ship pattern and reaches at most the explicit-comparative register. The measured case against building this *first* is §2.3 and §4.\n- **B. G10's conditional semantic tier** — the corpus's own designed shape: the LLM pass fires when a text yields value premises and the lexicon found nothing, asking whether an implicit weighing structures the argument. Cheap (conditional, batched: one call), targeted at the measured gap, leaves the lexicon primary where it works.\n- **C. LLM-primary detection** — the founder's question as asked: the LLM judges every extraction's claims (one batched call); the lexicon demotes to a fast-lane/baseline.\n- **D. Two-tier shadow, then distill.** Run both tiers, log agreements and divergences, keep the lexicon's firings as the measured baseline; after enough labeled data, *distill* a validated lexicon from LLM labels if a deterministic tier is still wanted — built from evidence rather than introspection, curing §2.3 permanently.\n\n**The observation that softens the fork: for Swedish, B and C are the same tier in practice** — the Swedish lexicon fires on nothing (0/9), so B's condition is always met and the LLM pass is the detector either way. Building B delivers full Swedish coverage for the election test *without* deciding whether to demote the English lexicon; the B-vs-C question for English can wait for D's shadow data.\n\n**The routing decision this opens (fresh, not covered by Jul 7):** the auto-open-on-source ratification was made under a deterministic, deliberately high-precision detector. Does an *LLM-detected* weighing auto-open the descent on source provenance, or propose/invite until its precision is measured? The conservative reading of \"propose, never assert\" plus the daemon design's co-stimulation instinct: **lexical hits keep today's ratified actions; LLM-only hits invite rather than auto-open in v1**, each carrying its span, dialect, confidence, and published reasoning (the terminus-classifier pattern) — revisited once the election test yields precision data. This is a founder call either way; both options are safe against the graph (a wrong invitation is declinable; the brake still downgrades sanctity-marked cases).\n\n## 6. The design gates, answered\n\n- **Does this flatten contestation?** Detection returns a span and a judgment; it rewrites nothing, so the paraphrase-flattening surface is near zero. The residual risk is *over-generalizing weighing-hood* — reading every priority-flavored sentence as a weighing and burying claims in invitations. Answer: propose-tier routing (§5), per-firing published reasoning, and precision measured against ratified/declined invitations.\n- **Prompt injection.** Detection reads source text — untrusted input — through an LLM. The ingestion-time prompt-injection guard's deterministic tier already screens the same text upstream; the detection prompt gets the same treat-as-data framing, and a detection can propose at most an invitation, which bounds the blast radius.\n- **Adversary class.** Today: none met (all sources machine- or operator-extracted); the detector faces entropy only, where publishing it is free. The strategy-class note for later: a *published* LLM detector is promptable-around by a motivated author in a way a secret one is not — the same publishing-arms-the-adversary trade every instrument here carries, no worse.\n- **Failure discipline.** The pass runs through `llm_call` (Apr-1 rules inherited); a detection failure degrades to \"no invitation flagged\" — exactly v1's miss behavior — and is logged, never silent.\n\n## 7. What the election test yields under the two-tier shape\n\nRun both tiers in shadow through the session (and the post-hoc M-spoken extraction): every firing logged with tier, span, dialect, confidence. The session then produces, as a by-product, **the corpus's first Swedish weighing dataset** — ground truth for calibrating the LLM tier's precision, for the routing decision (§5), and for distilling a validated Swedish lexicon later if wanted (D). The weighing descent that follows any ratified detection is unchanged — the ratified CQ1–8 apparatus, the provenance split, and the brake all sit downstream and untouched.\n\n## 8. Transfer (the corpus already asked for it)\n\nThe same analysis governs the two sibling lexicons, in the corpus's own words: `stance.py`'s reported-speech signal (\"the axis wants a semantic tier, not more strings\" — it misses run 6's \"took the contrary position\" today) and the sacredness brake's sanctity register (rebuilt once from measurement already; its language axis is equally unswept). Neither is built here; both inherit whatever routing rule the founder sets in §5.\n\n## Provenance\n\nThe live demonstration ran 2026-08-23 in a cloud session through the structured-output harness described in [claude-code-as-extraction-engine.md](claude-code-as-extraction-engine.md) §7 (which carries the harness verification record); the model output is quoted verbatim above. Read fully for this document, per the read-the-origin rule: `deliberus/weighing.py` (docstring, patterns, `detect_weighing`, `_WEIGHABLE_TYPES`), `docs/research/dogfood-run-2-orthogonal-experiments.md` G10, `docs/research/weighing-scheme-and-terminus-classifier-design.md` (trigger-policy resolution and addenda), the stance-conflicts limitation notes, and run 7's instrument measurements. The founder's question is quoted from the 2026-08-23 session; nothing here is ratified.\n"}