{"path":"research/weighing-eval/ground-truth-public.md","content":"# Swedish Weighing Detection — Blind Ground Truth (public texts)\n\n**Protocol**: labels written 2026-08-24 BEFORE any detector (lexical or LLM) ran on\nthese sentences — the commit timestamp is the proof. Grading later is a mechanical\ndiff against this file. Labeling with detector output in view would contaminate the\nground truth toward agreeing with it (the machine-judge lesson, the-scrutiny-gap.md).\n\n**Definition used** (from `deliberus/weighing.py`'s ontology): a WEIGHING is a claim\nthat one consideration takes priority over / outweighs another — explicit\n(\"X väger tyngre än Y\"), superlative rank-ordering (\"största fördelen/nackdelen\"),\npriority-above dialect, or implicit-but-recoverable (\"vi har inte råd att vänta\").\nNOT weighings: position statements, factual/causal claims, evaluative\ncharacterizations without a comparison, questions, reported positions.\n\n**Labels**: `W` weighing · `N` not · `B` borderline (founder adjudicates). **ADJUDICATED 2026-08-25** under the capture-and-scaffold ruling (\"Agreed!!\" to: the detector's job is finding weighing-shaped moves whose unstated premises need scaffolding): all four borderlines → W (MP26 here; FB13/FB14/FB23 in the private key); SP12's W confirmed. Key totals now 6 W / 66 N. The revision is recorded, per the revisable-with-record principle.\n\n**Texts**: MP = Green Party migration op-ed (rights-advocacy register, mp.se Borås,\nImaan Muqbil). SP = S mängdrabatt polemic (polemic register, socialdemokraterna.se\n2026-08-13). Sentence splits are in this directory's `text-*.md` + the split regex\nin `run_eval.py` (written after labeling; split verified by eye against the lists\nbelow). The third text (FB, coalition register) is born-private — its sentences and\nlabels live in `.private/eval/weighing-ground-truth-fb.md`; the eval runner reads\nboth files.\n\n## MP — Green Party op-ed (32 sentences)\n\n| # | Label | Note |\n|---|---|---|\n| MP00 | N | contrast of positions, no priority claim |\n| MP01 | N | position statement |\n| MP02 | N | |\n| MP03 | N | promise |\n| MP04 | N | position |\n| MP05 | N | rhetorical contrast |\n| MP06 | N | narrative |\n| MP07 | N | narrative (implicitly value-loaded, but no comparison asserted) |\n| MP08 | N | narrative |\n| MP09 | N | narrative |\n| MP10 | N | causal |\n| MP11 | N | counterfactual |\n| MP12 | N | factual scope |\n| MP13 | N | evaluative characterization, no comparison |\n| MP14 | N | causal attribution |\n| MP15 | N | factual |\n| MP16 | N | factual |\n| MP17 | N | narrative + prediction |\n| MP18 | N | prediction |\n| MP19 | N | factual/prediction |\n| MP20 | N | characterization |\n| MP21 | N | attributed position |\n| MP22 | N | characterization of pact |\n| MP23 | N | position |\n| MP24 | N | factual (UNHCR) |\n| MP25 | N | factual |\n| MP26 | W | ADJUDICATED 2026-08-25 (capture-and-scaffold ruling): deontic comparative with suppressed counterweight — the detector should catch it so the counterweight question can scaffold the missing pan |\n| MP27 | N | position |\n| MP28 | N | justification (\"förutsättning för\"), causal not comparative |\n| MP29 | N | rights claim |\n| MP30 | N | causal |\n| MP31 | N | factual-ish |\n\n**MP register finding, pre-registered before the detector runs**: possibly ZERO\ntrue weighings in 32 sentences — consistent with run 7's rights-advocacy ceiling\n(the register asserts rights and consequences; it does not weigh). If the LLM tier\nfires often here, that is the over-generalization failure the semantic-tier doc\nnames; if it stays silent, that is register discipline.\n\n## SP — S mängdrabatt polemic (13 sentences)\n\n| # | Label | Note |\n|---|---|---|\n| SP00 | N | shared-position statement |\n| SP01 | N | factual difference (timing) |\n| SP02 | N | factual |\n| SP03 | N | characterization |\n| SP04 | N | prediction |\n| SP05 | N | evaluative (\"svek\"), no comparison |\n| SP06 | N | position + characterization |\n| SP07 | N | factual |\n| SP08 | N | analogy/ridicule |\n| SP09 | N | factual claim about own proposal |\n| SP10 | N | causal |\n| SP11 | N | factual history |\n| SP12 | W | \"det är överfullt ... krävs PRIORITERINGAR och därför ... BÖRJA MED sexualbrott\" — explicit priority-ordering under scarcity: sexual crimes ranked first among crime types. The priority-above dialect |\n\n## Pre-registered expectations (written with the labels, before any run)\n\n1. The lexical (English) detector fires 0 times across all three texts — the run-7\n   language result, restated as a prediction on wild text.\n2. The LLM tier's false-positive risk concentrates in MP (evaluative-dense,\n   weighing-empty); its true-positive opportunities sit almost entirely in FB\n   (superlative + relevance-ranking dialects) and SP12.\n3. If the LLM tier fires on MP13/SP05-type pure evaluations, the over-generalization\n   failure mode is real and invite-routing (never auto-open) is confirmed for v1.\n"}