| Decision quality. Zero wrong or fabricated answers accepted for reliance under computed acceptance, across 2,520 graded answers; accept-on-claim accepted seven. | The mechanism changes which outputs become relied on work at the point of acceptance. | 2,520 graded answers, six language models, paired task-for-task; 7 vs 0 false accepts. | Recomputable, verifiable tasks; a soundness property of recomputation; not a general accuracy claim; no field result. |
| Evidence-qualified admission. Evidence presence, rather than answer correctness alone, determined admissibility. | Required evidence functions as a real admission condition, not decoration. | 0 of 270 correct, evidenced answers refused; 90 of 90 correct-but-unevidenced answers refused. | A separate 30-task legitimate-work battery of answerable tasks, run in two conditions across the same six models; the terse condition supplies no evidence by construction. How the 270 and 90 answers are allocated across those conditions is not specified in the cited v99 paper, so the counts are given as reported results rather than reconstructed from the task count. The 90 of 90 refusals are the designed cost of the discipline, published with the benefit. |
| Recorded route composition. Governed routing matched the strongest static comparator with one-third fewer model calls. | Some valid outcomes are reachable through cheaper, exact routes. | A fixed 36-task set; 24 model calls against 36; twelve tasks answered by exact computation with no frontier-model calls; every outcome correct in both arms. | The battery saturated: it cannot resolve a quality difference and licenses no parity claim. It does not show lower total cost. |
| Live routing: modeled savings and measured outcome limits. Most modeled savings occurred where the outcome was identical. On the same campaign, governed routing lost correct disposition against the frontier comparator. | Most of the modeled savings did not depend on accepting a worse outcome, and end-to-end superiority is not established on this campaign. | Paired 240-task live campaign; 75 identical correct refusals at less than half the comparator’s modeled cost under the declared list-price model; ~96.5% of total modeled savings attributable to same-outcome pairs: a share of savings, not a 96.5% cost reduction. On the same campaign: 21 false accepts against 1 for the frontier comparator, and 0.883 correct disposition against 0.988. | Declared list-price cost model, which diverges materially from observed usage cost; descriptive, not powered confirmatory evidence. The false-accept comparison is unstable at one comparator event and is reported as counts for that reason. The defensible conclusion is selective, calibrated routing under measured constraints, not that lower-cost routes preserve frontier-model outcome quality. This is one recorded campaign on one configuration; it establishes neither the performance of later configurations nor that the deficit was cured. |
| Synthetic hazard testing. Six planted hazard designs were blocked on every repeated trial, with no good work refused. | A control nobody attacked is a control nobody has tested. | Six hazard designs, one per scenario, × 60 repeated trials = 360 hazard trials; zero unsafe permissions on the hazard arm; zero good work refused on the matched clean arm. | A synthetic governed-world estate, not the live campaign. The repeated trials are repeatability evidence rather than sample size, over designs authored alongside the controls that caught them; 4 of 6 hazard scenarios share one window-opening mechanism. It is existence evidence about this estate, not a general safety property. |
| Release-path mutation testing. Every deliberately planted release regression was detected. | A regression suite that has never caught a regression has not been shown to work. | 46 of 46 release regression mutants detected through the product’s real release path; every declared output regenerated. | None of the 46 carried a matched clean control; causal attribution follows for a separate 8 paired mutants only. Author-side regeneration is not independent validation. |