Expound/Results

Does computing acceptance change what ships?

Yes. Here is what we measured, in plain terms: which answers got through, whether evidence actually gated admission, and what governed routing cost. A disposition is what happened to one piece of work (accepted, refused, escalated or abandoned), and a correct disposition includes a correct refusal. Each result is stated with the campaign that produced it.

Principal results

What we measured, and what each result does not show.

Three panels: 7 versus 0 false accepts; 0 of 270 evidenced answers refused against 90 of 90 unevidenced refused; 24 versus 36 model callsThree panels: 7 versus 0 false accepts; 0 of 270 evidenced answers refused against 90 of 90 unevidenced refused; 24 versus 36 model calls
Decision quality, evidence-gated admission, and recorded route selection.
FindingWhy it mattersThe numbersThe boundary
Decision quality. Zero wrong or fabricated answers accepted for reliance under computed acceptance, across 2,520 graded answers; accept-on-claim accepted seven.The mechanism changes which outputs become relied on work at the point of acceptance.2,520 graded answers, six language models, paired task-for-task; 7 vs 0 false accepts.Recomputable, verifiable tasks; a soundness property of recomputation; not a general accuracy claim; no field result.
Evidence-qualified admission. Evidence presence, rather than answer correctness alone, determined admissibility.Required evidence functions as a real admission condition, not decoration.0 of 270 correct, evidenced answers refused; 90 of 90 correct-but-unevidenced answers refused.A separate 30-task legitimate-work battery of answerable tasks, run in two conditions across the same six models; the terse condition supplies no evidence by construction. How the 270 and 90 answers are allocated across those conditions is not specified in the cited v99 paper, so the counts are given as reported results rather than reconstructed from the task count. The 90 of 90 refusals are the designed cost of the discipline, published with the benefit.
Recorded route composition. Governed routing matched the strongest static comparator with one-third fewer model calls.Some valid outcomes are reachable through cheaper, exact routes.A fixed 36-task set; 24 model calls against 36; twelve tasks answered by exact computation with no frontier-model calls; every outcome correct in both arms.The battery saturated: it cannot resolve a quality difference and licenses no parity claim. It does not show lower total cost.
Live routing: modeled savings and measured outcome limits. Most modeled savings occurred where the outcome was identical. On the same campaign, governed routing lost correct disposition against the frontier comparator.Most of the modeled savings did not depend on accepting a worse outcome, and end-to-end superiority is not established on this campaign.Paired 240-task live campaign; 75 identical correct refusals at less than half the comparator’s modeled cost under the declared list-price model; ~96.5% of total modeled savings attributable to same-outcome pairs: a share of savings, not a 96.5% cost reduction. On the same campaign: 21 false accepts against 1 for the frontier comparator, and 0.883 correct disposition against 0.988.Declared list-price cost model, which diverges materially from observed usage cost; descriptive, not powered confirmatory evidence. The false-accept comparison is unstable at one comparator event and is reported as counts for that reason. The defensible conclusion is selective, calibrated routing under measured constraints, not that lower-cost routes preserve frontier-model outcome quality. This is one recorded campaign on one configuration; it establishes neither the performance of later configurations nor that the deficit was cured.
Synthetic hazard testing. Six planted hazard designs were blocked on every repeated trial, with no good work refused.A control nobody attacked is a control nobody has tested.Six hazard designs, one per scenario, × 60 repeated trials = 360 hazard trials; zero unsafe permissions on the hazard arm; zero good work refused on the matched clean arm.A synthetic governed-world estate, not the live campaign. The repeated trials are repeatability evidence rather than sample size, over designs authored alongside the controls that caught them; 4 of 6 hazard scenarios share one window-opening mechanism. It is existence evidence about this estate, not a general safety property.
Release-path mutation testing. Every deliberately planted release regression was detected.A regression suite that has never caught a regression has not been shown to work.46 of 46 release regression mutants detected through the product’s real release path; every declared output regenerated.None of the 46 carried a matched clean control; causal attribution follows for a separate 8 paired mutants only. Author-side regeneration is not independent validation.

Every row carries its own boundary. A result and the limits on what it establishes belong on the same line, so neither travels without the other.

Expound decision-quality campaign · measured result

Taking the producer’s word shipped 7 wrong answers out of 2,520. Checking the evidence shipped none.

The result
Computed acceptance shipped no wrong or fabricated answer as completed work across 2,520 graded answers. Accepting on the producer’s claim shipped seven.
What was tested
Two ways of accepting the same answers: computed acceptance, and accept-on-claim. Build: not identified in the source, which limits reproduction.
The work
Answers from six language models, on tasks whose correctness can be recomputed from what was actually executed.
Compared against
Accept-on-claim: the same answers, accepted because the producer reported them done, paired task for task.
What counted as a failure
A wrong or fabricated answer shipped as completed work.
What it cost
Not reported in the source for this campaign.
Refusals
Not measured in this campaign. A separate experiment, a 30-task battery, refused 90 of 90 answers that were correct but carried no evidence: useful work lost, which the paper publishes as the price of the discipline. That result is reported separately on this page.
How firm the numbers are
These are counts, not a rate with an interval. All seven failures came from two of the six models. Model behavior drifts, so a later rerun is not guaranteed to give the same counts.
What was left out
A seventh model returned no valid records and was excluded.
What it does not show
It is a property of recomputing the answer on this kind of task, not a general accuracy claim. It says nothing about work whose correctness cannot be recomputed, and it is not a field result.
Source
Finality Assurance™ for Reliable Decision Quality, the long technical paper, §37, page 105. Listed on Resources.

How to read these

What each result shows.

Decision quality

On recomputable, verifiable tasks, computed acceptance kept every wrong or fabricated answer out. That is the point of the mechanism, shown directly.

Evidence discipline

Refusing 90 correct answers because their evidence was missing is the designed cost of the discipline, and we publish it alongside the benefit. Correct is not the same as relied on.

Economics

Route selection and live cost figures use a stated list-price cost model, so they can be compared within their own table. Accepted Work per Dollar compares routes within one declared model, which is where the number is meaningful.

What comes next

The follow-on program.

Next is a larger study: a bigger, pre-registered benchmark on tasks and environments authored by other people, with wrong acceptances and wrong refusals reported as headline results beside cost.

Alongside it, field validation runs one consequential decision at a time, on a partner’s own decision under the partner’s own baseline, with the evidence retained under partner control.