Expound/Platform/Ladder & Four Corners
Claim ceilings and sufficiency, kept apart from the verdict
The Confidence Ladder™ has one job: to bound what may be claimed about a result, from “reproduces under my own hand” up to “reproduces independently,” with every rung derived from retained evidence. The Four-Corners Test™ asks whether someone else could rebuild the conclusion from what is attached. Both inform the decision. Neither is the decision. Permission to rely comes from a verdict.
The Confidence Ladder
One job: bound what may be claimed.
The Confidence Ladder is a ranked ceiling, derived from evidence, on what may currently be claimed. It is not a probability and not a model-confidence score. Every rung is a ceiling on assertion, not a score. The same evidence can support “these bytes reproduce on my machine” and be quietly inflated into “this result is independently reproducible.” The ladder makes that inflation a rule violation rather than a judgment call.
Each rung is a test over kept evidence, evaluated for each campaign. The three rungs below R0 are not decoration: collapsing no package, broken package and package tested against bytes that no longer exist into one not-yet state destroys the information needed to act.
| Rung | Holds when | You may claim |
|---|---|---|
NOT_PACKAGED | The reproduction contract does not compile. | Nothing about reproduction. |
PACKAGED_ROUTES_FAIL | A package exists; its routes do not pass. | That the package exists. |
PACKAGED_NO_CURRENT_EVIDENCE | Routes passed, but against bytes that have since been superseded. | History only. |
| R0 | Routes A, B and C pass on hosts that one principal controls. | “Reproduces under my own hand.” |
| R1 | R0 and every term of the R1 predicate holds against an accepted governing root. | “An outsider could reproduce this.” |
| R2 | One positively identified independent principal did reproduce it. | “Someone else reproduced it.” |
| R3 | Two or more independent principals did. | “It reproduces independently.” |
What may be said about a result≠Who may act on it
Any score that licenses reliance becomes a score worth gaming. So the ladder bounds claims and informs disclosure; it never authorizes. The evidence ceiling on a reliance license is a separate scale doing a different job: the ladder caps what may be said, the ceiling gates what may be done.
What makes it sound
Five properties, each enforced in code.
Every rung is derived, never asserted
No literal ever stands in for a measurement. A count of external reproductions is read from a receipt store; if the store is empty the answer is zero because zero was counted, not because zero was typed.
Total and disjoint
Every campaign lands on exactly one rung, always, including on error. An exception is not a rung; an unreadable subject sits at a named state with a reason.
Currentness-bound
Evidence attaches to exact bytes. Rebuild the package and its rung falls, because the thing that was tested no longer exists.
No partial credit
R1 requires every one of its terms. Two terms passing is two terms passing, not R1.
It goes down
This is what makes it a ladder rather than a scoreboard. A ratchet that only rises records the high-water mark of optimism. Demotion is routine, silent and automatic: the honest cost of changing anything.
A hard-coded zero≠A computed count
An exception≠A FAIL verdict
The package was rebuilt≠The rebuilt package was tested
Routes pass and plan accepted≠R1 reached
What no internal act can move
Independence is a property of who, not how many.
R0 to R1 requires an accepted governing root. R1 to R2 requires a principal nobody here controls. R2 is not reachable by effort. More hosts, more interpreters, more architectures, more careful controls: all of it improves R1 and none of it touches R2.
The test for R2 requires positive identification of a different, independent party, and treats no identification as a no. A placeholder identifier that merely differs from the producer’s string is not a different principal; without that rule, a run on the producer’s own machine could be awarded independence.
Multi-host reproduction≠External reproducibility
An unidentified principal≠A different principal
Not known to be the same≠Known to be different
Each rung names what is unmet
A rung that reports only its level is unusable. It reports which terms are unmet and what evidence would change the answer, so the reading is “R0 because these three of fifteen terms fail,” not “R0.” The ladder tells you where you are and how to move.
The ladder is discriminating
Each rung’s test is shown to fail against a deliberately broken subject. A ladder whose R1 test cannot fail would assign R1 to everything and mean nothing, so the ability to return false is tested along with the rung.
What it refuses to be
Not a percentage. Not an average across campaigns. Not a summary number. Twenty-one campaigns at R1 and zero at R2 is not “50% confident”; it is exactly what it says, and the shape carries the meaning. Collapsing it to a scalar would destroy the only thing the ladder was built to protect.
The test passed≠The test can fail
When finality cannot be reached
A sensor fails on the line. Does the line stop?
The ladder is usually introduced as a ceiling on claims, which makes it sound purely restrictive. Its more useful job is the opposite one, and it is the one real operations need most: it says what remains permissible when a finality verdict cannot be reached, or when Continuous Finality withdraws one that previously held.
Consider the case that decides whether an architecture is usable. A sensor on an assembly line fails. The obligation that sensor discharged is now undischarged, so the verdict for anything depending on it can no longer be VERIFIED. The naive reading is that the line stops. That reading is wrong, and a design whose only response to imperfect evidence is to halt is one operators route around, which removes the governance rather than the risk.
What actually happens
The operational verdict is recomputed rather than the claim disappearing. That is the Finality Account at work (the graded verdict), and it is distinct from the ladder, which limits only what may publicly be claimed and never decides whether the line runs. The observation obligation becomes UNKNOWN, and the verdict takes the lowest grade across the list of requirements, so it can be no better than UNKNOWN while that obligation stands unmet.
How it becomes DEGRADED
Only where a declared compensating obligation, established independently of the failed source, is itself discharged, and that is not a verdict being raised. The compensating obligation discharges the roster obligation at a degraded grade, so the lowest grade is recomputed over changed evidence rather than overridden. A compensating check that shares the failure mode it compensates for raises nothing.
Who keeps their license
Relying parties whose decisions required the withdrawn evidence class are notified and their licenses narrowed. Relying parties whose decisions never depended on it keep theirs. Withdrawal of reliance need not mean withdrawal of all work.
Policy declared in advance, not improvised while the line is down, states what a DEGRADED verdict still allows at each level of risk:
Degradation is bounded in three ways, also declared in advance. Each level of risk carries a limit on how long degraded operation may persist, and on expiry the verdict falls rather than persisting. Each carries a limit on how many mandatory obligations may be degraded at once, above which the verdict falls to BLOCKED, because concurrent independent degradations erode defense in depth in a way the evidence roster alone cannot see. And entry into degraded operation is an authorized act with a named owner, not a state a system drifts into. These guards feed into the same lowest-grade rule and can only lower: they can carry a verdict down and never up.
Unknown≠False
Degraded≠Stop everything
A binary architecture has exactly two behaviors when evidence degrades: proceed as though nothing happened, or stop everything. Both are wrong most of the time. A graded verdict preserves the distinction between we no longer know this to the standard that decision required and we know this is unsafe, which is the distinction operators actually manage. The same mechanism covers evidence that expires, an authority that lapses, a dependency that is superseded, and a supplier attestation that is withdrawn. Degradation behavior is part of the specified benchmark program, because a governance architecture measured only on clean evidence has not been measured on the case that matters.
Not to be confused with
The Evidence Maturity Ladder is a different object.
Expound also classifies how mature the evidence behind a capability is, on a six-class scale: Defined, Implemented, Executed, Adversarially tested, Independently reproduced, Field validated. That is the Evidence Maturity Ladder. It answers how mature is this capability overall? over a program’s lifetime, and it generally rises as evidence accumulates.
The Confidence Ladder answers a different question (what reproduction state does this particular package and campaign support right now?) over exact bytes, and it falls as routinely as it rises. There is no rung-to-rung equivalence between the two. A capability can be field validated on the maturity axis while its current reproduction package sits at PACKAGED_NO_CURRENT_EVIDENCE, and a narrowly scoped package can reach R3 without the technology being field validated. Both bound claims; neither computes finality.
Confidence Ladder state≠Maturity class
R3≠Field validated
Independently reproduced (maturity)≠R2 or R3
The Four-Corners Test
Can a second party reconstruct it from what is attached?
Everything the answer rests on sits inside the four corners of the record, with no call to the producer and no access to the system that did the work. It is a sufficiency test, not the verdict.
What fails it
A summary fails, because a summary is the producer’s account of the evidence rather than the evidence. A decision that leans on a conversation nobody recorded fails, and so does one that leans on an assumption nobody wrote down.
What passes it
A bound set of records from which the determination recomputes to the same answer. Which is why evidence has to be bound rather than described.
What a pass does not mean
Not that the result was good. In the invoice example the test passes on a hold, meaning the reasoning for holding it can be rebuilt by someone who was not there.
Run it yourself. No tooling needed.
Take one determination your agents produced this week. Copy out everything attached to it and nothing else: no access to the system, no thread to scroll, no colleague to ask. Hand that to someone who was not involved and ask them what was decided and on what basis.
If they reach the same answer from what you gave them, it passes. If they come back asking who to talk to, it does not.
Then ask the half no amount of reading answers: if something that determination rested on were superseded this afternoon, what in your stack would move it, and who would learn that it had moved?
Next steps
Neither instrument grants permission.
Permission is a license, and it comes from a verdict.