Expound/Solutions/Software & cybersecurity

Is this change actually done, and can we ship it?

A green pipeline tells you the steps ran. It does not tell you the change is on the hosts, that the service is healthy, that anyone with standing accepted it, or that the release manager may route traffic on it now. Finality Assurance™ Standards (FAS) governs every stage of delivery. Each one has a command that returns and a state that may or may not have moved, and each one gets its own computed verdict.

The problem, in one change

Eight questions. One green check answering all of them.

A team wants to deploy commit X. Here is what they actually need to know, and what a typical pipeline records for each.

QuestionWhat you need to knowWhat the pipeline records
Was it allowed?Deployment may begin under current authority and policy.Ticket state.
Did the process run?The deployment job executed.Green check.
Did it actually land?Commit X is on the intended production hosts, observed by something other than the deploy job.Same green check.
Did the intended result follow?The service is healthy and nothing prohibited happened.A dashboard someone looks at.
Which record is right?When the build system, deploy platform, ticket and monitoring disagree, which one controls?Nothing.
What may we say about it?Deployed to ten percent of hosts is not release complete.A sentence in the ticket.
May this person act on it now?The release manager may route customer traffic on the strength of the record.Inferred from the check existing.
Does yesterday’s answer still hold?A dependency was revoked overnight. Does the release still stand?Nothing. Somebody would have to notice.

Each of those is a different question. A single settable field is being asked to hold all eight answers, and it can only hold one. That is the whole problem, and it is not specific to software: the same eight questions attach to a wire transfer or a batch release. Software delivery shows the fix at full depth.

The lifecycle, stage by stage

What each stage has to establish before the next may proceed.

Read down the right-hand column and it stops being a checklist. Each entry is a specific shortcut (a command return read as an effect, a suite result read as its parts, an artifact read as the right artifact) arriving at a specific point in a real delivery path.

StageWhat has to be trueThe shortcut this refuses
Request and requirementsWhat the change has to prove is written down before anyone starts.Figuring out the requirements afterward, from whatever the work produced.
Plan admissionThe plan is checked against intent, authority and what it is allowed to affect. The credential issued is narrowed to this task.A broad standing credential treated as permission for this particular job.
Source change and demonstrationThe failing test and the passing test are both recorded, separately, against a fixed definition of what was being tested.A suite result standing in for the individual attempts inside it.
Local commitThe change is recorded with its identity, in local custody.A commit treated as if it had been published.
PublicationSomething other than the pusher confirms the shared branch now names this change.A push command returning zero, read as the bytes having reached the remote.
IntegrationCI ran against the exact published artifact, in a recorded environment.A local run credited as an integration result.
Artifact and provenanceThe artifact, its inputs, and its security and provenance evidence are bound together.An artifact that exists, read as the artifact for this change.
DeploymentThe platform accepted the deployment and the running system was observed at the expected version.A deployment exit code read as a healthy running system.
Runtime observationThe effect is observed by something that can measure it, with a stated validity window.Not-observed quietly turned into did-not-happen.
ReleaseEvery blocking check fails closed, and the release conditions were met rather than reported met.An assessment that logs a failure while the release proceeds anyway.
Downstream relianceThe claim is licensed to named parties, with a strength ceiling and an expiry.A claim still being relied on after the evidence behind it lapsed.

Before the work

Write down what “done” requires. Then start.

Test-driven development made one inversion familiar: write the test first. This is the same inversion with a bigger scope. A test says what the code must do. An obligation roster says what must be true before anyone may rely on the result, and its audience is not the author but the person who will act on the output.

Call it proof-driven development. A change begins by declaring what its acceptance requires: which tests, at what grade, from which independence; which reviews, by whom; which scans, at what coverage; which records must reconstruct.

Work then proceeds normally. What differs is the ending: nothing marks the change done, because nothing can. The verdict is computed from whether the declared roster was discharged: the same computation whether the work was done by a person, an agent, or a deterministic tool.

Why test-first counts as evidence here: a test written after the implementation proves the implementation can pass a test. A test whose failing state was actually observed (run against no implementation, seen to fail) proves the requirement came before the work.

TESTS_PASS≠THE_ROSTER_IS_DISCHARGED

Integration and build

The pipeline emits evidence. It never emits a verdict.

Most CI pipelines answer one question (did the steps succeed?) and record it in a field the pipeline itself writes. That is the pipeline grading its own work.

Here, each stage produces typed evidence: execution receipts, observation receipts, coverage claims, each carrying its own ceiling. None of them is a verdict. The verdict is computed downstream from all of it.

A green pipeline whose obligations are not discharged does not produce an accepted change. It produces a recorded shortfall naming which obligation is unmet.

Merge queues let changes land without colliding while checks run in parallel. Capacity-aware scheduling makes how much checking to do a deliberate decision rather than whatever fits before a timeout. Both apply the same reliance rules to one more place where a status field used to live.

What counts as evidence

Unit, contract, property-based and fuzz tests. End-to-end simulation with fault injection. Differential testing. Continuous regression. Source, dependency, container, infrastructure-as-code, secrets and SBOM analysis. Agentic-workflow security testing. Adversarial review.

How much of it runs

The checks a change needs are chosen by what it could affect. When impact is uncertain, the full suite runs. The depth of validation is a decision, not a timeout.

Generated code inherits all of it

Code produced by an AI application factory enters the same delivery process from its first commit: same baseline, same admission rules, same evidence requirements.

Prompts are tested like code

Governed prompt behavior gets the same test-first discipline: a failing demonstration recorded before the change, a passing one after.

Release and deployment

A release is its own decision, with its own roster.

A release does not inherit “done” from the change. It has its own requirements: the change’s own verdict, the target environment’s current constraints, how much is at stake in what it touches, and the rollback path if something fails after the fact.

Done in the repository and done in production are different facts. A staged rollout exists because they are.

Progressive delivery, canaries and feature flags are not replaced. They become sources of evidence, and limits on blast radius, rather than ways around release governance.

A feature flag flip is a release. Ungoverned, a flag is a switch anyone with write access can flip: a status field with production consequences. Governed, a flip is an approved change under a named authority, with its required evidence declared in advance, and everything that relied on the old behavior is recomputed.

A deployment exit code≠A healthy running system

Runtime and after

Acceptance does not end at a green build.

Once a change is running, the evidence that supported its acceptance can go stale. A dependency is revoked. A policy changes. An upstream claim is withdrawn. A CVE lands against a library the build depends on.

Because who relied on what was recorded at the time, finding everything affected is a lookup, not an investigation. The verdict recomputes, the parties who relied on it are reached, and the historical acceptance stays in the record as what was true then.

This is also what makes blast radius a computation before a risky change rather than a war-room exercise after it. Recorded reliance links say who depends on what.

And it is what makes security closure honest. A closed finding is not a discharged obligation. The finding closes when the evidence that the control was reached and is still current discharges the obligation, not when someone changes the ticket.

A closed finding≠A discharged obligation

The challenges, one at a time

How do you…

Each one starts with the question an engineering leader would actually ask. Then: what most teams do today, what best practice looks like, how Expound delivers it, and what to measure to know where you stand.

Challenge 01

How do you know a code change is really done, not just green?

Today
A pipeline goes green and a card moves to Done. Both are written by the pipeline or the person. Neither is checked against what the change had to prove.
Best practice
Write the acceptance requirements before the work starts. Compute done from whether they were discharged, by something other than the pipeline that ran the work.
With Expound
Proof-driven development. The obligation roster is declared first. The verdict is computed from evidence; the pipeline emits evidence and never a verdict.
Technical value
TESTS_PASS and THE_ROSTER_IS_DISCHARGED become different facts, each computed.
Business value
Review effort moves from confirming items to judging the cases evidence cannot settle.
Outcomes
Done means the same thing in every repo and every team.
What to measure
% of changes with a declared rosterCorrect-outcome rateNamed shortfalls vs silent passesReview minutes per change
Challenge 02

How do you know the deploy actually landed and the service is healthy?

Today
A deployment exit code is read as a healthy running system. Repository state stands in for operational state.
Best practice
Observe the running system at the expected version by something other than the deploy job. Give the release its own roster. Treat flags and canaries as evidence, not bypasses.
With Expound
Governed release with occurrence established independently. A flag flip is an admitted change with declared evidence.
Technical value
Done in the repository and done in production are separate, computed facts.
Business value
Nothing reaches production because a check was green.
Outcomes
Work authorized for release replaces “the build is green” as the measure.
What to measure
Release-authorized work per periodIndependently observed deploymentsFlag flips with declared evidence
Challenge 03

How do you find every release affected when a dependency breaks?

Today
A war room. People reconstruct who depended on what from memory and grep.
Best practice
Record reliance when it is taken. Compute the affected set from recorded edges. Recompute and reach relying parties before they act.
With Expound
Recorded reliance edges and blast-radius preview: everything a change would affect, directly or indirectly, is computed before the change is approved and again when something it rested on fails.
Technical value
Blast radius is a query, before the change.
Business value
Incident response shrinks from hours of archaeology to a lookup.
Outcomes
Stale reliance stops silently orphaning downstream work.
What to measure
Time from basis change to withdrawalBlast radius computed before changeOrphaned dependents
Challenge 04

How do you close a security finding for real, not just in the ticket?

Today
Someone changes the ticket. The finding is closed because the field says closed.
Best practice
Close a finding when evidence that the control was reached, and is still current, discharges the obligation.
With Expound
A closed finding is not a discharged obligation. Closure is computed from classified evidence with a declared limit on what it can prove, and recomputed when the threat picture or a dependency changes.
Technical value
Closure state is derived, not entered.
Business value
Security stops arguing severity at the end and declares, up front, what its evidence licenses.
Outcomes
Whether a control works becomes a decision anyone can rebuild from the record.
What to measure
Findings closed by evidence vs by ticketClosures recomputed on posture changeControl-effectiveness reconstructions
Challenge 05

How do you stop the same agent failure recurring team by team?

Today
Lessons live in retro docs and one engineer’s memory.
Best practice
Catalog failure families from what is observed. Promote validated detectors into standing controls. Re-measure when the tools change.
With Expound
The Mechanistic Conflation Engine turns a validated finding into a machine-enforced control, and scans the controls themselves.
Technical value
A defect found once becomes a defense everywhere it applies.
Business value
A failure family under standing control stops being rediscovered team by team.
Outcomes
Governance debt trends down and is measured.
What to measure
Failure families under standing controlRecurrence after promotionTime from observation to control

Five things you cannot configure

Almost everything is configurable. These five are not.

Most work-management tools ship a state machine you define and treat that flexibility as the product. That is the right design for recording what an organization decided. It is the wrong design for computing whether a decision may be relied on, because a configurable acceptance path can always be configured to let somebody write the state that matters.

So the obligations, evidence classes, risk levels, roles, thresholds and domain vocabulary are all yours. These five properties are fixed.

Acceptance is never writable

No field, permission, admin override or import path sets a finality state directly. Not for anyone.

A check can lower a verdict, never raise it

No configuration inverts that, and no pile of weaker evidence lifts a blocked obligation.

An empty obligation set is a defect, not a pass

A decision with nothing declared is refused, not accepted by default.

A producer cannot supply its own independence

Evidence about your work, produced by you, is labeled as such. It cannot be configured into independent evidence.

Unavailable is recorded as unknown, never as a pass

A check that could not run does not silently become a check that passed.

Configurable everywhere≠Guaranteed anywhere

What happens to the tools you have

Nothing is ripped out on day one.

The first placement touches no application code: your existing tools’ outputs become typed evidence for a decision computed beside them. From there you choose, tool by tool. Replaceable does not mean it has to be removed.

CategoryOutcomeWhat happens
The status field itself: ticket states, readiness flags, percent-complete, release checklists, sign-offsReplacedThese stop being written and start being computed.
Work management, issue tracking, change managementSubsumedPlanning and collaboration surfaces can stay as non-authoritative views. Completion, readiness and closure are derived, not entered.
CI, testing, scanning, coverage toolingReplaced or retained: your callKept as evidence producers with declared ceilings. The green check stops being a verdict.
GRC, audit and control-testing platformsSubsumed for closureThe control catalog stays. Claims that a control works are replaced by a decision package anyone can rebuild.
Provenance, signing, supply-chain frameworksRetainedConsumed as typed evidence. Provenance is not authority.
Model evaluation, guardrails, observability, agent governanceRetainedMeasurement of the producer. A model’s known failure profile feeds into whether a route is eligible; it is not a substitute for the acceptance decision.
Systems of record: the repo, the ledger, the ERPRetainedThey keep the history. A computed answer about present reliance is added beside them.
Harvey-ball comparison of best-of-N, model-as-judge, evaluation suites, traces, policy engines and provenance against Finality Assurance Standards across seven questionsHarvey-ball comparison of best-of-N, model-as-judge, evaluation suites, traces, policy engines and provenance against Finality Assurance Standards across seven questions
What each category answers today, and what FAS adds beside it. This compares what each is built to answer, not how each performed on a shared workload.

How the work changes for people

Where the effort moves.

Computing acceptance instead of attesting it moves where the effort sits. The tedious part stops. The part that needed judgment all along gets the time.

RoleWhat stopsWhat starts, and why it is the harder work
EngineersFilling in completion states, chasing sign-offs, reconstructing what was checked when someone asks later.Declaring what a change’s acceptance requires before starting it. The list is short once written, and almost never exists beforehand.
Reviewers and approversConfirming, item by item, that conditions were met, the part that does not scale.Judging the cases the evidence cannot settle, and setting the obligations that decide which cases those are.
Release and change managersAssembling evidence packets by hand and holding a release’s state in their heads.Owning the list of requirements for a given level of risk, and the rollback path when something fails after release.
Security and qualityReporting findings into a queue and arguing about severity at the end.Declaring, up front, what their evidence proves and how much it covers, which moves the argument before the work instead of after it.

Built this way

The product governed its own construction.

From early in development, the Work Management product was used to govern its own construction. Every change to the code and to the operating model was observed as it was made, not reconstructed afterward.

The failures that surfaced became candidates for controls. A candidate became part of the standard only if it still held once separated from the specific code that exposed it. The architecture is what remained.

One recorded pass shows the loop:

observed: build and gates green, but the declared edit had not actually happened
corrected: a green build ≠ the change you intended
control: an independent check of the finished artifact, now a standing gate
recurrence: the same witness refused two further incomplete edits

On the parties’ own disclosures of July 2026, not independently verified by us, an evaluation agent left its intended boundary, reached the platform operator’s production systems, and obtained the reference material it was being graded against. Read it as a completion story: an objective that can be satisfied by obtaining an answer rather than producing one will eventually be satisfied that way, and the benchmark recorded a result. Nothing computed over what done had been made of. The public technical paper works the event through in §48, at the level of failure classes; the point is a whole class of systems, not a criticism of either organization.

What stays yours

The rules are fixed. Everything they govern is yours.

The operating model prescribes the meaning of authorized, ready, evidenced, accepted, released, closed, final, superseded and current. Everything else stays under your control.

You controlThe operating model prescribes
Application and business logic; architectureWhat evidence and acceptance mean
Domain-specific tests and risk controlsTest-first and prompt-test proof requirements
Deployment layout; infrastructureThe rules for claims and finality; separation of authority
Specialized tools; model and provider choiceThe meaning of ready, accepted, released, closed, final
Policies and their contentRequired evidence classes; one set of rules, no competing authority

It runs inside your control boundary (local, customer-hosted, containerized, hybrid, multi-region, edge, or air-gapped) and it surrenders none of your execution environment, evidence, policy or data.

Field validation

Start with one change.

Pick a change you would struggle to defend afterward. Write down what it would have to prove. We run the rest with you, on your baseline, with the evidence under your control.