Expound/Capabilities

How do you know the work your AI agents produce can be relied on?

That is the question underneath every other one on this page. Answering it well takes seven capabilities in your own organization. Each one starts with the question you would actually ask. Then it lays out what most organizations do today, what best practice looks like, how Expound delivers it, its technical and business value, the outcomes, and what to measure, so you can see where you stand, and what a high performer looks like. Where a capability is delivered by a product still in development, it is described as designed, and marked as such.

The maturity scale

Five levels. Your band is your lowest row.

Score one consequential decision, not your organization as a whole. Averages hide exactly the variation that matters. One unmet obligation defeats the rest, so your band is the lowest score across the capabilities, and your total is a progress measure for comparing the same decision against itself later.

0

Absent

The capability does not exist. Nobody could answer the question.

1

Asserted

Somebody declares it. The declaration is the record, and it is not checked.

2

Documented

It is written down and reviewable, but assembled after the fact, by the party being examined.

3

Evidenced

Evidence exists at the moment of change, is typed, and is separable from whoever produced the work.

4

Computed

The answer is derived from that evidence on demand, cannot be written by anyone, and is withdrawn when its basis fails.

Where does one of your decisions sit today? The point of scoring is to know what to work on next, not to grade the organization.

The seven capabilities

Challenge, current state, best practice, and how to get there.

Every harness, workflow tool, registry and deployment system plugs in underneath. These seven are the architecture. Read each one and ask which level you are at.

Capability 01 · Compute acceptance

How do you know the work an agent hands back is actually done, and that your team can act on it?

Today
The agent returns an answer and stops, and the answer arriving is taken as done. Where a system records anything at all, it is a ticket moving to a column, a pipeline going green, a reviewer signing, and nothing checks whether it still holds.
Best practice
Write down what the work has to prove before it starts. Compute whether it proved it, from evidence, by a party that did not do the work. Make that answer specific to who is asking and what they want to do with it.
With Expound
The Finality Kernel computes a verdict from an obligation roster and typed evidence: BLOCKED, UNKNOWN, DEGRADED or VERIFIED, for one exact question. Nobody can set it by hand.
Technical value
One vocabulary for done across every tool. The verdict is computed every time it is read, not stored in a column, so it cannot go stale silently.
Business value
Reviewers stop confirming items and start judging the cases evidence cannot settle. Approval chains collapse into computation.
Outcomes
Zero wrong or fabricated answers accepted for reliance across 2,520 graded answers. Work that counts is separated from work that was merely produced.
What to measure
% of consequential decisions with a declared rosterCorrect-outcome rateFalse acceptsExpert minutes per outcomeAccepted Work per expert hour
Capability 02 · Control who can change state

How do you stop anyone in your organization (person, agent or tool) from marking work as done?

Today
Any interface with write access can move the status. Admin overrides, imports, integrations and agents all have a path to the field that matters.
Best practice
Route every change of state through one gate that carries authorization and evidence. An unapproved change is refused, not recorded.
With Expound
Work Management has no field named state to set. State changes pass through one gate; a direct change that has not been approved for that exact target is refused. Every path to a decision is listed and governed.
Technical value
The write path, the real defect behind status fields, is closed. Adding more status values never fixed it; closing the write path does.
Business value
A misdirected agent reporting success in the voice of a correct one buys nothing. The cheapest route to a result that counts is to do the work.
Outcomes
Acceptance cannot be written by anyone. Ungoverned paths become defects you can count down to zero.
What to measure
Number of paths to each decision% of paths governedDirect-write attempts refusedTime from unadmitted change to refusal
Capability 03 · Produce evidence

How do you get evidence of what your systems actually did, as a by-product of the work, instead of a binder assembled afterward?

Today
Evidence is a PDF in a folder, a link in a ticket, a screenshot in a thread. It is gathered when someone asks, by the team being asked, and thrown away after.
Best practice
Label every piece of evidence with what it can prove, when it was observed, and who produced it. Capture it as the work runs. Give each class a ceiling so a thousand weak items never add up to one strong one.
With Expound
Typed evidence with declared ceilings. CI, scanners, observers and people all emit evidence, never a verdict. Evidence about your own work is labeled as exactly that.
Technical value
Each piece of evidence has an address and a version, so it can expire, be replaced, or be found not to support what cited it.
Business value
Audit preparation stops being a reconstruction project. The proof exists at the moment of change.
Outcomes
The Finality Account is the artifact you hand over. Reconstruction becomes a query.
What to measure
% of obligations discharged by typed evidenceReconstruction cost per relied-upon decisionAudit preparation hoursEvidence age at reliance
Capability 04 · Verify independently

How do you check an agent’s work without redoing it, and without taking the agent’s word for it?

Today
Two options, both bad: take the answer, or do the work again. There is rarely a third, because there is nothing to examine. So verification collapses into trust or repetition, and at machine speed it usually collapses into trust.
Best practice
Establish that the effect actually happened through a path that could contradict the producer. Give the verifier one power only: to lower a verdict, never to raise one. Make the acceptance itself re-runnable by a second party.
With Expound
An independent verifier that can narrow, hold or refuse and never raise. The Four-Corners Test™: can someone rebuild the decision from what is attached and nothing else? Independence measured against shared components, not assumed from an org chart.
Technical value
Occurrence is separate from execution. A call returning is not the change landing.
Business value
Reviewers read what is attached instead of redoing the work. Verification takes machine time, not a person’s time.
Outcomes
A third option appears between trust and repetition: check.
What to measure
Four-Corners pass rate% of occurrence evidence from an independent pathVerifier downgradesTime to verify vs time to redo
Capability 05 · Deploy accepted work

How do you make sure only accepted work from your agents and pipelines reaches production, the ledger, or the customer?

Today
A green pipeline releases. A flag flips. A card reaches Done. Repository state and operational state are treated as the same fact.
Best practice
Give the release its own roster: the change’s verdict, the environment’s current constraints, how much is at stake, and the rollback path. Treat flags, canaries and staged rollouts as sources of evidence, not as ways around release governance.
With Expound
Governed release. A feature flag flip is a release: an admitted change with declared evidence that triggers recomputation of everything relying on the prior behavior. Prerequisites are written in an order a machine can check.
Technical value
Done in the repository and done in production are separate facts, each computed.
Business value
Nothing reaches production because a check was green. It reaches production because the evidence for this release, in this environment, discharged this roster.
Outcomes
Work authorized for release replaces “the build is green” as the operating measure.
What to measure
Release-authorized work per periodFlag flips with declared evidenceBlast radius computed before changeTime to accepted work inside the window
Capability 06 · Observe runtime posture

How do you know a decision your systems made last month still holds today?

Today
It does not get checked. A result correct on Tuesday is quietly wrong on Thursday because a dependency changed, and nothing was watching. Someone finds out by asking, which requires already suspecting.
Best practice
Record who relied on what, at the moment they relied. When a basis changes, compute the affected set, recompute the verdict, and reach every relying party, before the corrected answer, not after.
With Expound
Continuous Finality™. Dependencies are recorded when reliance is taken. When something changes, you get a specific, named list of who is affected rather than a broadcast or a search. A withdrawal that cannot reach someone says so, by name.
Technical value
Whether a result applies is derived on every read, not stored. The value read after a change is a value no actor wrote.
Business value
Noticing scales with the system instead of with the number of people who remember. Catching stale decisions stops depending on memory and becomes a lookup.
Outcomes
Out of use before recomputed. The history stays true as history.
What to measure
Time from basis change to withdrawalRelying parties reached / relying partiesStale reliance incidents% of decisions with recorded reliance edges
Capability 07 · Preserve knowledge

How do you turn a failure your team found once into a control that runs forever?

Today
Lessons live in a retro doc, a Slack thread, or one engineer’s memory. The same class of failure is rediscovered team by team, session by session.
Best practice
Catalog failure families from what is observed. Turn a validated finding into a machine-enforced control. Re-measure when the model, agent or tool doing the work changes. Never let the thing that discovers a failure also decide it is a control.
With Expound
The Mechanistic Conflation Engine scans for the places where two things that must stay separate were treated as one (a call returning read as the effect landing, a valid signature read as an authorized signer) and turns confirmed detectors into standing controls through a governed approval step.
Technical value
It runs across the whole system: it limits what the other parts may do, and it scans the controls themselves.
Business value
A defect found once becomes a defense everywhere it applies. Known failures need not be rediscovered.
Outcomes
Governance debt becomes a measured quantity with a downward trend.
What to measure
Failure families under standing controlRecurrence after control promotionTime from observation to controlGovernance debt per capability

The capability as an object

What a capability carries.

In the Work Management product, a capability is not a roadmap label. It is a durable organizational ability with an owner, and the system computes its status from its health, evidence and governance debt.

OwnerOne accountable person.
Who it servesWho benefits from the ability.
Business valueWhat it is worth, stated.
Target maturityThe level it is meant to reach.
DependenciesWhat it rests on, and what rests on it.
Evidence obligationsWhat it must keep proving.
Exit criteriaWhat done means for the capability itself.
Health metricsHow its health is read.
Health & governance debtComputed, not reported.
Acceptance stateA rollup over the work inside it.
Contribution to Accepted WorkWhat it adds to the count of accepted work.

From health, evidence and governance debt, the system computes one of four statuses:

BUILDThe ability is needed and not yet mature. Invest.
STABILIZEIt exists but its evidence or health is slipping. Shore it up.
KEEPMature, evidenced, healthy. Maintain.
RETIRESuperseded or no longer needed. Retirement is a computed end state; when something replaces it, the history is kept.

Because dependencies form a governed graph, the system can compute which capabilities block which releases and what must mature first. Whether a capability is accepted rolls up from whether its work was accepted: one set of rules, not a second system.

Self-assessment

Score one decision that matters.

Pick a decision you would struggle to defend afterward: a release, a payment, an approval, a batch. Score each row from 0 to 4 for that decision. Your band appears below, with what to work on first.

01
Can you produce the list of what was required before this decision was made?
What a 4 looks like: A declared obligation roster exists, is versioned, and is written by an authority separate from the work.
02
Is permission to act kept separate from evidence that the act occurred?
What a 4 looks like: Authorization and occurrence are separate records. Neither can stand for the other.
03
Is occurrence established by something other than the actor’s own account?
What a 4 looks like: Proof that it happened comes from something that could have contradicted the system that did the work.
04
Is the producer prevented from influencing whether its own work is accepted?
What a 4 looks like: The producer has no write path to the verdict and no control over the evidence that grades it.
05
When a claim is made, is it bounded to a party, a purpose, and a period?
What a 4 looks like: Every claim licenses a named party for a named decision until a named condition changes.
06
If a supporting fact became false today, what would happen?
What a 4 looks like: Dependencies are tracked, the verdict is recomputed, and every relying party is notified and unwound.
07
Is each class of evidence capped so that quantity cannot substitute for strength?
What a 4 looks like: Each class of evidence has a fixed limit on what it can prove. No amount of it crosses that limit.
08
Are confidence scores, model outputs and self-reports kept out of the verdict?
What a 4 looks like: Confidence can inform the case. It cannot move the verdict.
09
How many paths exist by which this decision can be reached?
What a 4 looks like: Every path is enumerated and governed. An ungoverned path is a defect, not an exception.
10
What happens when evidence is missing rather than negative?
What a 4 looks like: A gap produces a hold that names the missing requirement, never a default to yes and never a blanket stop.
11
Is the independence of each check declared against the ways things fail together?
What a 4 looks like: Independence is declared for each way things could fail together, against the known shared weak points.
12
Do you know the full cost of reaching an outcome, including refused and reworked attempts?
What a 4 looks like: The full cost of reaching an outcome and the value accepted are both measured, and the ratio is a tracked operating metric.
13
Do the controls you rely on stay current with how the work actually fails?
What a 4 looks like: Failure families are established from what is observed, become controls that run without anyone remembering them, and are re-measured when the model, agent or tool changes.

This is a self-assessment and a conversation starter with your own teams. Read the lowest row first, and the total second. Nothing you enter leaves this page.

What to measure

Keep every input metric. Add one governed numerator.

Tokens, model calls, GPU hours and seats are real costs and stay in the denominator. What changes is what counts as output: not activity, but work the organization is entitled to rely on.

Input metric you already trackOutcome metric it feeds
Model calls, tokens, GPU timeAccepted Work per call, and per dollar
Tickets closed; builds greenCorrect outcomes; work authorized for release
Review countAccepted Work per expert hour
Model utilizationOpportunities captured without lowering the bar
Audit effortReconstruction cost per relied-upon decision

Seven channels of value, each with the field measure that tracks it:

ChannelWhat you would measure
Avoided lossIncident, rework and remedy cost against baseline
Expert time won backExpert minutes per outcome
Lower cost per outcomeFully loaded cost per accepted outcome
Latency and opportunity valueTime to accepted work inside the decision window
Accepted-Work throughputAccepted units per period, quality held
Work you can newly hand offWork newly eligible to delegate under computed acceptance
Trust and revenue premiumAudit effort, dispute rates, counterparty friction