Expound/How runtimes decide

An AI agent’s step involves eight decisions. Most agents just return an answer.

Before anyone acts on an agent’s work, someone has to settle eight things: whether the work was allowed to start, whether it ran, whether the change really happened, whether the intended result followed, which record wins, what may be said about it, who may act on it, and whether that still holds today. Most agents do not settle these eight questions separately. They return an answer and stop, and whoever receives the answer takes it as done. Where a runtime does keep a status field, one value stands in for all eight. Either way, your records cannot tell apart situations your organization is obliged to tell apart. The figures below follow one constructed example through each place this happens.

One step, eight decisions

What has to be decided, and what the agent leaves behind instead.

The eight decisions are not stages in a line. Each can come out differently from the others: a call can run without the refund happening, and a refund can happen without anyone being entitled to rely on it.

Figure 1 · Constructed example

A refund agent’s step: eight questions, and what the agent leaves behind for each.

Every row is a separate question with its own evidence. The right-hand column shows what this constructed workflow leaves behind instead: the agent’s answer and whatever its tools logged.

DecisionThe question for this refundWhat is there instead
AdmissionMay this agent issue this refund, under the current refund policy and limits?The ticket was assigned to the agent.
ExecutionDid the refund call run?The payments API returned 200 OK.
OccurrenceDid the money actually move, as seen from outside the agent?The agent’s own report that it did.
OutcomeIs the customer made whole, with nothing prohibited done, such as a second refund?Nothing. The agent’s answer is read as the outcome.
CanonicalityWhen the agent log, the payments ledger and the CRM disagree, which record controls?Nothing declared. The dashboard someone checks first wins.
AssertionWhat may support tell the customer?“Your refund is complete,” sent from a partial picture.
RelianceMay Finance close the case on this, and until when?An answer, addressed to whoever reads it.
ApplicabilityAfter a chargeback, does yesterday’s answer still hold?Found out when the incident report is written.

Constructed example; no customer or company is described. The eight decisions and their typical merges are set out in the core technical paper, section 3.

Execution is not occurrence

A successful call is not a refund.

The agent’s account of what it did is evidence of what it claims. It is not proof of what happened. That proof has to come from a path that could contradict the agent.

Figure 2 · Constructed example

Two ways to decide whether an agent’s refund happened.

On the agent’s word

  1. The agent calls the payments API.
  2. The API returns 200 OK.
  3. The agent answers “Refund issued” and stops.
  4. Whoever reads the answer takes the refund as done. If the runtime keeps a status, it turns green.

Nothing in this chain could have said no. A timeout that retried, a refund to the wrong account or a payment the processor later rejected all look the same here.

From a path that can contradict it

  1. The agent calls the payments API and reports what it did.
  2. The report is kept as the agent’s claim.
  3. The payments ledger, a system the agent does not write to, is read for this refund.
  4. Occurrence is recorded only if the ledger shows exactly one refund, to this customer, for this amount.

Constructed example. The rule it shows: a successful command is not an occurred effect (core technical paper, section 4.4).

Evidence has a ceiling

A thousand log lines are not an attestation.

Each kind of evidence can support a claim up to a limit that is set in advance. Piling up more of a weaker kind does not turn it into a stronger one, and no amount of evidence turns into authority.

Figure 3 · Constructed example

Evidence for the refund, by what each kind can show.

Log linesThe process ran.
An attestationA named party says it happened.
An independent observationThe ledger shows it happened.
AuthoritySomeone entitled to decide says it may be relied on. Not evidence, and not reached by stacking evidence.

Constructed example. The ledger stands for any source the agent cannot write to. Evidence classes and their limits: core technical paper, section 4.3.

Independence is measured

Three checkers that share a blind spot can agree and still miss the same thing.

Adding reviewers adds independent assurance only when they can fail differently. Checkers that run on the same model, read the same source or use the same parser can agree with each other about the thing none of them can see.

Figure 4 · Constructed example

Three checks on the refund agent’s work, and what they share.

Looks like three

  • A second agent reviews the first agent’s summary.
  • An evaluator model scores the summary.
  • A rules check parses the summary for “refund issued.”

All three read the agent’s own summary. If the summary is wrong, all three can agree it is right.

Counts as independent

  • The payments ledger, which the agent cannot write.
  • The customer’s account balance, read separately.

Independence is judged against how each check could fail: a shared model, provider, source, prompt lineage, parser or control plane is not independence.

Constructed example. Rule from the core technical paper, section 4.3.

The answer depends on the question

The same record can be enough for one use and not for another.

Even where “done” is recorded, it does not say done for what. The same evidence answers different questions differently, because each question has its own list of what must be shown.

Figure 5 · Constructed example

One refund record, two questions asked of it.

VERIFIED

May Support count this as a refund attempt in this week’s operations summary?

This use needs the agent’s record of the call and the ticket. Both are there.

BLOCKED

May Finance close the case and release the hold?

This use also needs the ledger to show the refund happened once, and Finance’s own approval. Neither is there yet.

Constructed example. No one sets these answers; each is computed for one question from the same record (core technical paper, section 4.2).

Applicability

Permission to rely on an answer can narrow or end. It is not forever.

If nobody records who relied on a result, for what purpose and on what basis, nobody can be told when that basis changes and the permission no longer applies. The chargeback in this example reaches Finance weeks later, as a reconciliation problem.

Figure 6 · Constructed example

What happens to the refund answer after a chargeback.

Relied onFinance is permitted to close the case, for this refund, for 30 days. The permission is recorded, with what it depends on.
Something changesThe card network reverses the refund. One thing the answer depended on has failed.
Everyone affected is foundThe decisions that depended on it are looked up from the record, not remembered.
Permission narrows, with noticeFinance is told before acting again. The original decision stays on record as what was known then.

Constructed example. A permission may narrow or be withdrawn inside its window; widening it needs a new decision, never a quiet rewrite (core technical paper, sections 3.4 and 4.5).

Five familiar mistakes

Five things a runtime can read as something they are not.

Each is a pair of things that look the same in a log and are not the same in fact.

Figure 7

Five pairs a runtime can treat as one, with a refund-world example of each.

Execution
The call succeeded.
≠
The change happened as authorized.
Method
The test of the refund limit failed as expected.
≠
The limit itself was reached and held.
Authority
An approval record exists. The signature on it is valid.
≠
Someone entitled to approve it issued it.
Trust domain
A separate process checked the refund.
≠
A separate party, able to fail differently, checked it.
Custody
The card token was removed from the current record.
≠
The token was never exposed in the history.

The five families are described in the core technical paper, section 4.7. Examples are constructed.

Where FAS comes in

Finality Assurance™ Standards (FAS) keeps the eight answers apart and computes each one.

Finality Assurance Standards (FAS) keeps a separate record for each requirement of a decision, so none of the eight answers is accepted just because the agent said so, and none is merged with another when it is written down. It computes the answer to each question from that record; no person or agent sets it. And it records who relied on what, so a permission can be narrowed or withdrawn with notice when its basis fails.

Take one decision your organization makes routinely and would struggle to defend afterward. Can its requirements be written down when the decision is made? Can you name the evidence that settles each one? Can you tell when that evidence, or the authority behind it, changes? Where the answers are yes, the decision is a candidate.