Expound/How runtimes decide
An AI agent’s step involves eight decisions. Most agents just return an answer.
Before anyone acts on an agent’s work, someone has to settle eight things: whether the work was allowed to start, whether it ran, whether the change really happened, whether the intended result followed, which record wins, what may be said about it, who may act on it, and whether that still holds today. Most agents do not settle these eight questions separately. They return an answer and stop, and whoever receives the answer takes it as done. Where a runtime does keep a status field, one value stands in for all eight. Either way, your records cannot tell apart situations your organization is obliged to tell apart. The figures below follow one constructed example through each place this happens.
One step, eight decisions
What has to be decided, and what the agent leaves behind instead.
The eight decisions are not stages in a line. Each can come out differently from the others: a call can run without the refund happening, and a refund can happen without anyone being entitled to rely on it.
Figure 1 · Constructed example
A refund agent’s step: eight questions, and what the agent leaves behind for each.
Every row is a separate question with its own evidence. The right-hand column shows what this constructed workflow leaves behind instead: the agent’s answer and whatever its tools logged.
| Decision | The question for this refund | What is there instead |
|---|---|---|
| Admission | May this agent issue this refund, under the current refund policy and limits? | The ticket was assigned to the agent. |
| Execution | Did the refund call run? | The payments API returned 200 OK. |
| Occurrence | Did the money actually move, as seen from outside the agent? | The agent’s own report that it did. |
| Outcome | Is the customer made whole, with nothing prohibited done, such as a second refund? | Nothing. The agent’s answer is read as the outcome. |
| Canonicality | When the agent log, the payments ledger and the CRM disagree, which record controls? | Nothing declared. The dashboard someone checks first wins. |
| Assertion | What may support tell the customer? | “Your refund is complete,” sent from a partial picture. |
| Reliance | May Finance close the case on this, and until when? | An answer, addressed to whoever reads it. |
| Applicability | After a chargeback, does yesterday’s answer still hold? | Found out when the incident report is written. |
Constructed example; no customer or company is described. The eight decisions and their typical merges are set out in the core technical paper, section 3.
Execution is not occurrence
A successful call is not a refund.
The agent’s account of what it did is evidence of what it claims. It is not proof of what happened. That proof has to come from a path that could contradict the agent.
Figure 2 · Constructed example
Two ways to decide whether an agent’s refund happened.
On the agent’s word
- The agent calls the payments API.
- The API returns 200 OK.
- The agent answers “Refund issued” and stops.
- Whoever reads the answer takes the refund as done. If the runtime keeps a status, it turns green.
Nothing in this chain could have said no. A timeout that retried, a refund to the wrong account or a payment the processor later rejected all look the same here.
From a path that can contradict it
- The agent calls the payments API and reports what it did.
- The report is kept as the agent’s claim.
- The payments ledger, a system the agent does not write to, is read for this refund.
- Occurrence is recorded only if the ledger shows exactly one refund, to this customer, for this amount.
Constructed example. The rule it shows: a successful command is not an occurred effect (core technical paper, section 4.4).
Evidence has a ceiling
A thousand log lines are not an attestation.
Each kind of evidence can support a claim up to a limit that is set in advance. Piling up more of a weaker kind does not turn it into a stronger one, and no amount of evidence turns into authority.
Figure 3 · Constructed example
Evidence for the refund, by what each kind can show.
Constructed example. The ledger stands for any source the agent cannot write to. Evidence classes and their limits: core technical paper, section 4.3.
Independence is measured
Three checkers that share a blind spot can agree and still miss the same thing.
Adding reviewers adds independent assurance only when they can fail differently. Checkers that run on the same model, read the same source or use the same parser can agree with each other about the thing none of them can see.
Figure 4 · Constructed example
Three checks on the refund agent’s work, and what they share.
Looks like three
- A second agent reviews the first agent’s summary.
- An evaluator model scores the summary.
- A rules check parses the summary for “refund issued.”
All three read the agent’s own summary. If the summary is wrong, all three can agree it is right.
Counts as independent
- The payments ledger, which the agent cannot write.
- The customer’s account balance, read separately.
Independence is judged against how each check could fail: a shared model, provider, source, prompt lineage, parser or control plane is not independence.
Constructed example. Rule from the core technical paper, section 4.3.
The answer depends on the question
The same record can be enough for one use and not for another.
Even where “done” is recorded, it does not say done for what. The same evidence answers different questions differently, because each question has its own list of what must be shown.
Figure 5 · Constructed example
One refund record, two questions asked of it.
May Support count this as a refund attempt in this week’s operations summary?
This use needs the agent’s record of the call and the ticket. Both are there.
May Finance close the case and release the hold?
This use also needs the ledger to show the refund happened once, and Finance’s own approval. Neither is there yet.
Constructed example. No one sets these answers; each is computed for one question from the same record (core technical paper, section 4.2).
Applicability
Permission to rely on an answer can narrow or end. It is not forever.
If nobody records who relied on a result, for what purpose and on what basis, nobody can be told when that basis changes and the permission no longer applies. The chargeback in this example reaches Finance weeks later, as a reconciliation problem.
Figure 6 · Constructed example
What happens to the refund answer after a chargeback.
Constructed example. A permission may narrow or be withdrawn inside its window; widening it needs a new decision, never a quiet rewrite (core technical paper, sections 3.4 and 4.5).
Five familiar mistakes
Five things a runtime can read as something they are not.
Each is a pair of things that look the same in a log and are not the same in fact.
Figure 7
Five pairs a runtime can treat as one, with a refund-world example of each.
The five families are described in the core technical paper, section 4.7. Examples are constructed.
Where FAS comes in
Finality Assurance™ Standards (FAS) keeps the eight answers apart and computes each one.
Finality Assurance Standards (FAS) keeps a separate record for each requirement of a decision, so none of the eight answers is accepted just because the agent said so, and none is merged with another when it is written down. It computes the answer to each question from that record; no person or agent sets it. And it records who relied on what, so a permission can be narrowed or withdrawn with notice when its basis fails.
Take one decision your organization makes routinely and would struggle to defend afterward. Can its requirements be written down when the decision is made? Can you name the evidence that settles each one? Can you tell when that evidence, or the authority behind it, changes? Where the answers are yes, the decision is a candidate.