How to prove a human approved an AI agent's action
The short answer: the human has to sign the approval with a key only they hold, and the record has to carry the public key needed to check that signature. Anything less is a claim, not proof.
You can check one of ours in your browser right now — and then try to forge it.
The question you will be asked
It always arrives in the same shape, from an auditor, a regulator, a customer's security team, or your own incident review:
"Show me that Priya — and not a compromised service account, and not an engineer with database access, and not the agent itself — approved this specific payment, for this specific amount, and that nobody edited the record afterwards."
Notice what is being asked for. Not "do you have a log?" — everyone has a log. The question is whether the log is worth anything given that the party producing it is the party under suspicion.
Why the usual answers fail
"We have an approved_by column." It is mutable by every person and process with write access to
that database — which includes your on-call engineer, your CI system, and anyone who compromises any
of them. Crucially, it includes the person you are trying to hold accountable. A row that the
suspect can edit is not evidence about the suspect.
"We have an append-only / immutable log." Better. But its immutability is a property of your infrastructure, and the person asking is asking precisely because they do not fully trust your infrastructure. "It's immutable because we configured it to be" is a claim you make about yourself.
"We ship logs to a SIEM." Now two systems you control agree with each other.
"Slack has the approval." Slack has a message. It does not have a cryptographic act. Whoever could post that message could approve — and six months later nobody can distinguish a real approval from a replayed webhook.
"Our vendor's dashboard shows it." If a vendor has to be alive, reachable, and cooperative for you to validate your own audit trail, it is not an audit trail. It is a receipt.
The common failure is that every one of these proves the record is consistent with itself. None of them proves who decided.
What actually proves it
A signature made at decision time, with a key the approver holds and nobody else does — including you, and including us.
{
"decision": "approve",
"decision_by": "human",
"authorized_amount": "9000.00",
"reason_note": "Called the vendor to confirm the bank-detail change. Legitimate.",
"approver_did": "did:key:z6Mkeh…", ← carries her PUBLIC KEY inside it
"approver_sig": "ed25519:588a988d…", ← made in her client, with her private key
"approver_mandate_ref": "mnd_approver_2f9b" ← the grant that gave her the authority
"emitter_did": "did:key:z6MkrE…", ← the gate that enforced it
"emitter_sig": "ed25519:7aa4c851…" ← countersigns her whole envelope
}
Three properties follow, and they are the three the other answers lack:
- Tamper-evidence. Change the amount, the decision, or even her stated reason, and the signature stops verifying. Not "we would detect it in the diff" — it stops verifying, arithmetically.
- Attribution that survives denial. She cannot later say "I never approved that", because producing that signature requires her private key. This is what non-repudiation means, and it is the property a database row can never have.
- Verifiability by a stranger. The public key is inside the record's own
did:key. An auditor, a court, or an angry customer verifies it themselves, offline, with no registry, no API key, and no cooperation from you. That is what makes it evidence rather than testimony.
The subtle part: two signatures, in order
The human signs first, at decision time, over a body that excludes the gate's envelope — she cannot attest to a countersignature that doesn't exist yet. The gate signs last, over a body that includes her signature.
That asymmetry is load-bearing. Sign them independently and an attacker can lift a valid $9,000 approval off one event and staple it onto a $9,000,000 one, with both signatures still verifying. Because the gate countersigns her whole envelope, an approval is bound to that action and no other.
(There is a test for exactly this. Try it on the live page: reprice the payment and watch both signatures fail at once.)
The attack everyone forgets: deletion
A signature proves an event was not altered. It proves nothing about an event you never received — and the easy attack on an audit trail was never forgery, it was deletion. Nobody forges a log entry; they quietly drop the inconvenient one.
So every decision carries a signed, monotonic sequence number issued by the gate. Suppress decision 42 and you leave a hole at 42 that cannot be renumbered around without forging the gate's signature over every subsequent event.
npx -p @authoxi/js authoxi-verify decisions.jsonl
# ✓ sequence did:key:z6Mkf…: seq 1–108, no holes — the record is whole
A verifier that checks signatures but not continuity is checking the wrong thing. Ask any vendor selling you an "immutable AI audit trail" how they detect a missing event. Most cannot.
What this still does not prove
An honest list, because you will be asked and you should not overclaim:
- It does not bind the key to the human. A signature binds an action to a key. Binding that key to Priya is your identity and key-custody process. What you get is: whoever holds that key cannot deny the decision.
- It does not prove the action executed. This is a record of an authorization, not an effect. Join it to your execution trace for that.
- It does not make a lying system honest. If the gate signs whatever the agent tells it, the signature is valid and the content is garbage. Signatures move trust; they don't create it.
See for yourself
Don't take this page's word for any of it — that would rather defeat the purpose.
Verify a real signed decision in your browser → · or npx -p @authoxi/js authoxi-verify ·
or read the open spec and implement your own.