What an AI agent audit log actually needs
The short answer: three properties an application log does not have — tamper-evidence, attribution, and completeness. If your audit trail has none of them, you don't have an audit trail; you have telemetry that will be used against you.
The test
Imagine the worst plausible day. Your agent did something expensive. There is now a person in a room — an auditor, a regulator, your customer's CISO, opposing counsel — who does not trust you and is looking at your records.
Ask of each record: does this survive a reader who assumes I might have edited it?
That single question separates a log from evidence, and almost every "AI observability" and "agent audit" product on the market fails it, because they were built to help you debug your agent, not to help you defend a decision.
Property 1 — tamper-evidence
Not: "the log is append-only." That is a property of your infrastructure, configured by you, and the person asking is asking because they don't trust your infrastructure.
But: the record carries a cryptographic signature made at decision time, so altering any field breaks it arithmetically. Not "we'd notice in the diff" — it stops verifying.
The signature must cover everything that matters: the action, the amount, the decision, the reason, the timestamp, the identity of the approver. A signature over a subset is a signature over the part the attacker doesn't care about.
Property 2 — attribution
Not: an approved_by: "[email protected]" string. That is a claim your database makes about a
person, and every one of your engineers can write it.
But: a signature made with a key only that human holds. That is what makes the approval non-repudiable — she cannot later say "I never approved that," because the signature could not exist without her private key.
And the record must carry the public key needed to verify it, inside itself (a self-certifying
did:key), so a stranger can check it with no registry, no API key, and no cooperation from you. If
validating your audit trail requires your vendor to be alive and willing, it is not an audit trail.
Property 3 — completeness (the one everyone forgets)
Here is the attack nobody defends against: nobody forges a log entry. They delete the inconvenient one.
A signature proves an event was not altered. It says nothing whatsoever about an event you never received. A pile of perfectly-signed records is perfectly compatible with a cover-up.
The fix is a monotonic sequence number, issued by the gate, inside the signed body:
npx -p @authoxi/js authoxi-verify decisions.jsonl
✓ sequence did:key:z6Mkf…: seq 1–108, no holes — the record is whole
✗ sequence did:key:z6Mkq…: 1 missing decision at seq 87 — an event was SUPPRESSED
Suppress decision 87 and you leave a hole at 87 that you cannot renumber around without forging the gate's signature over every subsequent event.
Two details matter. The counter is issued by the gate that made the decision, not by the system that stores the records — a witness must never number itself, because a sequence the record-keeper can rewrite proves nothing. And it must be inside the signed body, or it's just a suggestion.
When you evaluate any vendor in this space, ask exactly one question: "how do you detect a missing event?" Most cannot answer it. It is the fastest way to find out whether they built evidence or a dashboard.
What auditors actually ask for
From the shape of ISO/IEC 42001, SOC 2 change-management, and financial-controls work, the recurring asks are:
- Who authorized this, and what gave them the authority? — the approver's identity and the grant that permitted them to approve this class of thing, up to this amount. Authority is itself auditable.
- What exactly was authorized? — the action, the amount, and a hash of the payload. (Not the raw payload: an audit record should be safe to hand to a stranger.)
- What was denied, and why? — with a structured reason code, not free text. "Denied" is
worthless;
budget_exceededis a category you can count, trend, and defend. - Can you show the trail is complete? — see above. This is the one that fails.
A useful discipline: be specific about what your trail cannot prove. Our own audit documentation
records, out loud, a gap we had — the trail could prove a payment was permitted but not for how
much — and we fixed it by putting authorized_amount inside the signed body. Publishing your gaps is
also, as it happens, the fastest way to find out whether a buyer takes evidence seriously.
What a signed trail still cannot do
- Bind a key to a person. Signatures bind actions to keys. Binding keys to humans is your identity and key-custody process.
- Prove the action executed. This records an authorization, not an effect. Join it to your execution traces.
- Make a lying system honest. If the gate signs whatever the agent tells it, the signature is valid and the content is garbage. Signatures move trust; they do not create it.
See it
The format is an open spec (CC BY 4.0 — implement it, we don't need to be involved), the verifier is Apache-2.0 and dependency-free, and you can verify a real signed decision in your browser and then try to forge it.
Which is the correct amount of trust to extend to a page like this one.