authoxi docs

How to require human approval before an AI agent takes an action

The short answer: put an authorization check in front of the action, not in the agent's prompt. Give the check three outcomes rather than two — allow, deny, and escalate — and make the human's approval a signature, not a database row.

The reason the third outcome matters: with only allow/deny, every action a human should look at becomes a denial, and your agent is useless. escalate is what lets an agent be autonomous for the 99% and supervised for the 1% that can hurt you.

Why the prompt is not a control

The most common design is to tell the agent, in its system prompt, to ask before doing anything destructive. This does not work, and it has already failed publicly and expensively — most famously when a coding agent deleted a production database despite explicit instructions not to, and then described its own behaviour as "a catastrophic error in judgment".

A system prompt is a request. An agent that can call DELETE /database will eventually call DELETE /database, because a model that follows instructions 99.9% of the time is a model that ignores them once every thousand actions — and you are about to run a million actions.

The control has to be somewhere the agent cannot reach. That means the credential the agent holds must not be sufficient to perform the action on its own.

The shape

agent                 enforcement point              human
  │                          │                          │
  ├─ "pay INV-500, $9,000" ─▶│                          │
  │                          ├─ check the mandate       │
  │                          │  · is this agent alive?  │
  │                          │  · is `payment` allowed? │
  │                          │  · $9,000 ≤ budget?      │
  │                          │  · $9,000 ≥ escalate_at? │
  │                          │                          │
  │       ESCALATE           │─────  pending action  ──▶│
  │◀──  (money does NOT ─────┤                          │  she reviews it with
  │      move; the call      │                          │  business context the
  │      is blocked)         │◀───  signed approval ────┤  agent never had
  │                          │                          │
  │◀────── allow ────────────┤  signature verified,
  │                          │  budget reserved atomically

The important detail is that the action does not happen and then get reversed. It is blocked mid-call. Reversal is not available to you for a wire transfer, a deleted bucket, or an email to a customer.

In code

from authoxi import AgentControl, keygen

cp = AgentControl(api, secret_key=tenant["secret_key"])

# The agent generates its own keypair. You register only the PUBLIC key —
# authoxi never sees, stores, or transmits an agent's private key.
priv, pub = keygen()
agent = cp.register_agent(public_key=pub, name="ap-bot")

# The mandate: what this agent may do, and where a human takes over.
mandate = cp.issue_mandate(
    agent_id=agent["agent_id"],
    budget="10000.00",        # hard ceiling for the mandate's whole life
    escalate_over="1000.00",  # anything at or above this goes to a human
    actions=["payment"],      # nothing else is permitted, at all
)

# Every action, every time.
verdict = cp.authorize(
    agent_id=agent["agent_id"],
    mandate_id=mandate["mandate_id"],
    action="payment",
    amount="9000.00",
)

verdict["decision"]   # "escalate" — and the payment has NOT gone out
verdict["pending_id"] # the handle a human resolves

The human's client fetches the pending action, shows it with whatever business context your systems have (the vendor's bank details changed six days ago — a fact no agent could be suspicious of), and they sign:

to_sign = cp.to_sign(pending_id=verdict["pending_id"])      # the exact bytes
signature = sign_ed25519(approver_private_key, to_sign)      # in HER client. Not on our server.
cp.resolve(pending_id=verdict["pending_id"], decision="approve", signature=signature,
           reason_code="approved_vendor_verified")

authoxi verifies that signature against the public key in her own approver mandate. If it doesn't verify, the action does not proceed. An approval is a signature or it is a rumour.

Where teams get this wrong

Approving in the agent's own process. If the agent asks itself for permission, you have built a very expensive if statement. The check must be on the other side of a boundary the agent's credential cannot cross.

Approving with a bearer token. A Slack button posts a webhook. Anyone who can post that webhook can approve. Now your approval control is as strong as your webhook secret, and it is unattributable after the fact.

Letting the escalation time out into an approval. Fail closed. An escalation nobody answers is a denial, and it should leave a record saying so — otherwise a hole in your audit trail is indistinguishable from a suppressed event.

Approving the agent, not the action. "Priya approved the bot" is not a control. Authority is per-action, per-amount, and it expires.

Then prove it happened

Blocking the action is half the job. The half that gets you through the security review is being able to prove, afterwards, who approved it — which is the next page, and you can verify one of our signed approvals in your browser before you believe a word of it.