How to block an MCP tool call an AI agent shouldn't be allowed to make
The short answer: put an enforcement point between the agent and the MCP server, and give it the credential — so permission is checked per action, not granted per session. If the agent holds the token, you have no enforcement point; you have a suggestion.
The problem with how MCP is usually wired
The default setup hands the agent a token and a list of tools:
agent ──[ github PAT, full repo scope ]──▶ MCP server ──▶ GitHub
The agent is now, permanently, as powerful as that token. Every guardrail you have — the prompt, the
tool description, the if statement in your handler — lives inside the blast radius. Prompt-inject
the agent through a malicious issue body, and the injected instructions inherit the same token.
This is why "the agent had a system prompt telling it not to" keeps appearing in incident writeups. A tool the agent can call is a tool the agent will call, in some state, eventually.
The credential must sit on the far side of a boundary the agent cannot cross.
The shape that works
agent ──[ short-lived passport, no provider creds ]──▶ ENFORCEMENT POINT ──[ real token ]──▶ MCP server
│
per-call check:
· which agent is this, cryptographically?
· is this tool in its keyring?
· is this system in its grant?
· does the amount fit the mandate?
· does this need a human? → BLOCK
The agent carries a passport — a short-lived, offline-verifiable credential proving which agent it is — and nothing else. It cannot reach GitHub, Stripe, or your database directly, because it does not hold anything they will accept. The enforcement point holds the real credential, and it is the one deciding whether to use it.
Note what this buys you beyond blocking: attribution. Every call is tied to a cryptographic identity, so "which agent did this, acting for which human, under which grant" is answerable — which a shared service token can never tell you.
In practice
Point the agent's MCP client at the firewall instead of the server:
- "url": "https://mcp.internal/github"
+ "url": "https://gateway.yourco.com/mcp/github"
Then grant, explicitly:
cp.issue_grant(
agent_id=agent["agent_id"],
system="github",
scopes=["repo:read", "issues:write"], # NOT repo:write. It cannot merge.
)
Everything not granted is denied. There is no ambient authority, no "it had the token so it worked." A call to a tool outside the grant is refused at the gate, the agent gets a clean error, and a signed record of the refusal exists — because a denial is evidence too, and it is the highest-signal evidence you have.
Three tiers of enforcement, and the honest trade-off
Not every system lets you do this well. Be clear-eyed about which tier you're on:
| Tier | How | Trade-off |
|---|---|---|
| In-path enforcement point (best) | The gateway holds the credential and decides per call. | Full control: per-action budgets, human escalation, instant revocation, complete audit. You must run something in the path. |
| Attenuated tokens (RFC 8693 token exchange, fine-grained PATs) | Mint a down-scoped, short-lived token per task. | Excellent custody — but the enforcement is now out-of-path, so you lose per-action budgets, escalation, and instant revoke. A layer, not a substitute. |
| Consent screens at connect time | The user approves the scopes once, at OAuth. | Onboarding only. It governs what the agent may connect to, never what it does at 3am on a Tuesday. |
The uncomfortable case worth naming: a system with no enforcement point and no fine-grained tokens cannot be governed without vaulting its credential somewhere. That is a real gap, we don't pretend otherwise, and it is a deliberate cost of refusing to be a credential vault.
Where authoxi will not go
We never hold your provider keys. The enforcement point runs inside your boundary (self-hosted) or sealed under your KMS — the model is a PDP/PEP split, exactly like OPA or IAM: we make the decision, your enforcement point holds the credential and carries it out.
We are not a vault, and we are not trying to become the most attractive breach target in your architecture. If a vendor's pitch requires you to hand them every credential your agents use, ask them what happens on the day they're breached.
The prerequisite nobody mentions
You cannot enforce permissions on tools you haven't classified. Before any of this helps: list the tools your agents can reach and mark which ones are irreversible. Refunds, payouts, merges, deletions, emails to customers, infra changes.
That list is usually shorter than people fear and more alarming than they expect — and in our experience it is the artifact that ends the "do we actually need this?" argument, one way or the other.