Leet Force
Back to Blog
DevelopmentUpdated

Your Agent Has a Permissions Problem, Not a Prompt Problem

Prompt injection is a confused deputy attack, described in 1988 and unfixable at the model layer for the same reason it was unfixable then. The containment architecture already exists, most of it is standardised, and almost nobody ships it.

Abdullah Khan
By Abdullah KhanCEO, Leet Force
Your Agent Has a Permissions Problem, Not a Prompt Problem

In 1988, Norm Hardy wrote three pages about a compiler.

The compiler ran on a shared system and billed its users. To do that it needed write access to a billing file, so it was given that authority permanently. It also let callers name an output file for debugging statistics. A caller who passed the billing file's name as the output path destroyed the billing records — not because the caller had permission, but because the compiler did. The compiler had been handed a name by someone untrusted and had applied its own authority to it, with no way to tell that it was being used.

Hardy called it the confused deputy, and gave the paper a subtitle that reads differently now: or why capabilities might have been invented.

This is not an analogy for what happens when an agent reads a malicious web page. It is the same defect, in the same place, for the same reason. A system holds standing authority. It accepts instructions from a channel it cannot authenticate. It applies the former to the latter. Everything the security community learned about this between 1975 and 1990 applies without modification, and almost none of it is present in the agent frameworks shipping today.

Which means the dominant industry response — detect the injection, filter the input, train the model to resist — is a category error. You cannot fix a confused deputy by making the deputy smarter. You fix it by taking away the authority it can be confused into using.

Why detection is the wrong layer

Start with the mechanism, because the mechanism explains the futility.

A language model receives one stream of tokens. The system prompt, the user's request, the contents of a retrieved document, the body of an email, the text of a GitHub issue, the output of a tool — all of it arrives as the same kind of thing. There is no channel separation, no out-of-band signalling, no structural marker that says these tokens are commands and those are data. The architecture that makes the model useful is the architecture that makes the attack possible.

You can train the model to be suspicious of imperative text in retrieved content. That raises the attack's cost. It does not change the shape of the problem, because the defender is now running a classifier against an adversary who gets unlimited attempts, pays nothing per attempt, and only has to succeed once. OWASP's work through 2026 places prompt injection at the top of its LLM risk list, and the honest framing across the security literature is not "unsolved yet" but "unsolvable at this layer." The useful question is what happens after an injection lands.

The reason the industry keeps building detectors anyway is worth naming plainly: detection is the only mitigation that preserves the thing that makes agents feel magical. An agent with broad standing access to your mail, your repositories, your calendar and your database can do impressive, open-ended work. Every real containment measure makes it less impressive. Filtering promises safety without that cost, which is why it sells.

Ambient authority is the actual vulnerability

The term for what the 1988 compiler had, and what your agent has, is ambient authority: permission that applies automatically to any request the program makes, simply because of who the program is, without the requester having to present anything.

When an agent holds an OAuth token with repo scope, it does not have permission to touch a repository. It has permission to touch every repository, on every call, forever, and nothing in the call itself distinguishes the one the user asked about from the one an injected instruction named. The token answers "who is calling." It says nothing about "what was this specific call authorised to do."

This is precisely the distinction between access control lists and capabilities. In an ACL system, you name a resource and the system checks whether your identity may touch it — so possession of a name, plus someone else's identity, is enough. In a capability system, the name is the permission: you cannot refer to a thing you were not handed, and handing something over is an explicit, visible act. Designation and authority are the same operation, which is exactly why the confused deputy cannot occur. The compiler could not have been tricked, because it would have had no way to name the billing file that did not come with the right to write it.

The same failure in 1988 and 2026: an untrusted caller supplies the target, while the privileged program supplies the authority. 1988 — THE CONFUSED DEPUTY An ordinary caller with no special privileges. It supplies a filename. caller names a file The compiler holds standing write access to the billing file. It applies that authority to whatever name it is handed. compiler holds standing write authority The billing file is destroyed. The caller never had permission to do this. The compiler did. billing file, destroyed The caller supplies the target. The deputy supplies the authority. 2026 — THE SAME DEFECT A web page, an email, a calendar invite, a GitHub issue, a tool response. Anything the agent reads that an attacker can write. untrusted text names an action The agent holds the user's token with broad scopes. It applies that authority to whatever instruction reaches its context. agent holds your token, broadly scoped Data exfiltrated, a payment issued, a repository modified. Carried out with the user's full standing authority, and indistinguishable in the audit log from work the user asked for. your data, exfiltrated Untrusted text supplies the instruction. The agent supplies the authority. Nothing about this is new. The deputy is just non-deterministic now, and there are more of them.
Hover any box. The structure is identical; only the deputy changed.

What the frameworks actually enforce

If the theory above sounds abstract, a 2026 audit makes it concrete. In Capability Gates Are Not Authorization, David Mellafe Zuvic examined the default dispatch behaviour of LangChain and LangGraph, LlamaIndex, and the Stripe Agent Toolkit, and drew a distinction that most teams have never articulated.

These frameworks give you a capability gate: control over which tools exist in the agent's toolbox. What they do not give you by default is a per-call authorization gate — a deterministic, fail-closed check on the concrete argument values of a specific call. Deciding that an agent may issue refunds is not the same as deciding that it may issue this refund, for this amount, to this recipient, on behalf of this user, right now.

The paper demonstrates an unauthorised payout executing under LangChain's default dispatch, with a matching proof of concept on LlamaIndex. Not a theoretical weakness in an exotic configuration — the default path. The gap between "this tool is available" and "this call is permitted" is where the money leaves.

That gap exists because the frameworks inherited the ACL mental model. They ask who the agent is. The question that matters is what this particular request was authorised to do, and that question has to be answerable without consulting the model, because the model is the component under attack.

The experiment that worked, and the number nobody wants

In 2025, researchers from Google, Google DeepMind and ETH Zürich published CaMeL — capabilities for machine learning — and it is the most important result in this field precisely because it is unglamorous. CaMeL does not make the model resistant to injection. It assumes the model will be fooled and arranges the system so that being fooled does not matter.

The design has two halves. A privileged model sees only the trusted user request and emits code describing the plan. It never reads untrusted data at all. A quarantined model processes untrusted content but holds no tools and can only return structured values against a fixed schema. The plan executes in a custom interpreter that maintains a dataflow graph, tracking the provenance of every value, so a capability policy can refuse an action whose arguments are tainted by untrusted input.

On the AgentDojo benchmark, across 949 attacks, CaMeL recorded zero successful attacks against the next-best defence's eight. That is the headline. Here is the number that explains why your vendor has not shipped it: task completion fell from 84.5% undefended to 77.3%.

Seven points of utility. That is the price of containment, and it is a very good price — paid in the currency nobody wants to spend, because utility is what the demo measures. The authors are careful rather than triumphant, stating directly that prompt injection is not fully solved: side channels remain, data-dependent control flow remains unaddressed, the interpreter's security properties are not formally verified, and the released implementation is a research artefact.

The gap between a published, benchmarked, principled defence and what ships in production is not a research gap. It is a willingness gap.

Why adding a human does not close it either

The reflex at this point is human approval. Put a person in front of every consequential action and the deputy can no longer be confused, because a human reads the request before it executes.

This works at low volume and degrades in a specific, measurable way as volume rises. Automation bias — the tendency to accept machine recommendations without independent scrutiny — is one of the better-replicated findings in human factors research, and it gets worse as the queue gets longer. A reviewer processing hundreds of approvals per shift anchors on the first few, applies less scrutiny as fatigue accumulates, and converges on the agent's output as the default answer. Oversight quietly becomes a rubber stamp while continuing to look like oversight in the org chart.

A 2026 preprint by Emre Turan, Oversight Has a Capacity, models this explicitly: reviewer reliability holds until a capacity threshold and then declines linearly with cumulative load. The conclusion from that model is counterintuitive enough to be worth sitting with.

Modelled rate at which dangerous actions get through, comparing an optimally tuned escalation rate against escalating every action, at three levels of reviewer capacity. DANGEROUS ACTIONS THAT GET THROUGH (LOWER IS BETTER) Purple: escalate selectively, tuned to the reviewer. Grey: escalate everything. Reviewer capacity: 10 Escalating 64% of actions: 56% of dangerous actions still get through. 56% Escalating 100% of actions: 69% get through. Worse, because the reviewer is past capacity on every one. 69% Reviewer capacity: 25 Escalating 64% of actions: 42% get through. 42% Escalating 100%: 57% get through. 57% Reviewer capacity: 50 Escalating 72% of actions: 22% get through. 22% Escalating 100%: 39% get through. 39% Simulation from a stated fatigue model, not field measurement — the direction is the finding, not the digits.
At every capacity level tested, escalating everything was worse than escalating selectively. More oversight, less safety.

Escalating every action was strictly worse than a tuned escalation rate at all three capacity levels. The mechanism is simple once stated: every unnecessary approval request spends reviewer attention that is not available for the request that mattered.

Two implications follow, and the second one is nastier. First, approval is a budget, not a policy — it has a finite size and spending it carelessly makes the system less safe, not more. Second, the same fatigue curve is an attack surface. An adversary who can generate benign-looking escalations can exhaust the reviewer and then submit the real one. Flooding the approval queue is a denial-of-attention attack, and systems that escalate indiscriminately are the easiest to flood.

Treat these figures as a model rather than a measurement — it is a single-author preprint with chosen parameters, and the paper is explicit that reviewers themselves agree only moderately on what is dangerous. The direction is robust. The decimal places are not.

The mental model: stop asking who, start asking what

Every authorization question about an agent reduces to three, and most production systems can only answer the first.

QuestionWhat it establishesTypical state of practice
As whom?Which human principal this action is attributable toUsually a shared service account, which makes the audit log useless
With what authority?What this specific call may do, independent of what the agent could doRarely expressed at all; the token's scopes are the only answer
Provable to whom?Whether a verifier downstream can reconstruct the delegation chainAlmost never; the chain is lost at the first hop

The first row is where most organisations are blocked, and the fix is unglamorous. If every agent action arrives at your systems as svc-automation, you have no attribution, no revocation granularity and no forensic story. The incident review will establish that the automation did it. That is not an answer.

The architecture that contains it

Five mechanisms, ordered by how much they reduce blast radius per unit of engineering effort.

1. Designation carries authority

The structural fix. Do not give the agent a token and let it name targets; give it a handle that is permission for one target, obtained through a path the untrusted content cannot influence.

// Ambient authority: the agent holds broad power and supplies a name.
// An injected instruction only has to change the name.
await github.repos.update({ owner, repo, private: false });

// Designated authority: the caller is handed one capability, scoped
// at the moment of issue. There is no argument an injection can
// rewrite to reach a different repository.
const handle = await issueCapability({
  principal: user.id,          // attributable to a human
  resource: "repo:acme/web",   // bound at issue time, not call time
  actions: ["issues:comment"], // nothing else is reachable
  expiresIn: "10m",
});
await handle.comment(body);

The test for whether you have done this correctly: if an attacker fully controls every argument the agent passes, what is the worst reachable outcome? With ambient authority the answer is the union of your token scopes. With designated authority it is bounded by what you handed over.

2. Attenuation at every hop

Multi-agent systems need delegation, and delegation must only ever narrow. The mechanism for this was published in 2014 by Birgisson and colleagues at Google: macaroons, bearer credentials built from chained HMACs that carry caveats. Anyone holding a macaroon can add caveats and pass it on; nobody can remove one, and no round trip to the issuer is required.

// The orchestrator holds a broad credential and must call a sub-agent.
// It can only hand over something strictly weaker than what it holds.
const sub = parent
  .addCaveat("resource = repo:acme/web")
  .addCaveat("action = issues:read")
  .addCaveat("expires = " + isoMinutesFromNow(5))
  .addCaveat("calls <= 20");

// Third-party caveat: this credential is only valid if an external
// authority will vouch for the condition, checked at verification time.
const gated = sub.addThirdPartyCaveat(
  "https://approvals.internal",
  "ticket = OPS-4821 approved"
);

This is the property that makes agent chains tractable. A sub-agent that gets compromised cannot do more than its parent handed it, and the verifier can see exactly what was narrowed, by whom, at each step.

3. Delegation, not impersonation — and the standard already has this

Here is the part most discussions of agent identity get wrong. The common claim is that OAuth cannot express agent delegation chains. It can. RFC 8693, the OAuth 2.0 Token Exchange specification, draws the distinction explicitly and has for years.

Under impersonation, the RFC says, principal A "is given all the rights that B has within some defined rights context and is indistinguishable from B in that context." Under delegation, A "still has its own identity separate from B, and it is explicitly understood that while B may have delegated some of its rights to A, any actions taken are being taken by A representing B."

Impersonation is what an agent holding a user's access token is doing. It is also what destroys your audit trail, by construction — indistinguishability is the stated property.

The specification provides the alternative. The act claim "provides a means within a JWT to express that delegation has occurred and identify the acting party to whom authority has been delegated," and chains are first-class: "A chain of delegation can be expressed by nesting one act claim within another. The outermost act claim represents the current actor while nested act claims represent prior actors."

{
  "sub": "user:a.khan",              // the human this is attributable to
  "aud": "https://api.internal",
  "exp": 1793040000,
  "scope": "issues:comment",
  "act": {                           // current actor
    "sub": "agent:triage-worker",
    "act": {                         // who delegated to it
      "sub": "agent:orchestrator",
      "act": { "sub": "app:support-console" }
    }
  }
}

Every hop is named. The resource server can authorise on the whole chain rather than on the subject alone — refusing, for instance, any write whose chain contains a worker that touched untrusted content. The primitive has been standardised since 2020. It is mostly unused, and that is an implementation failure rather than a specification gap.

4. A fail-closed gate on concrete values

Between the model's intent and the side effect, put a deterministic check that does not consult the model. It sees the resolved arguments, the principal, the delegation chain and the taint status, and it defaults to refusal.

function authorize(call: ToolCall, ctx: AgentContext): Decision {
  const rule = POLICY[call.tool];
  if (!rule) return deny("no policy for tool");          // default deny

  // Arguments derived from untrusted content cannot drive side effects.
  if (call.taintedArgs.length > 0 && rule.sideEffecting) {
    return deny("argument tainted by untrusted input");
  }

  // Authorise the value, not the capability.
  if (!rule.allows(call.args, ctx.principal)) return deny("out of scope");

  // Spend from a budget that the whole task shares.
  if (!ctx.budget.tryConsume(rule.cost(call.args))) return deny("budget");

  return allow();
}

Two details matter more than they look. Default-deny means a tool added next quarter without a policy entry is unreachable rather than unguarded. And the budget is per task, not per call — the limit that contains a compromised agent is the one on aggregate effect, since nothing stops it from making a thousand individually reasonable requests.

5. Never assemble the trifecta in one context

Simon Willison's framing of the lethal trifecta is the most useful operational heuristic in this area: an agent becomes an exfiltration tool when it simultaneously has access to private data, exposure to untrusted content, and the ability to communicate externally. Any two are survivable. All three in one context is a loaded weapon pointed at whoever can write to your inputs.

Three properties that are individually safe and jointly dangerous: access to private data, exposure to untrusted content, and the ability to send data outward. ANY TWO ARE SURVIVABLE. ALL THREE IS AN EXFILTRATION TOOL. Access to private data: your mail, your repositories, your customer records, your internal wiki. Exposure to untrusted content: anything an attacker can write that your agent will read. Web pages, emails, issues, documents, tool responses. Ability to communicate externally: HTTP requests, sending mail, writing to a public resource, even rendering a remote image URL. danger private data untrusted input a way out Remove any one circle and the injection has nothing to do. Egress is usually the cheapest one to remove, and the one most often left open.
Hover each circle. Note that "a way out" includes a rendered image URL — exfiltration does not require an HTTP tool.

Egress is usually the cheapest of the three to remove and the one most often overlooked, because it hides in places that do not look like network access. A markdown renderer that fetches images will carry data out in a URL. A tool that writes to a shared document hands it to anyone who can read that document. Enumerate the paths by which bytes can leave the context, not the tools that are labelled as network tools.

Deciding how much authority to grant

Autonomy is not a dial from zero to ten. It is a function of reversibility and blast radius, and those are properties of the action rather than of the agent.

Action classExampleAuthority modelHuman involvement
Reversible, boundedDraft a reply, label an issue, propose a diffDesignated capability, short expiryNone — review the artefact, not the action
Irreversible, boundedSend mail, post a comment, merge to a branchDesignated plus per-call value gateSpot-check by sampling, not per action
Reversible, broadBulk relabel, reindex, backfillTask budget with aggregate ceilingApprove the budget once, monitor the rate
Irreversible, broadPayments, deletions, access grants, deploysDelegation chain plus third-party caveatApprove every one — and keep the volume low enough that approval stays real

The bottom row is where the oversight-capacity finding does its work. If that row has high volume, you do not have an approval problem, you have a design problem: the right move is to restructure the work so fewer actions land there, not to hire more approvers.

What this costs, and where it fails

Four honest limits.

You will lose capability, and it will be visible. CaMeL's seven points of task completion is the measured version of a general truth. An agent that cannot reach a resource it was not handed cannot improvise its way to a solution, and improvisation is what impresses people in demos. Expect the containment work to make your agent look worse before it makes it trustworthy.

Capability discipline does not make an agent correct. It bounds what a confused agent can reach. An agent that sends a badly worded email to exactly the right recipient has done something stupid entirely within its authority. Containment is about the ceiling on damage, not the quality of judgment.

Caveats and policies are code, with bugs. A third-party caveat pointing at a service that fails open, a value gate with an off-by-one on a currency boundary, a taint tracker that loses provenance through a serialisation hop — each reintroduces exactly what you were defending against, with the added hazard that everyone now believes the system is safe.

The research problems are real. CaMeL's authors name them: side channels, data-dependent control flow, no formal verification of the interpreter. An agent that reveals information through which of two permitted actions it takes has leaked without violating a single policy. Nobody has a complete answer to that, and anyone selling one should be asked for their threat model.

What to ask before you deploy anything

  1. If an attacker controls every tool argument, what is the worst reachable outcome? If the answer is the union of your token scopes, you have ambient authority.
  2. Which named human is each action attributable to? If the answer is a service account, your audit trail is decorative.
  3. Can a verifier reconstruct the delegation chain? If the chain dies at the first hop, you cannot reason about multi-agent behaviour at all.
  4. List every path by which bytes can leave the context. Include image rendering, logs, shared documents, error messages and analytics.
  5. What is the per-task aggregate budget? Not the per-call limit — the ceiling on total effect.
  6. How many approvals per reviewer per hour does this design generate? If the honest answer is "a lot," your oversight is already notional.
  7. What happens when a tool is added without a policy entry? If it becomes reachable, you are failing open.

Frequently asked

Can prompt injection be fixed?

Not at the model layer, and the reason is structural rather than a maturity problem: instructions and data arrive in the same token stream with no mechanism to separate them. What can be fixed is the consequence. The working strategy is to assume injections land and ensure a landed injection reaches nothing worth reaching — which is an authorization problem, not a prompt problem.

What is the confused deputy problem, and why does it apply to AI agents?

A privileged program is tricked into misusing its own authority on behalf of a less-privileged caller. Norm Hardy described it in 1988 using a compiler that held standing write access and accepted a filename from its caller. An agent holding a broadly scoped token and accepting instructions from retrieved content is the same structure: the untrusted party supplies the target, the privileged party supplies the authority.

Is a service account for my agent acceptable?

Only for work with no user-specific data and no irreversible effect. A shared service account collapses every agent action into one identity, so you lose attribution, you cannot revoke narrowly, and your incident review cannot establish which user's request caused what. Use token exchange to obtain a delegated credential that keeps the human principal in the subject and the agent in the actor chain.

Does OAuth support agent delegation chains?

Yes. RFC 8693 defines the act claim for expressing delegation and states that nesting act claims represents a chain of prior actors. The common belief that OAuth cannot express this is wrong; the primitive is standardised and simply under-adopted. What the ecosystem genuinely lacks is consistent enforcement of chains at resource servers.

Do guardrail and injection-detection products help at all?

They raise attacker cost and catch unsophisticated attempts, which has value. Treat them as a filter, not a control. A detector is a classifier facing an adversary with unlimited free retries who needs to succeed once, so it should never be the thing standing between an agent and an irreversible action.

How do I keep human approval meaningful?

Treat reviewer attention as a budget with a finite size. Escalate selectively rather than universally, route by irreversibility rather than by model confidence, show the resolved arguments rather than the model's explanation of them, and monitor approval latency — when it collapses, oversight has become a rubber stamp and the control is gone whether or not the dashboard says so.

Where should a team start?

Inventory the trifecta. For every agent in production, list what private data it can reach, what untrusted content enters its context, and every path by which bytes can leave. Most teams find at least one agent holding all three, and removing egress is usually a day of work. Do that before anything else on this page.

The short version

The industry is treating a 1988 access-control defect as a 2026 machine-learning problem, which is why the proposed cures keep landing in the wrong layer. No amount of model training resolves a confused deputy, because the deputy's confusion was never the vulnerability. The standing authority was.

The components of the fix are old and well understood: designation that carries authority, attenuation that only narrows, delegation that preserves a provable chain back to a human, deterministic value checks that fail closed, and a hard rule against assembling private data, untrusted input and egress in one context. CaMeL demonstrated that this works against a real benchmark and priced it at roughly seven points of task completion.

That price is the actual obstacle. It is also a bargain, and the organisations that pay it early will be the ones still running agents after the first serious incident makes everyone else turn theirs off.

If you are putting an agent anywhere near a system that holds real data, the authority model is the part worth getting right before anything else. Tell us what you are building and we will walk the trifecta with you. Our work on systems and API integration and on verifying what AI produces approaches the same problem from the other direction.

Frequently asked questions

Can prompt injection be fixed?

Not at the model layer, and the reason is structural rather than a maturity problem: instructions and data arrive in the same token stream with no mechanism to separate them. What can be fixed is the consequence. The working strategy is to assume injections land and ensure a landed injection reaches nothing worth reaching — which is an authorization problem, not a prompt problem.

What is the confused deputy problem, and why does it apply to AI agents?

A privileged program is tricked into misusing its own authority on behalf of a less-privileged caller. Norm Hardy described it in 1988 using a compiler that held standing write access and accepted a filename from its caller. An agent holding a broadly scoped token and accepting instructions from retrieved content is the same structure: the untrusted party supplies the target, the privileged party supplies the authority.

Is a service account for my agent acceptable?

Only for work with no user-specific data and no irreversible effect. A shared service account collapses every agent action into one identity, so you lose attribution, you cannot revoke narrowly, and your incident review cannot establish which user's request caused what. Use token exchange to obtain a delegated credential that keeps the human principal in the subject and the agent in the actor chain.

Does OAuth support agent delegation chains?

Yes. RFC 8693 defines the act claim for expressing delegation and states that nesting act claims represents a chain of prior actors. The common belief that OAuth cannot express this is wrong; the primitive is standardised and simply under-adopted. What the ecosystem genuinely lacks is consistent enforcement of chains at resource servers.

Do guardrail and injection-detection products help at all?

They raise attacker cost and catch unsophisticated attempts, which has value. Treat them as a filter, not a control. A detector is a classifier facing an adversary with unlimited free retries who needs to succeed once, so it should never be the thing standing between an agent and an irreversible action.

How do I keep human approval meaningful?

Treat reviewer attention as a budget with a finite size. Escalate selectively rather than universally, route by irreversibility rather than by model confidence, show the resolved arguments rather than the model's explanation of them, and monitor approval latency — when it collapses, oversight has become a rubber stamp and the control is gone whether or not the dashboard says so.

Where should a team start?

Inventory the trifecta. For every agent in production, list what private data it can reach, what untrusted content enters its context, and every path by which bytes can leave. Most teams find at least one agent holding all three, and removing egress is usually a day of work. Do that before anything else on this page.

#ai-agents#security#prompt-injection#authorization#architecture