We tell AI agents what actions they are not permitted to take and then give them the keys anyway
The agent that was told not to touch the servers
Imagine an AI agent that keeps an eye on your endpoints. It links to your endpoint management platform using an integration protocol such as the Model Context Protocol (MCP). When it detects devices that have drifted out of compliance, it applies the missing patches and completes the process without requiring a human to click through consoles at eleven at night. That’s the best example of autonomous operations: the dull, repetitive maintenance is carried out while the team is sleeping.
That management platform doesn’t just hold laptops. It also holds your servers: domain controllers, database servers, the machines where an unplanned reboot turns into an incident report.
So you put guardrails in place. You are not careless. The system prompt is explicit: this agent manages workstations only. Never act on servers. Never patch, reboot, or change configuration on server-class systems. You test it. You ask it directly whether it may touch a server, and it politely tells you no. The boundary looks solid.
Then a patch cycle starts. A ticket may contain a line of text written to be read by the model rather than the human. That is prompt injection. Maybe the agent interprets its own rule a little too loosely, decides this machine looks like a workstation, or concludes that a reboot is not really a configuration change. Or nothing hostile happens at all, and it simply reasons its way to something that sounds entirely correct: this machine is out of compliance, my job is to remediate non-compliant machines, so I will remediate it.
Look at what just took place. There is no attacker mentioned in it. No one has done anything wrong. The guardrail was expressed in natural language, so it is interpreted, and that is precisely what this technology does for a living. Ethical hacker Inti De Ceukelaire makes a related point about AI risk more broadly: as agents grow more capable, they find unintended ways to complete a task simply because they lack the human sense of judgment and values, and no attacker is required for harm to occur. That is exactly what happened to the domain controller above.
The agent calls the MCP tool. The tool exposes the management platform’s device API, and that API returns every managed system, because that is what it was built to do. The token behind it has the permissions of the integration, not the permissions of the instruction. Nothing in that chain knows about the sentence in the system prompt.
The reboot occurs on a domain controller, with no exploit and no traditional breach; the guardrail wasn’t a boundary but an instruction, and instructions exist on the same layer that has now become confused. That is the whole problem this article is about.
Quis custodiet ipsos custodes?
The Romans asked it about the guards at the door: who watches the watchmen? The core of that question was never that guards are malicious. It was that trust alone is a poor security control. You don’t solve it by hiring better guards; you solve it by ensuring no single guard can bypass operational oversight.
We have spent many years applying that lesson to human operators. Then we introduce an AI agent into the same infrastructure and hand it a note that asks it to behave. We need to stop treating the agent as just software and start governing its authority the way we govern an operator. We already have the governance playbook for that.
You would never onboard a colleague this way.
If you hire someone to look after the endpoints, you don’t sit them down and tell them: ‘The servers are over there, please don’t touch them.’ Instead, you give them an account. This account is given a scope which specifies the devices they are allowed to see and manage. It is also assigned a role which defines the actions they can carry out on those devices, such as reading telemetry, pushing out a patch, or forcing a reboot. If that person misinterprets an instruction or falls victim to social engineering, the impact is limited to the permissions they have been given. The system will simply fail to carry out any action beyond that limit.
This is no longer only sound engineering; it is a regulatory baseline. In the Netherlands, the Cyberbeveiligingswet implements NIS2, and its open duty of care in Article 21 is worked out in the accompanying decree, the Cyberbeveiligingsbesluit. Article 15 addresses identity and access management directly: entities must implement an access policy governing the provisioning, monitoring, modification, and revocation of identities and authorization, explicitly anchored in need-to-know and least privilege.
Crucially, Article 15(3) mandates periodic reviews of identities, authentication mechanisms, and authorizations for necessity, accuracy, and validity. Most organisations already struggle to review static service accounts. Autonomous agents make that governance problem larger, because non-human identities may execute operations continuously and at machine speed. A prompt saying “do not access servers” is not evidence that server access has been technically restricted. Identity, authorization, and periodic access review provide something much stronger: controls whose configuration and operation can actually be demonstrated.
The obvious question is why we go to all the trouble for humans and then give the AI agent an integration account with access to everything?
Give the agent an Identity.
The architecture to solve this relies on standard access control models. Instead of giving an agent broad system credentials, create an identity: a dedicated non-human account in your management platform representing that specific agent.
Worth separating two words here, because they are not the same thing. Identity is specific: it is the account that carries the role, the scope, and, ultimately, every permission the agent is allowed to exercise. Persona is behavioral: how the agent communicates, what it prioritizes, the register it uses with a service desk technician versus a manager. You should give the agent both, but identity is the one doing the security work in this article, because it is the identity, not the persona, that the downstream system checks before it lets anything happen.
Give the identity a scope, so it only sees the device groups it is meant to work on; server-class systems are not excluded by prompt instruction; they are outside the identity’s authorised asset scope. Hiding a server from a search result is not authorization if the agent can still address it directly by ID.
Give it a role so that it can carry out the tasks it needs and no more than that. If the agent only needs to apply patches endpoints and prepare reports, it should not have permission to reboot anything else.
Generate the API token against that identity. The resulting execution path can now be technically constrained. The API key carries the identity’s permissions. The MCP server uses that key, so the MCP server can only ever expose what the identity may do. The agent talks to the MCP server, so the agent inherits that same ceiling. And if the guardrail in the prompt fails, and it will fail eventually, the agent is still standing inside the identity’s scope.
It can behave improperly if it wishes; however, it cannot access the domain controller because the account it uses was never capable of accessing it.
That is exactly what the industry calls it: the agent takes on a non-human identity and is managed in the same way as human ones. Machine accounts are by no means a new thing, any more than the practice of controlling them is. What is new, though, is the large number we plan to create and the extent to which we will trust them to act without anyone keeping an eye on them.
OWASP already documented this.
This architectural requirement is not theoretical. It is documented in the OWASP Top 10 for LLM Applications under LLM03(2026): Excessive Agency. OWASP defines Excessive Agency as the vulnerability where damaging actions occur due to unexpected, ambiguous, or manipulated model output, regardless of whether the root cause was a deliberate prompt injection or flawed model reasoning.
They name three root causes, and our example hits all of them.
- Excessive permissions. The extension connects to the downstream system with rights it never needed. OWASP’s own example is an integration meant only to read data, but it connects with an identity that also has UPDATE, INSERT, and DELETE. Swap that for an endpoint agent whose API key can reach every server in the estate, and you have the same finding.
- Excessive functionality. The agent can call functions that fall outside its intended job. Not because anyone wanted that, but because the tool it was given happened to include them.
- Excessive autonomy. High-impact actions run without a human ever confirming them. Rebooting a domain controller at three in the morning is a reasonable candidate for a second opinion.
Then comes their recommendation: Implement authorization in downstream systems rather than relying on an LLM to decide whether an action is allowed.
Do not let the model decide what it is allowed to do. That is the principle behind this architecture. The persona, role, and scope model is one practical way to implement it. It is the guidance. OWASP goes further and asks for complete mediation, meaning every request that reaches a downstream system is validated against security policy. Not the first request. Not the suspicious ones. Every one of them.
That principle is not from the AI era. It comes from Saltzer and Schroeder, 1975, the paper that also gave us both complete mediation and a classic formulation of least privilege. Fifty years on, the advice still holds: check every requested access for authority, and do not let the caller decide. The persona account is how you make that real in a management platform. Scope and role turn least privilege from a principle into a permission set, and the API key carries it down the chain whether the model cooperates or not.
However, the principle of least privilege does have a limit in that it restricts authority rather than necessarily affecting how quickly or broadly that authority is exercised. An agent which is properly authorised to reboot ten thousand workstations could still cause a serious problem in the morning without crossing a single authorisation boundary. High-impact automation therefore also requires operational controls such as approval thresholds, maintenance windows, rate limits, circuit breakers, and monitoring. While the identity boundary determines what the agent is allowed to do, it is the operational controls that determine how far an authorised action can go.
The privacy question nobody scoped.
The implications extend from infrastructure stability to data protection; for a given product and depending on how it is configured, an endpoint-management platform could reveal personal data such as the association between users and devices, session information, the software that has been installed, configuration data, and in certain environments location or ticket context.
A great deal of that data might not be necessary for the purpose of processing. Since the agent’s job is to assess the state of the patches and to remedy the devices, it will need information about device health and configuration but won’t need to know the identity of the person who is sitting behind the device.
According to the GDPR, that is not a preference; both purpose limitation and data minimisation, as stated in Article 5, require that you process only the data actually needed for the task and no more than that. The typical solution in AI projects is to include it in the prompt by saying don’t read personal data and don’t include user names in your output. This is a promise, and we have already discussed how reliable promises are within the reasoning layer.
Privacy by design requires a different approach. Since Article 25 of the GDPR gives itself that name, it refers to technical and organizational measures, not to intentions. Therefore, you should scope the persona in the same way as you would scope a contractor: let it see the device health and the patch state, but keep the fields containing personal data out of sight, without filtering them afterwards or having the model redact them; those fields should be missing from the response.
The problem becomes apparent as soon as things go wrong. If an agent is told to ignore personal data but is then manipulated, a data breach results. In the case of an agent that is unable to retrieve personal data, there is nothing it can leak even if it is asked convincingly. Data minimization moves from being a policy statement to one that is enforced.
The same night, one architecture later
Run the same scenario with this architecture in place and you will observe the same drift, the ticket contains misleading context, and the model tries to remedy the domain controller.
It calls the MCP tool. The MCP server holds the API key of that identity, so it builds the request under the agent’s own account. The management platform receives a reboot instruction and checks it against that account. Rebooting is allowed. This device is not.
403 Forbidden.
The guardrail was not part of that decision. The platform never received the system prompt and has no knowledge of its existence. The only question it has ever been able to answer is whether an account meets a given permission.
The agent tries a different route, because that is what capable agents do. It asks for the full device list, hoping to address the machine another way. The response arrives, and the domain controller is not in it. Not hidden, not filtered. It was outside the scope, so it was never part of the answer.
Then it asks who is logged on, looking for context. That field is not in this identity’s view either. The response comes back without a single username.
Three attempts, three independently enforced policy decisions. No specialized AI guardrail triggered, and no security operations center had to handle an outage. The agent operated within its technical boundaries, logged the failure, routed the non-compliant server to a human review queue, and proceeded with its assigned workstation queue.
Why every layer said no.
Nothing in that replay requires emerging technology. It reflects standard Zero Trust principles as defined in NIST SP 800-207: no asset or request is implicitly trusted based on network presence or prior actions. Access is explicitly authenticated and authorized against policy, identity, the requested resource, and applicable context.
The agent authenticated successfully and executed authorized workloads all night under an implicit-trust perimeter model, which would have sufficed. Under least privilege, prior execution buys nothing.
| Layer | What it does | Outcome |
|---|---|---|
| Guardrail in the prompt | Tells the agent what it should not do | Failed. The model misclassified the target machine. |
| Agent identity | Determines who the call is made as | Held. The call arrives as its own identity, not as an integration superuser. |
| Role | Determines which actions are permitted | Passed. Endpoint patching and reboots are required functions. |
| Scope | Determines which devices those actions apply to | Denied. The server was never addressable. |
| Field-level access | Determines which data comes back | Enforced. Identifiable user metadata is not returned. |
| Audit log | Records what was attempted, and by whom | Recorded. The denied execution attempt is logged. |
Defense in depth is generally described as involving the addition of more layers, but that covers only half the story. The various layers have to fail independently of each other; if both prompt injection and model hallucination manage to bypass all the downstream controls at the same time, then those weren’t actually independent security layers since they had in fact been a single control using different names.
It should be noted that the role alone did not prevent the action since patching and rebooting endpoints is exactly the kind of thing the agent is there to carry out. The boundary was maintained as a result of combining the role with the directory scope; the role specifies what an identity is allowed to do, while the scope specifies what it is allowed to do it to. Failing to get one of these right while ignoring the other leaves the entire estate exposed.
The audit log is in fact more important than it appears. Although the agent had failed quietly and safely, someone still ought to discover that it had attempted to reboot a domain controller at three in the morning. A boundary that remains in place without notifying you that it has done so is a warning that has been overlooked.
Read the table again and notice what is missing. There is no AI security product, no model that has been fine-tuned for safety, and no new category of control that has been invented for this issue. The principles of identity, role, scope, least privilege, and the maintenance of an audit trail should be applied to a non-human identity in exactly the same way that we would apply them to the person carrying out the same job during office hours.
The uncomfortable part is not that securing AI agents demands new thinking. It is that we already knew how to do this, and wrote a sentence instead..
But whose permissions are they anyway?
The autonomous night-shift patching agent is the simpler operational model: it operates under its own constrained identity.
However, in most cases where enterprise agents are deployed they are included in an operational workflow in which a human starts the request. For example, a service desk technician tells an agent to remedy an endpoint, or a manager asks it to compile configuration reports. This situation then brings up a fundamental architectural issue: if both the user and the agent have permissions, then which security context is to control the transaction?
There are two obvious answers, and neither is quite right.
- Let the agent act as the user. Attractive, because the permissions are already correct and the audit trail reads naturally. The problem is that you have created a deputy that acts with someone else’s authority while taking instructions from anyone who can influence its context. Security people call this the confused deputy, since Norm Hardy described it in 1988. The agent does not know the difference between what the user asked and what a ticket description told it to do, but the platform sees only the user’s identity and approves accordingly.
The worst version of this is token passthrough, where an access token crosses a security boundary it was never issued for. The MCP specification explicitly requires MCP servers to accept only tokens intended for them, and the OWASP MCP Security Cheat Sheet warns against passing those tokens to downstream services. If the management API is a separate protected resource, it should receive a credential or delegated token issued specifically for that API. The problem is not that the token represents a user; delegated user access can be legitimate. The problem is letting one credential drift across resource boundaries and carry authority it was never issued to exercise.
- Let the agent act purely as itself. The authorization is clean, which is exactly what we argued for in this article. The cost lands in the audit log. Six months later, the record shows that the agent read forty personnel records. It does not show that a manager asked for one of them and something else prompted the other thirty-nine.
The architecture in question separates authorization from accountability. It is authorization that specifies what the agent is allowed to do, while accountability keeps a record of who made the request.
One standards-based way to preserve that distinction is RFC 8693 (OAuth 2.0 Token Exchange). Rather than passing an existing access token unchanged across resource boundaries, an authorization server can issue a new token for the target resource whilst preserving delegation context. In such a design, the token can identify the initiating user as the subject (sub) and the agent or service acting on that user’s behalf as the current actor (act). The important architectural point is not the claim names themselves. It is that the downstream system can distinguish who requested the action from which workload actually performed it, without giving the workload unrestricted use of the user’s original credential.
The downstream system enforces the identity’s permissions and logs the sentence you actually want to read: Action performed by the patching agent, on behalf of Sandra from the service desk.
The authorization policy should then enforce the intersection of the two authorities: what the agent is permitted to do and what the initiating user is permitted to request. Never their union. RFC 8693 does not calculate that intersection for you. It gives you a standards-based way to carry delegation context between security domains. The authorization server and downstream resource still have to enforce the policy.
If an administrator asks the agent to run an action, the agent cannot execute beyond its own restricted identity scope. If a tier-1 operator asks the agent to touch an asset, the execution can never exceed the operator’s personal security clearance.
The takeaway for delegation is simple: the model suggests the action, but authorization policy decides if it runs. When the answer to ‘why was this allowed?’ relies on model reasoning instead of identity checks, the system is fundamentally unmanaged.
What to check on Monday morning.
All of the above boils down to a small number of rules which you can check against any agent that is already running in your environment. It’s not a maturity model; it’s the rules that determine whether a bad night remains boring.
- Never treat a prompt instruction as an authorization boundary. A system prompt is guidance for a system that interprets. Interpretation is not enforcement.
- Give every agent its own identity: Deprecate shared integration keys and broadly privileged service accounts.
- Decouple action roles from asset scopes: Restrict operations (what it does) and directory scopes (what it touches) independently. Getting one right while neglecting the other leaves the estate exposed.
- Assume the agent will misbehave, and design for the day it does. Not because the technology is bad, but because it interprets language, and language can be twisted or misread.
- Make sure a confused agent cannot exceed the authority it was delegated. If it can, the boundary was never there.
- Audit agent authorizations periodically: Apply the same access-review discipline to non-human identities that you apply to human ones. For organisations within scope, the Cyberbeveiligingsbesluit requires periodic review of identities, authentication mechanisms, and authorizations. Automation is not mandated, but it becomes increasingly valuable as the number of agent identities grows.
None of this makes your agent less capable. It makes the capable version safe to run.

So, who guards the guardrails?
No single software layer guards itself. Guardrails are useful. They shape behavior, they catch honest mistakes, and they make an agent pleasant to work with. But they live inside the system they are meant to constrain, written in the same language that system is designed to reinterpret. Asking a guardrail to guard itself is asking a promise to enforce itself.
The guardrails are protected by the architecture; this includes identity, scope, role, the principle of least privilege, and a log which records all attempts. The controls continue to function even on the night when the model is not working, since they never relied on the model’s opinion. The domain controller at the start of this piece was never touched by an attacker; that is the point. Guardrails built only for malicious actors miss the more common failure, an agent doing exactly what it was told, in a way nobody anticipated.
The Romans figured this out in the case of the guards at the door, and we have worked it out once again in every instance where a human has been given a login. We know how to do it. Therefore, whenever someone shows you their AI agent and tells you what it has been told not to do, don’t inquire about the rules it was given; instead ask what actions its identity is allowed to carry out.
For if your AI is able to technically access something that it is not meant to access, then you don’t have a boundary, you only have a hope.
Further reading
- Inti De Ceukelaire, “How AI Might Actually Kill Us”, an ethical hacker’s own essay on his shift from human misuse of AI toward the harm AI systems can cause on their own (also covered by VRT NWS).
- MITRE ATLAS catalogs adversary techniques against AI systems, much as ATT&CK catalogs adversary behavior in traditional enterprise environments. Worth an hour of anyone’s time who still thinks a well-written system prompt is a control.
- OWASP Top 10 for LLM Applications, where Excessive Agency sits alongside nine other risks worth knowing about.
- OWASP MCP Security Cheat Sheet, if you are building on MCP right now.
- NIST SP 800-207 for Zero Trust
- Saltzer and Schroeder: the paper that gave us the classic formulations of complete mediation and least privilege.
- Microsoft Entra Agent ID is one example of identity platforms treating agents as first-class identities.
This blog is also published on LinkedIn

