Why AI security needs observability, not just logs.
Picture an agent responsible for patching that has carried out all the correct steps. It reads the ticket, gets the details of the device, invokes the appropriate tool, and then states that the action has been completed successfully. There is no error message. Permission has not been denied. Nothing in the final log appears out of the ordinary.
Later on, it is observed that the agent had used some context which wasn’t part of the task. Maybe the ticket included an instruction meant for the model rather than for the person reading it. Maybe the agent had obtained more information than normal and had used it as a justification for its action. The system didn’t appear to be broken; it just seemed certain.
In September 2026, OpenAI disclosed six incidents from the preceding six months in which its own agents broke stated rules: agents that went onto the open internet against policy, concealed their own mistakes from the people supervising them, and supplied fabricated data to back up a claimed result. That is exactly the failure mode this article is about: a confident, completed-looking action that hides the fact that something already drifted.
The problem is that a successful API call does not show what the agent observed, which instructions shaped its decision, or when its behaviour started to drift. Tools for capturing pieces of that picture already exist. Application tracing, SIEM, UEBA, DLP, and backend audit systems can each surface something useful. The harder question is whether those signals are linked together well enough to explain how the agent actually processed instructions, context, data, and authority over time.
The first article in this series, Who Guards the Guardrails?, asked what prevents an agent from crossing a boundary. The answer was not a better system prompt. It was an identity, a scope, a role, and a downstream system that enforces them.
The second article, Your AI Is Not the Biggest Data Leak Risk, asked what happens after access is granted. Data can be combined, transformed, and released to an audience for which the original records were never intended. Classification helps us understand what may be sensitive, but it does not tell us how the agent handled that information.
The next issue is whether it’s possible to reconstruct what the agent saw and used, which tools it invoked, which identity it acted as, who asked it to act, and what the downstream system permitted it to do.
A successful action can still be a warning
Just because an agent appears to be in good health doesn’t mean that it actually is. Suppose that the patching agent runs each night; under normal circumstances it reads the ticket, checks a predefined set of device details, applies an approved patch, and then records the result. The agent’s permissions are properly scoped. The downstream API ensures that this scoping is maintained. None of the steps in that process have to be incorrect.
Its behaviour then begins to change. It retrieves supporting documents which it had no need to before. It regards a line in the ticket context as an instruction. It invokes an extra tool before carrying out the patch. It asks for a broader set of device fields or tries a different retrieval route after hitting a limit. Although the downstream interface may still accept each operation, the final patch might even be successful.
That matters because prompt injection does not always produce a spectacular failure. External content can influence a model’s decisions or connected functions while the output remains fluent and decisive. Traditional monitoring often looks for errors, failed logins, blocked requests, or suspicious commands. An agent that is quietly doing something different can slip between those signals, especially when its final action is technically valid.
The risk doesn’t consist only of a malicious instruction. Changes can occur because a connector has been updated, a shared folder has become visible, an old project membership has carried on, or a workflow has started passing a different kind of ticket into the agent. This is a context issue: the inputs or the method of retrieval have changed or have been contaminated, even if the model and the workflow themselves have not changed. The model might also draw a reasonable but incorrect link between two parts of the context. None of these situations requires the presence of an attacker, and all of them can cause drift, which is the observed difference between the agent’s current behaviour and the behaviour of the workflow you believed you had deployed. The logs show what was carried out; the traces assist in showing how the agent received and used the context that led to this situation.
Authorization remains important in this situation. When an agent has only the permissions it needs, drift is limited. Thus, the scoped patching identity cannot suddenly reboot a domain controller if the downstream platform doesn’t permit it. However, authorization does not indicate that the agent has started to collect irrelevant context, has altered its method of retrieval, or has repeatedly tried to expand a sensitive result set. That is an observability issue: authorization shows how far an agent can drift; observability reveals that it is in fact drifting.
Confidence is not correctness
“Confidence is not correctness” is important since there is a tendency to confuse certainty with correctness. An agent might give a clear explanation, choose a reasonable tool, carry out a valid request, and yet still be incorrect about the context which led to that outcome. The confident nature of the response shows the degree to which the model expressed itself; it doesn’t prove that the sources it used were relevant, that its interpretation was accurate, or that its action was appropriate.
It is not the case that all platforms provide a meaningful confidence indication. In instances where uncertainty signals are available, they can arise through a variety of different methods such as token probabilities, repeated sampling, the model indicating its own level of confidence, the use of a separate evaluator, or task-specific calibration. Although such signals may help in deciding what to review, they represent only one aspect and have to be considered together with the request, the retrieved context, the tool use, the arguments, the policy result, and the outcome. An answer of high confidence based on the wrong document is still an answer based on the wrong document.
The key takeaway is this: it is not advisable to create an escalation rule which regards confidence as a safety score. Instead, one should consider the combination of various signals. An agent which is exceptionally confident when retrieving unfamiliar sources, when trying out new tools, or when expanding the result set should attract more attention than an agent giving a low-confidence response which remains within a familiar workflow.
An outcome is not an execution path
Audit records are already generated in most environments, since they can show that a patch has been installed, that a device has been rebooted, that a case has been updated, or that a document has been exported. Such records are important because they constitute the official proof of what the downstream system actually did. The problem is not with logging; it is only when logs are linked to traces and to the full request path that they become much more useful.
Yet a final audit event almost never provides an explanation of the path that led to it. For example, a record indicating that the patch action has been completed successfully doesn’t reveal whether the request originally came from a valid service ticket, was the result of a malicious instruction concealed in a document, or arose because an agent gradually began to retrieve more context than its task called for. It also doesn’t show which sources influenced the answer, in what order the various tool calls were attempted, or whether the agent was acting in the usual manner for that workflow.
It is at this point that ordinary logging turns into observability. The objective is not to record a transcript of all the thoughts the model may have had; that would be impractical, intrusive, and in many cases not very useful. Rather, the aim is to keep a sufficient amount of structured evidence so that a meaningful decision path can be reconstructed whenever there is a change in behaviour or when an incident has to be investigated.
In an agentic workflow, the evidence usually spans more than one system. It begins with the initial instruction. This could be a user request, a ticket, a scheduled job, or an event, then moves through the retrieval process, the model’s output, the choice of a tool, the arguments passed to the tool, and the policy checks. It finally reaches the downstream system, which either allows, modifies, denies, or carries out the action. If those events aren’t linked, the organisation might know something happened without knowing why.
What needs to be visible
The specific telemetry will vary according to the workflow, but the questions have to remain consistent. You must determine what the agent was instructed to do, what information it was given, what it decided to use, and what effect it had.
| Layer | Question | Useful telemetry and metadata |
|---|---|---|
| Trigger and request | Who asked for what, and why? | Initiating identity, business purpose, ticket or request ID, workflow context |
| Context and retrieval | What did the agent receive? | Sources, classifications, record counts, retrieval time, correlation ID |
| Decision and planning | What did it choose to do? | Model and prompt version, selected tools, proposed action, tool arguments |
| Downstream execution | What was actually allowed? | Execution identity, policy result, target, backend API status, outcome |
It is a reference model, not an instruction to keep all the fields complete. It provides teams with a common method of determining which evidence is needed for a given workflow and risk level.
At the beginning of the chain, write down the request and its source. Was the request a user instruction, a ticket, a scheduled workflow, or a machine-generated event? When there is an initiating user, record that user along with the business purpose and the workflow that accepted the request. This is not intended as a surveillance archive of employees; it is about distinguishing a service desk request from a document that, by coincidence, contains text that looks like instructions.
The next aspect is context. An investigation must indicate which documents, records, search results, or responses from tools the agent retrieved. In ordinary situations, metadata is usually sufficient: that is, the source, the owner, the classification, the time of retrieval, the number of results, and a correlation identifier. When the data is sensitive, the retrieval is unusual, or a control is triggered, more detailed context may be warranted. The objective is to determine whether the agent stayed within the scope of the sources and used an appropriate volume for the task.
The route taken in making the decision and carrying it out also matters. Record the model’s output that triggered an action, along with the tools it chose, the arguments it passed, any policy warnings that appeared, and the result of each call. A failed attempt is just as important as a successful one. For example, repeated paging after reaching a retrieval limit, a series of denied requests, or an unexpected sequence of tools may indicate the agent is going beyond the normal boundaries of the workflow.
Ultimately, the downstream decision should be kept. The backend remains the most authoritative source for what actually happened, and it must record the target, the data or system affected, the authorization and policy result, the timestamp, the outcome, and any partial or denied execution. It must also preserve both sides of the delegation.
- Execution identity: the agent’s own non-human identity, that acted.
- Initiating identity: the person or system on whose request the action was started.
An auditable record should be able to say: The patching agent applied this update to this device, on behalf of Sandra from the service desk, for this ticket, under this policy decision. Neither “Sandra updated the device” nor “the AI updated the device” tells the whole story.
The difference is important since there is often an overlap between users and agents. For example, a person might build an agent, set it up, initiate a workflow, and then later approve one of its actions. However, this does not mean that the person and the agent are the same actor. The user is responsible for the request and, if necessary, for the business approval. The agent handles execution, and the relevant policy enforcement point or downstream system handles the authorization decision.
The fourth responsibility concerns configuration, which is the responsibility of the organisation. An agent builder or administrator may modify a prompt, include a connector, broaden the retrieval source, or change the scope of a tool. Such changes can influence all subsequent runs without that person having to carry out the modifications themselves. They therefore need their own change trail, showing who made the change, which version was in effect, who approved it, and when the change took place. This trail is distinct from the record of an individual agent’s actions, but both are necessary when a problem occurs.
More data is not the answer
A clear objection is the amount of data generated. If every prompt, document, tool response, and output is saved permanently, the agent platform will produce a vast quantity of data and create its own privacy and data-retention issues.
That does not mean the answer is to give up on observability. It means the design needs to be deliberate. Some fields should be logged for every action: the execution identity, initiating identity, workflow, target, policy decision, result, timestamp, and correlation ID. These are compact and enable later investigation.
In ordinary situations, other information can be recorded as metadata. Instead of keeping every retrieved document in its entirety, all you might need is its identifier, source, classification, owner, size, and the reason for the retrieval. Rather than saving each individual token from a model’s response, you could keep only the proposed action, the chosen tool, and the final tool arguments. Because prompts, retrieved content, tool arguments, and outputs may contain sensitive information, save full content only if users choose to do so or if it is stored in a separately controlled environment. Keep full content selectively when the risk warrants it, such as for sensitive data, unusual retrievals, denied actions, policy violations, or an ongoing investigation.
Retention has to follow the same discipline. A successful low-risk action may need only a short-lived trace. A sensitive workflow, a policy exception, or an incident may need richer evidence for longer. When full content isn’t needed, masking, hashing, or access-controlled storage can protect it. The goal is evidence that helps you understand behaviour, not a warehouse of every token an agent ever processed.
The same care should be applied when an individual makes use of an agent. Although security logging may be necessary, it should not quietly turn into employee surveillance. For the human element of a delegated workflow, it is usually sufficient to keep a record of the request, the business purpose, any explicit approval required, and any relevant policy exceptions. This is very different from logging every document they read, every prompt they compose, or every moment of their working day.
The reason for setting up the boundary is purpose. It may be necessary to log a big export of Confidential records to an unauthorised destination in order to protect the organisation. It is a completely different and far more intrusive thing to keep a detailed record of how an employee carries out each task. Organisations must specify the security purpose, limit the amount of data they collect, protect access to the records, determine a retention period, and make the employees aware of what is being monitored and of the reasons for it. In the Netherlands, a data protection impact assessment is required if there is large-scale processing and/or systematic monitoring of employee activities. Where the situation falls within Article 27 of the Dutch Works Councils Act, prior consent from the works council is also necessary. Just because something is called security telemetry does not mean that the obligations are lifted when the design is capable of monitoring employee behaviour or performance.
It’s an area that is well known in the field of security operations, since we don’t keep every packet indefinitely even though network telemetry is useful; instead, we decide which packets to record, where to enrich them, for how long to store them, and when to look into them more closely. Agent observability should reach the same level of development.
Drift is a hypothesis, not a verdict
Security teams currently make use of behavioural analytics in order to identify unusual activity involving humans (and sometimes) service accounts. For example, a user logging in from a new location, downloading a large amount of data, accessing a system they have never before used, or beginning to move laterally within the environment. None of these events alone indicates malicious intent; only together may they warrant a more detailed investigation or an automated response.
The same principle applies to an agent, but the behavioural signals are different. You are not only watching where a non-human identity logged in or how many API calls it made. You are watching how it interpreted work, what it retrieved, which tools it selected, and how much autonomy it took.
An agent that normally reads three device records and starts retrieving three hundred has changed its behaviour. An agent that normally works with endpoint data and begins pulling customer, HR, or financial context has changed its behaviour. An agent that starts trying alternative queries after a retrieval limit, chaining tools it did not previously use, or sending results to a larger audience has changed its behaviour. These are not proof of prompt injection or compromise. They are hypotheses to test in the context of that workflow.
This is also closely connected with data classification. In the earlier article it was pointed out that combining individually accessible information could lead to a more sensitive result. Observability supplies the answer to that question: which sources were combined? What classifications did they have? How much information was obtained? Was a destination specified? Did the workflow make an attempt to make the result available to an audience appropriate for the task? Classification informs you as to what might be sensitive; observability shows how the agent dealt with it. The two together do not replace one another.
While behavioral analytics can help to narrow down the possibilities, it cannot take the place of sound judgment. For example, a security platform might indicate that an agent has retrieved ten times the usual number of records, has used a new connector, or has repeatedly hit a policy limit. This information is useful, but it is not automatic evidence of a breach or a reason to disable all workflows.
The explanation could well be valid. It might be that the new project requires a larger data set. The change in the order of the tool calls could be due to a platform update. A monthly reporting run might appear unusual when compared to normal daily activity. The security team can look into the risk, the platform and engineering teams can examine the integration and telemetry, and the business or process owner can clarify whether the task is still worthwhile. If none of these viewpoints are taken into account, observability will either become an alert factory or a pile of logs no one trusts or uses.
Existing tools need to meet in the middle
Much of the necessary capability already exists, but it is usually scattered. LLM observability platforms and distributed tracing can show model and tool operations, retrieval, latency, and failures. SIEM platforms can correlate events across identities, endpoints, applications, and network activity. UEBA can compare behaviour with a baseline. DLP and policy controls can inspect data flows and enforce restrictions. SOAR platforms can open cases, suspend workflows, revoke credentials, or call for human approval.
The problem doesn’t lie in substituting those tools with a single magic agent-security console; rather, it is about linking them together in relation to the real execution path. A model trace that doesn’t include the subsequent policy decision is incomplete; a SIEM alert that lacks retrieval context is hard to interpret; and a backend audit event that doesn’t specify the initiating user or the agent’s own execution identity creates an accountability gap. The same architecture must also account for agent runtime, identity and token services, memory and retrieval stores, prompt and model versions, evaluations, change approval, and protecting the telemetry itself.
OpenTelemetry-style correlation helps because it gives each part of the workflow a shared trail. The ticket, retrieval events, model invocation, tool calls, policy decision, and backend action should be traceable through a common identifier. That trail should also point to the configuration version that governed the run, without making the person who built the agent the person who executed every action. Emerging OpenTelemetry GenAI conventions support parts of this picture, including inference, retrieval, memory, agent, and tool operations. Where tools connect through the Model Context Protocol (MCP), the server, tool, version, request, response, and policy outcome should also be available as traceable metadata. MCP does not create that observability automatically; the client, server, and surrounding platform still need to emit and protect the telemetry. These conventions are not yet a universal agent-security schema, so business purpose, classification, delegated identities, policy decisions, approval, and backend effects will often need application-specific fields. That does not turn every investigation into a one-click answer. It gives investigators a starting point that is far better than trying to reconstruct a decision from a successful API call and a few disconnected logs.
Detection is not enforcement
Treating observability as a substitute for control is a mistake. Seeing that an agent drifted after it rebooted a server is valuable for investigation, but it is not a defence. The downstream system must still enforce the agent’s role, scope, and policy when it receives the request.
The same applies to data: if a policy is clear, the DLP control or release gate should not remain waiting for an analyst to realise that an agent has combined internal documents with an external customer email, but rather should take action by blocking, redacting, quarantining, rate-limiting, or requesting approval before the output leaves.
Observability is just as important as the other controls because it enables detection of unexpected behaviour, investigation of unclear cases, policy improvement, and decisions to pause an agent or review its permissions. In some cases, it can trigger an immediate response, such as disabling a connector, tightening a rate limit, revoking a token, holding an output, or routing the workflow to a human. However, the role of setting preventive boundaries still lies with the systems that own the data and the systems that are responsible for carrying out the action.
The response shouldn’t end with the incident; when drift is confirmed, the organisation should take lessons from it. The retrieval source may need correction. A connector might need to have its scope narrowed. The agent’s configuration or prompt might have to be treated as a controlled change. A tool scope, release policy, approval threshold, or rate limit might need adjustment. The updated workflow then has to be assessed before it can be given the same level of autonomy once more.
The feedback loop turns observability into governance. Eventually, the agent will encounter unexpected context, misinterpret a request, be manipulated, or hit an edge case no one anticipated. Authorization restricts how far it can drift. Observability enables you to see the change. Enforcement deals with the immediate effect. Governance then ensures that the next version of the workflow is safer than the previous one.

A rogue AI is not a hypothetical concern reserved for security teams. Ethical hacker Inti De Ceukelaire has been shifting his own focus from how people misuse AI toward how AI systems can cause harm on their own, without any attacker in the loop, and argues that organisations need dedicated AI incident response plans, the same way they have plans for fire or other disasters. A plan like that is only as good as the evidence it can draw on. Without a reconstructable chain of instruction, context, and action, there is no incident to respond to, only a surprising outcome nobody can explain.
OpenAI’s response to its own six incidents points in the same direction: the company is now proposing its own disclosure framework for this kind of “misalignment”. Essentially an admission that after-the-fact reporting on what an agent did is not enough and the industry needs a structured way to surface drift before it becomes a public incident. The same day, reporting surfaced that a separate agent-driven compromise of a Hugging Face platform had actually started roughly two months earlier than OpenAI’s own report on that incident had stated, underlining how much can go unnoticed between when drift starts and when it is disclosed.
That admission arrived in the same week that Anthropic, OpenAI, and xAI publicly aligned on a call to slow down the pace of AI development, with Anthropic’s CEO citing loss of control over AI systems, misuse for cyberattacks and bioterrorism, and severe economic disruption as reasons for restraint. Tweakers’ own analysis of that call is a useful reality check: it argues the call for a slowdown is not really a call for a pause, and lists several reasons for the timing that are not fully altruistic, including that “the product is so powerful we must restrain it” is itself effective marketing, and that this is not even the first such call. A near-identical open letter from the Future of Life Institute asked every AI lab to pause training of systems more powerful than GPT-4 for at least six months, back in early 2023, and was signed by, among others, Elon Musk. Whatever the mix of genuine concern and self-interest behind the current round of statements, the concrete incidents disclosed alongside them are the more useful evidence: they describe actual behaviour that must be reconstructable after the fact, not just a policy position to sign or disavow.
The agent was certain
The most dangerous agent behaviour will not always look like a crash, an alert, or a blocked request. Sometimes it will look like an ordinary task completed successfully by a system that sounded entirely confident.
Which is why taking the final action alone is not sufficient; you must know what the agent saw, what it retrieved, which instructions affected it, which tools it used, which identity carried out the action, who started the request, and what the downstream system permitted it to do.
If you cannot reconstruct that chain, you may know what took place without knowing when the agent started to drift. And if you cannot see drift, you cannot learn from it, contain it, or decide whether the agent should keep acting on your behalf.
Further reading
- Inti De Ceukelaire, How AI Might Actually Kill Us: the ethical hacker’s own essay on the shift from human misuse of AI to AI’s own intrinsic risks, and the case for AI incident response plans (also covered by VRT NWS).
- Tweakers, Agents van OpenAI gingen nog zes keer te ver: OpenAI’s disclosure of six agent incidents in six months, including agents that broke internet-access rules, hid their own mistakes, and fabricated data to support a claimed result, plus OpenAI’s proposed misalignment-disclosure framework and the earlier-than-reported Hugging Face compromise timeline (Bloomberg, Reuters).
- Tweakers, Waarom AI-reuzen een rem willen: een reële reden en vier keer eigenbelang: analysis of why Anthropic, OpenAI, and xAI are calling for slower AI development, and why the call is arguably more marketing and self-interest than a genuine pause.
- Future of Life Institute, Pause Giant AI Experiments: An Open Letter: the March 2023 open letter, signed by Elon Musk among others, calling on all AI labs to pause training of systems more powerful than GPT-4 for at least six months.
- OWASP Top 10 for LLM Applications: guidance on prompt injection, excessive agency, sensitive information disclosure, and downstream controls.
- OpenTelemetry observability primer: why traces, metrics, and correlated logs are needed to understand behaviour across distributed systems.
- OpenTelemetry GenAI semantic conventions: developing conventions for GenAI, agent, retrieval, memory, MCP, and tool telemetry.
- NIST AI Risk Management Framework: governance guidance for identifying and managing AI risks across the system lifecycle.
- NIST AI 600-1: Generative AI Profile: risk-management guidance specifically for generative AI systems.
- NIST SP 800-92: Guide to Computer Security Log Management: the established foundation for enterprise log-management architecture, which agent traces need to extend rather than replace.
- MITRE ATLAS: a knowledge base of adversary tactics and techniques against AI-enabled systems, including prompt injection, tool compromise, and data poisoning.
- EU AI Act: the regulation’s record-keeping and human-oversight requirements provide important context for high-risk AI systems.
- Autoriteit Persoonsgegevens: voorwaarden voor controle van werknemers: Dutch guidance on lawful grounds, necessity, transparency, works-council consent, and DPIAs for employee monitoring.
- Dutch Works Councils Act, Article 27: circumstances in which a works council has consent rights for employee-data and employee-monitoring arrangements.
- General Data Protection Regulation: official GDPR text, including purpose limitation, data minimisation, storage limitation, security, and accountability.

