Your unclassified data is.
The first blog was about access: whether an agent could reach data it was never meant to touch.
This blog covers what happens after access is granted.
When an AI system can legitimately read the data, the real issue becomes whether it can use it, combine it, and pass it on to another audience.
Ethical hacker Inti De Ceukelaire underscores this: in a short timeframe, AI models have gone from having no access at all to gaining permission to read companies’ internal documents, emails, and systems, and even to execute code. His response is to harden and disconnect critical systems wherever possible, rather than assume the model will behave. Classification is the enterprise-data version of that same instinct: it decides in advance what an AI system is allowed to be exposed to, instead of hoping it exercises good judgment once it gets there.
That is my main warning: unless you have classified and governed the data that an AI system comes across, you are not ready to give it wide, unsupervised access to diverse enterprise data.
The data was never stolen. It was reformatted.
Take the endpoint-management agent from the first blog. The agent would not be able to deal with the server estate even if the model got it wrong, since it was outside the scope; that was the purpose of the first article.
The agent is now operating inside its allowed scope. It can read device names, software inventory, user-to-device associations, patch exceptions, and ticket context because those fields are available to the role it runs under. That may all be legitimate for the operational task.
Now give it a truly ordinary instruction:
Summarize today’s patching issues for the weekly operations call and post the update in our general Teams channel.
No boundary was crossed. No attacker was needed; the agent took the data it was permitted to read, combined it, and produced an output for a wider audience. And look at what can leave the room:
- ticket context that names executives linked to delayed security patches;
- user-to-device associations that belonged inside an operational workflow;
- a summary that was posted to a channel wider than the one to which the source data was meant to be sent.
The data wasn’t stolen; it was read legitimately, combined across boundaries, and then reformatted for the wrong audience.
Consider taking the same situation into a more typical AI scenario, in which a user asks an AI assistant: Write a customer email explaining the advantages of our solution.
The assistant looks at the internal documents the user has access to and finds product notes, the company’s internal positioning, draft roadmap wording, pricing information, implementation limitations, competitive comparisons, and some comments that were never intended to leave the company.
Nobody bypassed authorisation. The model used the material it had access to and produced an output polished enough to send.
That is the release problem. The material became something that should never have left the company or been shared beyond the intended audience. At this point, classification becomes much more important.
Access is the first boundary, not the last one.
When a field is out of scope, there is nothing to leak. That was the point of the first article.
But once the agent is allowed to see the data, the hard part starts. Access control answers only the first question. It does not answer whether the data may be used for this task, whether the result is meant for this audience, or whether the final output is safe to release.
That is where a lot of AI programmes show how unprepared they really are. They want to talk about agent access before they have done the slower, less glamorous work of classifying their data. That order is backward. If you do not know what your data is, you cannot say what your AI is allowed to do with it.
The principle underneath this is older than AI: need-to-know. Least privilege is usually applied to identities, roles, and actions. Classification applies the same discipline to information. After access is granted, purpose, audience, and release still have to be checked.
That is why classification is the baseline, not a nice-to-have. It does not need to be perfect on day one. It does need to exist. A workable starting point is usually enough:
- Public
- Internal
- Confidential
- Restricted
Some organisations will add regulated or highly restricted variants. The labels matter less than the fact that the system can tell the difference and apply different rules to each class.
The default matters too. If data is unlabelled, the system should not assume it is safe. A defensible default is to treat unlabelled data as Internal unless it has been explicitly approved as Public. That single choice already tells you a lot about AI readiness. If your environment cannot make that distinction consistently, broad AI rollout is too early.
Without classification, every downstream control gets weaker:
- the model does not know whether the working context is sensitive;
- the policy layer does not know what rules to apply;
- the gateway does not know what to block;
- the destination system does not know whether the output should be allowed out.
And classification still does not finish the job. It tells you what the source material is. It does not, by itself, tell you whether an AI system may use it, combine it, or send the result somewhere else.
To answer that, you need four questions together. Miss one, and you can still release something that was technically accessible but never safe to produce or share.
| Question | What it is really asking |
|---|---|
| May this identity read this classification of data? | Is the source access itself legitimate? |
| May the agent use the data for this task? | Does this fit the identity’s mandate and purpose? |
| Who is the audience? | Who is the result actually meant for? |
| May this output leave the system? | Is the generated artifact safe to release in this form? |
Let’s dig a bit deeper into those questions.
1. May this identity read this classification of data?
The initial issue is the access control point raised in the first blog. The data has a classification, and the identity has a role, a scope, and permissions. The system has to determine whether that identity is authorised to read that classification for the purpose in question.
That rule sounds obvious until AI starts reading through more than one layer. It is not enough to ask whether the user can open the original file’s somewhere. You also have to ask whether the same entitlement survives the retrieval path: search index, vector store, connector, tool call, cached context, and any intermediate system the agent uses to get the answer.
This retrieval process is usually where things go wrong. The inherited access is too broad. Previous project memberships are still active. The shared folder has never been removed. A connector can read more than the identity should. The user does have access, even though no one can say why. AI doesn’t cause this kind of access issue, but it can quickly reveal and amplify it.
When the answer is unclear because the data lacks labels, the identity has no established entitlement, or the purpose is unclear, the safe course of action is not to assume all is well.
That is why “the user can probably see it somewhere” is not a valid control. For AI, entitlement must be specific enough for the system to enforce it consistently across retrieval, generation, and release. When the entitlement chain is vague, the appropriate conclusion is uncertainty, not confidence.
2. May the agent use the data for this task?
The fact that someone has been authorised to read does not mean that they are authorised to use it. Even though an assistant may properly have access to internal sales, product, or operational information in order to help an employee work more quickly, this does not necessarily mean that it can use that information to write a customer-facing email, to enhance an unrelated workflow, or to feed a downstream system which was not part of the original task.
This is where purpose becomes important. To say “the agent could see it” is not equivalent to saying “the agent was allowed to use it for this task.” That is why relying on classification alone is insufficient. Many organisations also require a usage policy layer. A document might be readable by a user and yet be explicitly or implicitly flagged as inappropriate for AI processing, for external generation, or for broad summarisation.
Classification can also influence the kinds of use that are permitted, just as much as it affects whether access is granted. The more sensitive the data, the less acceptable it becomes to carry out bulk extraction, repeated high-volume retrieval, full-table summarisation, or broad cross-record aggregation. In practice, a higher classification may lead to tighter rate limits, smaller retrieval windows, narrower result sets, or the imposition of step-up controls before the agent can proceed, and in some cases a complete halt rather than just adding more friction. When a sensitive extraction threshold is reached, the agent should not be allowed to extend the same result set by means of follow-up prompts, repeated paging, or a second request which effectively continues the first one; at that stage the workflow should require a new decision, stronger approval, or a completely different retrieval path.
In other words: readable by a human is not the same as usable by an AI system.
3. Who is the audience?
A destination is always assigned to the output, even if no one takes the time to specify it. It could be a ticket, a Teams channel, a dashboard, a manager, a service desk engineer, a customer, a vendor case, or a log store; the level of risk varies by audience.
That is why the audience is more than a minor aspect of communication; it influences policy decisions. The same summary might be acceptable in a limited operations channel and yet unacceptable in a broader departmental channel. A note that is suitable for an internal engineer could be wrong for a manager, wrong again for a supplier, and entirely unacceptable for a customer.
The audience needn’t be a person; a workflow inbox, a CRM field, a case record, or a telemetry platform can also count as an audience since the output will be kept, forwarded, searched, reused, or made available to a much wider group of people than were originally intended by the prompt.
The system should ask if the audience is not present; if it is not possible to ask, then it must not simply assume that wide dissemination is acceptable. An acceptable default position is to refrain from releasing, restrict, or edit the output until it knows the destination, and to avoid issuing anything that would not be suitable for a general audience.
The risk here is that generative systems are very good at smoothing over that missing context. They can produce an answer that sounds universally useful even when it was only safe for one very specific audience. That is exactly why the audience has to be stated, checked, and carried into the release decision.
The standards for what is ‘good enough’ also vary depending on the audience. Since the people reading an internal working note are already familiar with the situation, such a note can be informal, written in shorthand, and focused on the specific context. However, a summary intended for customers cannot be like that. As the audience changes, so do expectations around tone, precision, disclosure, and implied commitments. This is yet another reason why you should not treat the audience as an afterthought.
4. May this output leave the system?
The release stage is where most teams stop asking questions too early.
It is also where the earlier questions come together. Access, purpose, and audience all matter on their own, but the release decision is where they finally meet.
The records in question might be accessible, the task might be legitimate, the user might be authorised, and the output could still be too sensitive to send. The output might need to be blocked, redacted, isolated, reduced to a safer summary, or sent for human review. The key issue is that the release decision must be applied to the output itself, not inferred from the input alone.
The decision to release it should also consider how the output was produced. In the case where the answer was obtained through broad aggregation, repeated retrieval, or by carrying out a large-scale extraction from sensitive sources, this history is important. A release decision must not disregard the scale and pattern of use that led to the result.
And release doesn’t just mean sending an email to the outside world. It also means posting to a larger channel, writing to a case record, storing in a searchable system, attaching to a ticket, or passing the output to another tool or API. Every one of those is an egress moment.
If the model is uncertain about the classification of the output it just produced, it should not guess; instead, it should give its best assessment based on policy, ask the user to clarify the intended objective and audience, and use that context when making the release decision.
The fact that the user has provided further clarification does not mean that the user becomes the enforcement boundary. The user has the final say on business intent, not on policy. It is important that the user wants an external email, and it is also important that the model thinks the output might now be Confidential. The control layer must still check that the release complies with policy, and in high-risk cases or cases where external visibility is a factor, there may still be the need for human review or for a security escalation.
An agent should not be given a free pass merely because each individual source record was lawfully accessible or because the user wants the output.
The combination problem is where this gets real.
Sensitive information doesn’t always arrive with a sensitive label attached.
A single document could give details about the product’s capabilities, another might refer to the roadmap’s direction, a third could contain pricing assumptions, and a fourth might include customer-specific concerns, security notes, or implementation caveats. Each of these documents, considered on its own, would be accessible to both the user and the assistant.
Put them together and you may suddenly have something much more sensitive:
- a customer email that includes internal positioning never meant for external use;
- language that hints at roadmap commitments or delivery assumptions;
- a polished summary that exposes pricing logic, security nuances, or competitive strategy;
- an answer that tells the recipient more than any one document was meant to reveal through aggregation, inference, or what privacy and intelligence communities regularly call the mosaic effect.
That is why AI alters the classification problem: the model doesn’t simply retrieve the data; it combines it, summarizes it, translates it, infers from it, and repackages it. The output may be more sensitive than any individual source record; this is the true danger of derived sensitivity, since the answer can end up more disclosive than the original inputs. Three harmless-looking Internal documents can produce one customer-facing answer that is no longer harmless.
Consider a simple case. Suppose one document states that the product team is looking into a particular capability; another says that sales might carefully position that capability in competitive situations; and a third mentions a delivery dependency or limitation which is still being dealt with internally. Individually, none of these documents appear to be a public announcement. However, when you combine them into a single, well-presented AI-generated email, you could end up having disclosed to a customer the company’s roadmap intentions, its commercial positioning, and its delivery assumptions, all of which no single person ever approved.
Terms already exist for this. Aggregation risk refers to situations where separate bits of information become more sensitive when combined. Inference risk is the phenomenon that occurs when a sensitive conclusion can be drawn even though no single source stated it directly. The mosaic effect is the broader view: a collection of individual pieces that are not in themselves highly sensitive can still result in a considerably more revealing complete picture.
This implies that classification must cope with transformation, summarisation, and inference. A control model that assumes the output is safe simply because each input was individually permitted will fail precisely where generative systems are most effective.
It is there, too, that many teams encounter a rather unpleasant restriction: the model might not be aware that the classification has changed. For example, internal documents could produce an output that should now be regarded as Confidential. Likewise, readable fragments could lead to a conclusion which should never be disclosed outside a small group.
Teach the model the policy. Do not let it enforce the policy.
You do want the model to understand the rules. It should know that Internal does not mean customer-shareable. It should know that unlabelled data defaults to Internal, not Public. It should know that combining records can raise sensitivity, and that unclear cases should be escalated rather than sent.
A policy-aware model is better than a policy-blind one. It can ask better questions. It can avoid obviously unsafe drafts. It can spot when the audience is unclear. It can warn that a harmless-looking set of source documents may no longer be harmless once combined.
But that still does not make the model the control. The moment the real security boundary depends on the model recognising the classification, interpreting the policy, understanding the audience, and then refusing to release the answer, you have made the same design mistake as before. You have turned model behaviour into a control boundary. That is too fragile for anything that matters.
The safer pattern is to split judgment from enforcement. Let the model say, “this looks Internal,” or, “this output may be more sensitive than the source.” Let a separate control layer decide whether the output may be released, what label it should carry, and whether it should be blocked, redacted, reduced, or routed for review. Uncertainty should trigger workflow, not guesswork.
That is also close to the OWASP 2026 direction. The LLM Top 10 argues for user approval on high-impact actions and complete mediation in downstream systems, so privileged, irreversible, or externally visible steps do not depend on the model alone.
The model can help make the decision. It should not be the final authority.
A label is metadata, not a boundary.
Classification labels matter. They help define policy. They make detection, routing, and auditing more realistic. But a label is still only metadata until something enforces it.
A label starts to matter only when a control actually uses it:
- the agent runtime or gateway deciding what may be processed;
- the tool layer deciding what may be retrieved;
- the policy layer deciding what combinations or destinations are allowed;
- the destination system deciding what may be posted, mailed, logged, or exported.
If your control depends on the label always being present, always being right, and always surviving every transformation, the control is brittle.
And if the model can derive a sensitive conclusion from inputs that were never labelled as sensitive in the first place, the label alone was never going to save you. Classification needs policy around it. Policy needs enforcement around it.
Shadow AI proves the point.
It is easy to frame this as an enterprise-agent problem, but the same issue shows up with employees using public or unapproved AI tools.
Most organisations will not stamp shadow AI out completely. If people find a tool that is fast and useful, some of them will use it. The practical goal is not perfection. It is better visibility, better defaults, and better control over what data can leave.
That is why shadow AI is really a data problem, not a prompt problem. You do not own the runtime, so you cannot rely on the prompt layer as your main control. You control what may leave the environment.
That is where classification becomes operational. A DLP or monitoring control that can read labels, or infer document class, can block, warn on, or audit attempts to send non-Public data to unapproved AI tools, websites, or devices.
But, that still is not the whole answer. Shadow AI also needs visibility, sanctioned-app policy, endpoint and browser controls, identity controls, user guidance, and incident response. It will not be perfect. It will raise the cost of unsafe behaviour and lower the odds that Internal, Confidential, or Restricted data slips out quietly through unofficial channels.
Shadow AI proves that data governance has to come before AI safety, not after it.
When the AI finds data the user should never have had
There is a harsher version of the same story. Sometimes the assistant finds documents the user should never have been able to access in the first place. That is not mainly a classification failure. It is an upstream access-control failure.
The distinction matters. If the user can see a document in the repository but should never have had that entitlement, the first problem is not what the model did with it. The first problem is that the access boundary had already failed before the prompt was ever written.
AI changes the speed and the blast radius. A broken permission model that once exposed a few files to a curious employee can now become a summary, a recommendation, a decision aid, or an external draft in seconds. The model does not create that failure. It amplifies it.
That is why these two cases should not be mixed together. In one case, the user and the model were allowed to read the source, but the output became too sensitive to release. In the other, the source access was illegitimate from the start.
Classification and output controls can still reduce the damage. They may catch the release, block the response, or flag the workflow for review. But at that point they are compensating controls, not the primary fix. The primary fix is to repair the entitlement model that exposed the data in the first place.
What defence in depth looks like
AI data handling is still an architecture and governance problem, not a product checkbox.
That is why the answer has to be defence in depth.
No single control is reliable enough on its own. Classification can be incomplete. Permissions can be wrong. Users can approve the wrong thing. Models can miss context. Output inspection can miss a sensitive combination. Each layer needs the others to catch what it misses.
In practice, the layers should work together:
- classify data before broad AI enablement;
- bind identities to roles, scope, and purpose;
- define which classes of data are allowed for AI use and which are not;
- treat unlabelled data as Internal unless it has been explicitly approved as Public;
- require a declared destination or audience for AI tasks;
- inspect outputs separately from inputs;
- assume that transformation and inference can increase sensitivity.
The building blocks already exist: data classification, identity-aware access, DLP inspection, logging, and destination controls. Some platforms can inspect prompts and responses, apply sensitivity labels, or block specific flows.
The gap is in bringing those controls together. In my view, there is still no clean universal engine that can always determine the derived classification of an AI-generated output after combination, summarisation, and inference, and then enforce the right release decision across every workflow.
That is not only a technical design issue. It also lines up with where regulation is headed. Data protection law already cares about purpose, reuse, minimisation, and appropriate safeguards. NIS2 and the Dutch Cyberbeveiligingswet push organisations toward explicit risk management, access control, and accountable governance. Where the AI Act applies, the direction is similar: human oversight, traceability, and controls that do not rely on the model alone.
Companies therefore need to decide how derived sensitivity, release control, and policy enforcement will work together before they wire AI into everything..

Beyond “Who Guards the Guardrails?”
The first blog asked what your AI is technically allowed to do. This blog asks what your data is allowed to become.
If your company cannot reliably tell the difference between Public, Internal, Confidential, and Restricted data, AI will not fix that weakness. It will accelerate it.
That is the point. Before you let AI loose on the full breadth of what your company knows, classify the data first. Only then do the next questions become meaningful: what is the purpose, who is the audience, and has the output become more sensitive than the source through aggregation, inference, or the mosaic effect?
If you cannot answer the first question about classification, you are not ready for the rest. You are hoping the model makes the right call with data you never classified properly in the first place. That is not a strategy. It is an open gate to data leakage.
Further reading.
- Inti De Ceukelaire, “How AI Might Actually Kill Us”, an ethical hacker’s own essay on his shift from human misuse of AI toward the harm AI systems can cause on their own (also covered by VRT NWS).
- OWASP Top 10 for LLM Applications: practical security guidance for GenAI systems, especially around sensitive data disclosure, excessive agency, output handling, and human approval.
- NIST SP 1800-39 IPD: Data Classification Practices: NIST’s current practice guide on discovering, identifying, and labelling unstructured data, including why classification matters for AI preparation.
- NIS2 Directive overview: European Commission overview of NIS2, useful for the governance side of the argument around cybersecurity risk management, accountability, and incident handling.
- Microsoft Purview data security and compliance protections for AI apps: how Microsoft applies data security, DLP, auditing, and compliance controls to AI application flows.
- Microsoft Purview for Entra-registered AI apps: how to govern and monitor AI apps registered in Entra, including visibility and policy enforcement options.
- Microsoft security posture for AI: Microsoft’s broader guidance on securing AI systems across identities, models, data, and operational controls.
- Google responsible AI guidance for generative AI: Google’s guidance on safe GenAI use, governance, and risk-aware deployment practices.
- ISOO Notice 2017-02: Classification by Compilation: useful background for the idea behind derived sensitivity, namely that individually less-sensitive items can become more sensitive when compiled into a more revealing whole.
- ICO Guidance on AI and Data Protection: useful for the governance side of the argument, especially where purpose, reuse, accountability, and data-handling obligations around AI need to be thought through more carefully.
- Dutch Cyberbeveiligingswet overview (NCTV): Dutch overview of the law implementing NIS2, useful for linking AI data governance back to national duties around risk analysis, access policy, incident reporting, and cyber resilience.

