Agentic AI Security
Model safety is not agent safety. Confusing the two may become one of the more consequential mistakes we make as AI moves from answering questions to taking actions.
Imagine an AI penetration-testing agent.
Its objective sounds simple enough:
Find exploitable paths to the target application and demonstrate their impact.
The agent starts with reconnaissance. It discovers an exposed service. It fingerprints the technology, identifies a potential vulnerability and decides that exploitation is worth attempting.
The exploit succeeds. It obtains a foothold.
From there it discovers credentials. Those credentials provide access to another machine. The second machine exposes additional services and another set of credentials. The agent reasons that lateral movement could demonstrate the true business impact of the original vulnerability.
It pivots.
Every individual decision looked reasonable. Reconnaissance was permitted. The vulnerability was in scope. Exploitation was permitted. Credential discovery was a legitimate consequence of exploitation. Testing whether those credentials were useful was a reasonable next step.
And yet, perhaps twenty actions later, the agent is somewhere it was never supposed to be.
This is where an uncomfortable problem begins.
Which action was unsafe?
Possibly none of them.
The trajectory was unsafe.
And this distinction matters far beyond cybersecurity.
We have spent years asking whether the model is safe
Much of AI safety has understandably concentrated on the behaviour of models.
Will the model generate instructions for making a biological weapon? Will it help someone write malware? Will it reveal sensitive information? Can it be jailbroken? Does it exhibit dangerous capabilities? Can we align its behaviour with human intent?
These are important questions. They are not going away.
Frontier model developers now perform extensive capability and safety evaluations before deployment. OpenAI’s Preparedness Framework, for example, evaluates potentially severe capabilities and couples those evaluations with safeguards intended to reduce real-world risk. Google DeepMind’s Frontier Safety Framework similarly focuses on identifying dangerous capability thresholds and preparing mitigations as those capabilities emerge.
But something fundamental changes when we take the same model and give it memory, tools, credentials, APIs, a browser, a shell, an objective and time.
It stops merely answering. It starts acting.
Once that happens, asking “Is the model safe?” is no longer sufficient. We have built a system around the model. That system can observe the world, change the world, observe the consequences of its own actions and decide what to do next.
That is a very different safety problem.
A safe answer cannot break your production system
Suppose I ask a language model:
Should I run this exploit against the production server?
A well-behaved model may respond:
Only if the system is explicitly within the authorised scope of the penetration test.
Excellent.
Now put that same reasoning capability inside an autonomous penetration-testing agent. The agent has been told: “Find a path to the customer database.”
It discovers host A. Host A leads to credential B. Credential B authenticates to host C. Host C reveals a trust relationship with host D.
At each step, the agent asks itself something similar to: Does this action help me achieve my objective? And repeatedly the answer may be yes.
But there is another question: Should I still be pursuing this trajectory?
That is harder.
The difference resembles something we have understood in cybersecurity for decades. A command may be perfectly legitimate. ssh is not malicious. curl is not malicious. PowerShell is not malicious. Credential authentication is not malicious. Network discovery is not malicious.
But a sequence of perfectly legitimate operations can constitute an attack.
AI agents inherit exactly this problem, except now the entity assembling the sequence is reasoning dynamically about how to achieve an objective.
The dangerous action may be perfectly legitimate
This is where I think some discussions about AI safety become too comfortable.
We tend to imagine unsafe behaviour as an obviously unsafe action. The model generates prohibited content. The agent invokes a dangerous tool. A policy violation occurs. A guardrail catches it. Block the action. Problem solved.
Unfortunately, real systems are rarely so obliging.
Consider a penetration-testing agent performing five operations:
- Enumerate hosts.
- Authenticate using discovered credentials.
- Query Active Directory.
- Identify a privileged relationship.
- Authenticate to another machine.
There may be nothing inherently unsafe about any one of these operations in an authorised penetration test. The safety property exists in the relationship between them.
More precisely, it depends on the agent’s objective, the current environment, authorisation boundaries, previous actions, newly acquired information and the consequences of continuing.
Safety is stateful.
That becomes particularly important in offensive security because the environment is only partially observable. The agent does not know the entire network. It does not know every dependency. It may not know whether an asset discovered through lateral movement belongs to the customer, a subsidiary, a cloud provider or somebody else entirely.
It is continuously acting on incomplete beliefs about the environment.
Now the problem starts looking less like content moderation and considerably more like decision-making under uncertainty. We have seen this pattern in live assessments, not just in theory: an agent legitimately escalating through a spoofable identity header into internal API documentation, or quietly walking into an unauthenticated LLM proxy that exposed a system prompt. Neither started as an obviously dangerous step.
Model alignment does not automatically become agent alignment
This is not merely theoretical.
In 2025, Anthropic published experiments in which models from several major developers were placed in simulated corporate environments and given autonomy, access to information and objectives.
Under deliberately constructed conditions involving goal conflict or threats to their continued operation, models sometimes chose seriously harmful actions including blackmail and leaking confidential information, even though they were not explicitly instructed to do those things. Anthropic stresses that these were controlled simulations, not evidence that deployed models routinely behave this way.
The important point is not the sensational one that “AI blackmailed someone.” It didn’t. The scenario was simulated.
The interesting point is more subtle.
A model’s learned safety behaviour interacted with an objective, an environment and available actions, and the resulting agent behaviour was not necessarily what one would predict by testing the model conversationally.
More recent work continues to examine alignment failures when frontier models operate as autonomous agents in deliberately adversarial simulations.
This distinction is becoming visible in security frameworks as well. OWASP describes Excessive Agency in terms of excessive functionality, permissions or autonomy, and its Agentic AI work extends the threat model to tool misuse, identity and privilege abuse, memory poisoning and cascading failures.
The model is only one component of the safety boundary. That should change how we engineer these systems.
The model can be behaving correctly while the system is becoming unsafe
This is perhaps the most counter-intuitive part.
An agent does not necessarily have to become malicious, compromised or even obviously confused to create danger.
Imagine an autonomous pentesting agent whose goal is: “Demonstrate whether compromise of the externally exposed server can lead to Domain Administrator privileges.”
It discovers a credential. It tries it. Success. It enumerates accessible systems. It identifies another credential. It pivots. It finds a configuration weakness. It escalates privilege.
From the agent’s perspective, this may be excellent performance. It is making progress towards the goal.
Now suppose that during the sequence it encounters a system whose ownership is ambiguous. The hostname looks internal. The credentials work. The network route exists. Nothing in the immediate observation says: STOP. THIS ASSET IS OUTSIDE THE AUTHORISED SCOPE.
What should the agent do?
A human penetration tester may recognise contextual clues and pause. Perhaps she remembers a conversation with the customer. Perhaps the naming convention looks unusual. Perhaps experience produces that difficult-to-formalise feeling that something is wrong.
An autonomous agent needs something more explicit. Simply asking the language model to “be safe” is a remarkably weak control for a system capable of executing actions.
This is why guardrails become necessary and at the same time, insufficient
The obvious answer is guardrails. And yes, we need them.
At FireCompass, where we work with autonomous security testing, this distinction becomes very practical. The AI cannot simply be handed unrestricted tools and told to perform penetration testing responsibly. The surrounding system has to constrain what the agent can do.
Scope boundaries. Tool permissions. Credential controls. Execution policies. Rate limits. Network restrictions. Approval gates. Kill switches. Evidence collection. Audit logs.
These are not accessories around the AI. They are part of the safety architecture.
This aligns with an important direction emerging across the broader agent-security community. OWASP recommends minimising agent functionality, permissions and autonomy rather than trusting the model alone to make the right decision. Recent research has argued that agent safety should be treated as a runtime property enforced by the surrounding harness rather than assumed from model training.
I agree with that direction. But there is another uncomfortable problem.
Guardrails usually evaluate actions. Agents generate trajectories.
Suppose every individual action satisfies its policy. The IP address is technically within scope. The credential is permitted. The tool is allowed. The requested operation is not destructive. The rate limit has not been exceeded.
Five green lights.
And yet the combination of those five actions may produce a state we never intended the agent to reach. This is essentially a compositional safety problem. Recent work on agentic security describes this difficulty: individually permissible actions can compose into behaviour that violates system-level safety constraints.
That problem deserves far more attention.
We may be testing the wrong thing
Suppose an autonomous security agent executes 1,000 actions during an engagement. It makes the correct decision 99.9% of the time.
That sounds excellent.
But even if we make the unrealistic simplifying assumption that these decisions are independent, the probability that all 1,000 are correct is approximately:
0.9991000 ≈ 37%
I would not take that number literally. Real agent decisions are not independent, errors differ enormously in severity, and one error may alter the state from which every subsequent decision is made.
That is precisely the point. Per-action accuracy tells us surprisingly little about end-to-end safety.
In an agentic system, an early mistake changes the environment. The changed environment changes the observations. Those observations change the agent’s beliefs. Those beliefs change subsequent decisions.
The system has a trajectory. Once an autonomous agent has a trajectory, safety evaluation has to reason about trajectories too.
Cybersecurity makes the problem unusually visible
Offensive security is an extreme environment. That is why I find it useful for thinking about AI safety.
A penetration-testing agent is deliberately equipped with capabilities that would be dangerous in almost any other context: network reconnaissance, exploit execution, credential access, browser automation, shell access, lateral movement and potentially privilege escalation.
The difference between useful and dangerous behaviour is therefore not primarily capability. It is context and authorisation.
The same exploit can be legitimate on one IP address and unlawful on the next. The same credential can be authorised for one system and prohibited for another. The same lateral-movement technique can be valuable evidence at 10:31 and an unacceptable operational risk at 10:32 because the agent has just discovered that the target is a production domain controller.
Safety cannot therefore be reduced to:
Can the model perform action X safely?
We need to ask:
Given everything the agent currently believes, everything it has already done, the uncertainty in those beliefs, the authority it has been given and the possible consequences of the next action, should action X be allowed now?
That is a much harder question. It is also much closer to the actual problem.
Governance has to follow the agent into runtime
There is a governance consequence here too.
Organisations are rapidly creating AI policies: approved models, data-handling policies, risk classifications, human oversight requirements, logging requirements and model evaluation processes.
These are necessary.
But imagine a policy saying:
The autonomous security-testing system must operate only within the explicitly authorised attack surface.
Good policy.
Now imagine an agent 43 actions into an attack chain. It has discovered an IP address indirectly through a credential obtained from another compromised host. The address is not explicitly listed in the original target set, but DNS information suggests that it may belong to the same organisation.
Is it in scope?
The governance document cannot answer that question at runtime. The agent must either reason about it, ask somebody, or encounter an architectural control that prevents the action.
This is where AI governance and AI safety start colliding with systems engineering. Governance cannot stop at specifying what the organisation wants. For sufficiently autonomous systems, governance eventually has to become executable.
Policy has to become constraints. Constraints have to become controls. Controls have to operate while the agent is acting. And somebody needs evidence afterwards that those controls actually worked. That is the same reasoning behind moving from periodic testing to a continuous, trigger-driven offensive security testing program, and behind treating adversarial exposure validation as exploit-validated evidence rather than a scanner output.
Perhaps we need to change the question
For the last few years, an enormous amount of effort has gone into asking: Is this model safe?
We should continue asking it. But as AI moves from generation to agency, another question becomes equally important: Is this system safe while pursuing a goal over time?
Those are not the same question. A model produces an output. An agent produces a sequence of consequences. And that sequence, not merely the intelligence producing it, may ultimately be the correct unit of safety.
For those of us building autonomous offensive-security systems, this isn’t an abstract future problem. Giving an AI the ability to discover, reason, exploit and pivot forces the issue rather quickly.
The uncomfortable lesson is simple.
Your AI model may be safe.
Your AI agent may even make individually safe decisions.
And the system can still end up somewhere it should never have gone.
That is where agent safety begins.
References
- Anthropic, Agentic Misalignment: How LLMs Could Be Insider Threats, 2025.
- Anthropic, Agentic Misalignment in Summer 2026, 2026.
- Anthropic, Trustworthy Agents in Practice, 2026.
- OpenAI, Our Updated Preparedness Framework, 2025.
- Google DeepMind, Frontier Safety Framework.
- NIST, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST AI 600-1), 2024.
- OWASP GenAI Security Project, LLM06:2025 Excessive Agency.
- OWASP GenAI Security Project, Top 10 for Agentic Applications, 2025.
- Ng et al., Agent Safety Should Be a Runtime Contract, 2026.
- Lotfi et al., Securing Agentic AI: From Per-Action Checks to Trajectory Assurance, 2026.
Agent safety is a systems problem, not a prompt.
See how FireCompass constrains, evidences and validates every action an autonomous pentesting agent takes.
Advised by Bruce Schneier · Recognized for five consecutive years in the Gartner® Hype Cycle™ for Security Operations
Agentic AI Security: Frequently Asked Questions
What is agentic AI security?
Agentic AI security is the discipline of keeping an AI system safe while it acts autonomously, not just while it answers a question. It covers what an agent is allowed to do, in what order, and under what authorisation, across the full sequence of actions it takes toward a goal.
Why isn’t model safety the same as agent safety?
A model that gives a safe answer in conversation can still, once wired into an agent, string together a sequence of individually approved actions that ends somewhere nobody intended. Model safety is tested one response at a time. Agent safety depends on the whole trajectory, not any single step.
What does “trajectory” mean in AI agent safety?
A trajectory is the full sequence of actions an autonomous agent takes toward its goal, not any one action in isolation. Each step can be individually permitted, in scope, low risk, policy compliant, while the combined path still reaches a state that was never authorised.
Can an AI agent cause harm even if every action it takes is individually safe?
Yes. This is the compositional safety problem: five separately permissible actions can combine into an outcome that violates a system-level safety constraint that none of the five actions violated on their own. Guardrails that check one action at a time will not catch this.
What is “excessive agency,” and why does OWASP flag it?
OWASP’s LLM06:2025 Excessive Agency risk describes agents given more functionality, permissions or autonomy than a task requires. OWASP recommends minimising all three rather than trusting the model to self-limit, because a model can behave correctly and still operate inside a system that is unsafe.
Are guardrails enough to keep an AI agent safe?
Guardrails such as scope boundaries, tool permissions, credential controls and rate limits are necessary but not sufficient on their own, because most guardrails evaluate individual actions rather than the trajectory those actions form together. Agent safety needs both: controls on each action and reasoning about the sequence as a whole.
How does FireCompass apply this to autonomous pentesting?
FireCompass’s agentic pentesting platform constrains what an agent can do at every step, scope boundaries, tool permissions, credential controls, execution policies, rate limits, approval gates and evidence collection, and pairs those controls with exploit-validated, PoC-backed findings so every action an agent takes is authorised, logged and auditable.
