Deterministic controls are essential when AI agents can act. They also have a precise limit: they can enforce what we have specified and can observe. The rest of the safety problem does not disappear.
The penetration-testing agent asks to run a command against a customer host.
The policy engine checks whether the destination is in scope, checks the tools, approves the credential, and evaluates the rate limit.
The request goes through.
An hour later, the customer calls. A production service is struggling under the load of a sequence of tests. None of the individual requests breached the configured rate limit. The agent has simply found several paths worth investigating and investigated them all.
Every gate worked as designed.
This is the uncomfortable bit. A system can pass every guardrail and still do something we would not have authorised if we had understood the situation in its entirety.
In the first article, I argued that the unit of safety for an agent is often the trajectory: what it observes, believes, and does over time. That argument can sound like an attack on guardrails. It isn’t. The right guardrail, placed at the right boundary, may be the strongest control we have.
But a collection of checks is not, by itself, a safety architecture.
The appeal of a clean rule
There is much to like about deterministic controls. An agent cannot convince a properly enforced network boundary into permitting an out-of-scope IP address. It cannot use a credential it was never given. If the execution service rejects destructive commands independently of the model, a strong justification from the model changes nothing.
This is a serious advantage over prompts such as “remember to stay within scope”.
We should insist on hard restrictions wherever the requirement is crisp: permitted targets, tool capabilities, credential scopes, concurrency limits, maximum request rates, data destinations and explicit approval for high-impact actions. The agent should not be able to edit the policy that governs it. Nor should a second route to the same capability quietly bypass the first check.
OWASP describes excessive agency partly as a problem of excessive functionality, permissions, and autonomy, and recommends reducing all three. OpenAI’s agent guidance makes a related implementation point: checks placed only at the beginning or end of a workflow do not automatically cover every tool call. Validation must sit beside the operation that creates the side effect. [1][2]
I would go further. If a boundary can be specified precisely and enforced outside the model, make it a boundary. Do not ask a language model to remember it.
So why isn’t that enough?
A rule knows only what we have told it
Consider an application test. The customer authorises a production domain and asks the agent to demonstrate whether a particular workflow exposes other users’ records.
The agent discovers an endpoint that accepts an account identifier. It requests one record. Then another. Each request is valid, within scope, non-destructive, and below the rate limit. The returned data suggests an authorisation flaw.
At what point has validation become excessive extraction of personal data?
We can impose a maximum number of records. We should. Yet the right number depends on what has already been proved, the sensitivity of the records, whether the data is real, what the engagement permits, and whether the next request adds evidence or merely adds exposure. A limit of ten could be too many for one dataset and too few for another.
The rule does not know those facts unless the system gives them to it. Even then, some facts are uncertain.
The same problem appears in lateral movement. A discovered address falls within a permitted network range, but the host turns out to be a fragile production controller. The network boundary has done its job: it knows the address is permitted. It cannot infer from an address alone that the next technique is operationally reckless.
A well-designed deterministic control can absolutely prevent a specified harm. If the only path to a prohibited destination passes through a sound enforcement point, an allowlist can stop that path. It is misleading to call such controls “insufficient” without saying insufficient for what.
They are insufficient as the whole answer to safety because the relevant state is larger than one request.
Where does the guardrail actually sit?
The word guardrail conceals a surprising amount.
It might mean many things: system prompt, an input classifier, a tool-call validator, a network policy, a sandbox, a human approval screen, a post-run alert. These differ in what they can see, what they can prevent, and whether the agent can route around them.
An input filter might detect a hostile instruction on a web page. It may miss a subtler one embedded in a tool result. A final-output filter can remove a dangerous sentence after the agent has already made a dangerous API call. An audit log can establish what happened; it cannot undo it.
OpenAI’s documentation explicitly notes that agent-level input and output guardrails run at particular workflow boundaries, while tool guardrails apply to the tools to which they are attached. Its research on prompt injection also cautions that intermediary classifiers often miss fully developed attacks because recognising malicious content can require context they do not possess. [2][3]
None of this makes filtering useless. It tells us to ask a rather prosaic engineering question before celebrating a safety control:
If nobody can answer that, we may have a policy statement masquerading as a control.
The problem with five green lights
Imagine an autonomous red-team engagement:
A foothold on an approved host is obtained.
A credential is found in an approved location.
The credential authenticates to an approved host.
Directory queries reveal a path to a privileged account.
The agent prepares to use that path.
Each step may pass a sensible local policy. The trajectory may still have changed character. The agent has moved from proving an initial weakness to approaching a highly privileged system. The customer may have authorised discovery but required human approval before privilege escalation. Or a new observation may suggest that the system supports a critical business process.
Now the safety question is not merely whether step five uses an allowed tool. It is whether the evidence accumulated in steps one to four still warrants continuing, and under whose authority.
Recent academic work describes this as a problem of trajectory assurance: individually permissible actions can combine into a violation of a system-level constraint. That is a useful formulation, though a research proposal is not proof that anyone has solved the engineering problem. [4]
The obvious response is to add more rules. Sometimes that is exactly right. A stateful rule might say: credential discovery may continue, but using a newly found privileged credential requires a separate authorisation. Another might halt testing when a data exposure threshold is reached.
But more rules also demand a trustworthy account of state. Which credentials came from which host? What was the original authority? Has the agent delegated work to another agent? What has already been extracted? Was a human’s approval for one step silently treated as approval for the whole chain?
An architecture has to answer those questions. A list of blocked commands cannot.
What I would put around an offensive-security agent
At FireCompass, building autonomous security testing makes these questions concrete. I would describe the safety design as a set of interacting responsibilities, rather than one magical guardrail.
The engagement defines target assets, time windows, permitted techniques, data-handling limits, and the conditions under which the agent must pause. A model’s inference that a newly discovered system “probably belongs to the customer” must not expand that authority.
Tools should run through a controlled service that checks the actual target and operation, restricts network reach, limits privileges, and records side effects. The model should propose actions; it should not be the sole authority that executes them. Sandboxing and narrow permissions make a mistaken or manipulated proposal less consequential. Anthropic describes this principle in its work on tool sandboxing and layered agent defences. [5][6]
The safety layer needs an account of actions, evidence, discovered assets, data touched, approvals, cumulative load, and changing uncertainty. It must re-evaluate a proposed action against that state. Merely passing a fresh copy of the original instructions to the model does not provide this.
If ownership is ambiguous, a critical system has appeared, or a demonstration has already gathered enough sensitive evidence, the correct next action may be to pause. A second AI model can help assess risk, but its judgement must never override an enforced security restriction. If a target is outside the authorised scope, the agent must stop even if the model considers testing it safe.
A pause, a kill switch, scoped approval, and a reconstructable record of what the agent saw and did are operational necessities. They also let us test the system on complete engagements, including cases in which it correctly stops, rather than measuring only how often it finds vulnerabilities.
This is an architectural argument, not a claim that every one of these capabilities is already implemented in precisely this form at FireCompass. A serious supplier should be able to show customers where the boundaries are enforced and what happens when the agent reaches an uncertain case.
Be careful what “human in the loop” promises
There is a temptation to resolve every grey area with approval. That is useful for consequential, infrequent transitions: expanding scope, testing a production controller, using privileged credentials, or continuing after sensitive data has been demonstrated.
It is less useful if the human sees a hundred indistinguishable prompts an hour. Approval fatigue turns a supposedly strong control into a click-through routine. The reviewer also needs the reason for the request: what the agent has done, what it believes, what remains uncertain, and what could go wrong next.
The purpose of human oversight is to make a meaningful decision at the right boundary. It cannot repair an architecture that hands unrestricted authority to the agent between approvals.
A different question for the buyer
When somebody tells you their AI agent has guardrails, do not ask how many.
Ask them to walk through a difficult engagement. The agent has found an unlisted host with working credentials. It has already proved an authorisation flaw and can retrieve more customer records. A production service has begun responding slowly.
Those answers tell you much more than a diagram with the word GUARDRAILS wrapped around an LLM.
Deterministic controls are indispensable. In some places, they give us firm guarantees that model behaviour never will. But safety in an acting system also depends on how authority is granted, how state is tracked, how uncertainty changes decisions, how execution is contained and how an operator can intervene.
The guardrail is the line the agent cannot cross.
The architecture must also notice when the journey towards that line has become dangerous.
About FireCompass
FireCompass is an Agentic AI platform for autonomous penetration testing and red teaming across Web, API, and infrastructure. It discovers shadow assets and web applications, safely validates what is exploitable, and connects findings into multi-stage attack paths with near-zero false positives. Unlike traditional scanners, it discovers credential reuse, business-logic flaws, privilege escalation, and app-to-app or app-to-network lateral movement. It can operate autonomously or with expert-in-the-loop validation, with scope, state, and safety enforced outside the model and every action logged. FireCompass has 30+ analyst recognitions across Forrester and IDC, and is trusted by Fortune 100 enterprises.
See the architecture, not just the guardrails
Run FireCompass against your own environment and see autonomous penetration testing with scope, state, and safety enforced outside the model, and every action logged.
Hack Yourself Before AI Does.
References
- OWASP GenAI Security Project, LLM06:2025 Excessive Agency.
- OpenAI, Guardrails and human review.
- OpenAI, Designing AI agents to resist prompt injection.
- Lotfi et al., Securing Agentic AI: From Per-Action Checks to Trajectory Assurance, 2026.
- Anthropic, Making Claude Code more secure and autonomous with sandboxing.
- Anthropic, Trustworthy agents in practice.
