Three vulnerability classes I’ve spent most of my career testing for turned up at Black Hat and DEF CON this year, wearing AI costumes. CSP bypass. SSRF. HTML injection race conditions. Every one of them was written off as solved; every one of them was reachable again.
The clearest example was SearchLeak, disclosed by Varonis Threat Labs and presented at DEF CON 34. The chain has three links. A URL parameter in Microsoft 365 Copilot Enterprise Search gets passed to the model as an executable prompt. The model emits an <img> tag that the browser renders mid-stream, before the output sanitizer runs. The image source points to Bing’s Search-by-Image endpoint, which is allowlisted in the page CSP and performs a server-side fetch to an attacker-controlled host.
Victim clicks a microsoft.com link. Copilot searches their mailbox, calendar, SharePoint, and OneDrive. The data leaves through Microsoft’s own infrastructure. Microsoft rated it critical and assigned CVE-2026-42824.
Pull that chain apart, and two of the three links are bugs any of us would have triaged as low ten years ago. A sanitizer that runs after render. An allowlisted domain that will fetch arbitrary URLs. On their own, barely worth a ticket. Composed with a model that follows instructions, they exfiltrate an inbox.
Chaining for less
The second thing worth your attention is how cheap chain discovery got.
James Kettle’s HTTP Terminator, an AI research system built around his own desync methodology, explored 30,000 candidate desync vectors and tested across 30,000 authorized sites, flagging roughly 700 vulnerable targets. Banks, government infrastructure, security products, and an airport. The system generated new desync techniques on its own, but Kettle’s strongest results came when he stepped back into the loop. The model proposed a new vulnerability class, which Kettle calls Shared-Parser Confusion, and he confirmed and generalized it. One researcher’s method, with AI doing the legwork, reached internet scale.
Two weeks before Black Hat, a Searchlight Cyber researcher published a pre-authentication RCE chain in WordPress core found with about $25 of AI usage and just over ten hours of runtime. A batch API validation desync opened a pre-auth SQL injection. From there, the model poisoned the post cache, abused a customize_changeset to borrow the administrator’s identity, and replayed the request through a hook to create a new admin account. The researcher’s own verdict is that no security researcher could have found and finished that chain in ten hours without AI.
For anyone running an AppSec program, the gap between “an attacker could theoretically chain these four lows” and “someone did it on a lunch budget” is now closed.
What the numbers say
Akamai’s 2026 State of the Internet report on apps, APIs and DDoS backs this up:
- 61.18% of API attacks in 2025 involved unauthorized workflows and abnormal activity, up from 30.01% in 2024. Akamai reads it as attackers moving off familiar exploits like SQL injection toward logic-based abuse of how APIs handle business processes.
- 87% of organizations had an API-related security incident in 2025.
- Average API attacks per organization per day hit 258, up 113% from 121 the year before.
Behavior-based abuse doubled its share in twelve months. None of it produces a malformed request, which means signature-based tooling has nothing to fire on.
Where this actually lands
The request is syntactically perfect. Clean 200, valid JSON. The only problem is whose record came back. No payload to fingerprint. Proving it means holding two authenticated identities at once, harvesting real object IDs from the first, and replaying them with the second. Scanners that log in as one user and crawl can’t reach this class at all.
A profile endpoint over-returns and leaks the internal user ID format. That seeds the object corpus. The corpus yields a BOLA against billing. The billing record holds a service API key. That key authenticates to a second internal service. Four findings, every one of them rated low in isolation.
/api/v2/orders enforces authorization. /api/v1/orders is still routed, still live, never got the fix. Staging hosts share a production database; same story. Certificate transparency and DNS enumeration surface these in minutes.
Testing cadence is the real gap
Annual and semi-annual testing was designed around annual and semi-annual releases. Nobody ships that way now. A two-week engagement scoped to three applications tells you those three applications were secure in March, while the team pushed to production two hundred times between then and now.
Most organizations do have continuous coverage of some kind. WAF, SAST, dependency scanning. What they don’t have is anything proving exploitability across the rest of the surface between engagements, so scanners fill the calendar with volume instead. Long lists of theoretical findings, mostly unreachable, which is how a triage backlog reaches four figures and stays there.
Testing needs to run at deployment speed. That’s the problem FireCompass was built for.
How our agents handle it
Continuous external discovery through certificate transparency, DNS, and passive sources surfaces the forgotten v1 hosts and unlinked staging environments, which can then be loaded into our AI Agents for on-demand assessments instead of them sitting in an inventory report that nobody opens. For internal assessments, you can deploy FireCompass’s Internal Appliance behind the perimeter for service-to-service APIs and internal web apps.
Agents hold concurrent sessions across OAuth2/OIDC, SAML, and flows behind MFA, in multiple roles and multiple tenants at the same time. That concurrency is the prerequisite for cross-identity replay, which is the only way to prove BOLA and BFLA rather than guess at them.
A candidate gets filed when the agent has pulled identity A’s record using identity B’s token and can show both requests and both bodies with a reproduction script attached. Our published false positive rate is under 2%, and the reason is that unproven candidates never leave the pipeline.
We put that last claim in public rather than in a datasheet. Three months on a live authorized bug bounty program, $5,000 a month all in, including tokens, cloud, and human review. Our agents reached #1 on HackerOne’s OWASP A01 board for the quarter, #2 for highest critical reputation, and #3 on the US country board. Critical and high findings were 64.4% of everything severity-rated.
A01 is broken access control. Business logic, scored against human researchers, on production targets.
Guardrails
Agents chaining exploits against production need limits, and this is the question that comes up in every second meeting.
Try it against your own surface
SearchLeak chained two low-severity web bugs into an inbox breach. Kettle’s AI-assisted system found roughly 700 vulnerable hosts. A WordPress pre-auth RCE cost $25 to discover.
Your perimeter is being tested at that cadence whether or not you’re testing it yourself.
FireCompass offers a free trial. Point it at your external attack surface and see what comes back with a working proof of concept attached.
