Talk recap: why ASM and CTEM tools are about to fall behind
At the September 17, 2026 CISO Platform Turbo Talk, FireCompass founder and CEO Bikash Barai argued that traditional ASM and CTEM tools will fail in the next 12 months. His reason is economics. A vendor can afford about half a cent per asset assessment. An attacker running AI agents can now spend a few hundred dollars on the same asset and test it as deeply as a human pentester.
The argument carries weight because of who made it. Barai’s team coined Continuous Automated Red Teaming (CART) and helped build the ASM and CTEM category. In a 20-minute session, he made the case against the tools his own company helped define.
This post summarizes the talk: the math, the experiment behind it, and the depth-first order Barai recommends.
The gist
- Daily ASM on 200 assets is 73,000 assessments a year. At typical prices, that leaves a vendor about half a cent to 11 cents per assessment.
- FireCompass AI agents ranked in the top 3 across HackerOne categories in one quarter, on a $5,000 monthly token budget.
- Barai’s fix is to flip the order: depth first, then breadth, then continuity.
- His closing line: “You cannot outscan an attacker who can now afford to out-think you.”
The half-cent math behind traditional ASM and CTEM
Barai opened with a worked example. Take any ASM, CTEM or CART product built before the current wave of AI agents, which means anything designed more than a year ago. Then run the numbers on a mid-sized estate.
Say an organization has 200 internet-facing assets and monitors them daily. That is 200 × 365, or 73,000 asset assessments a year.
Pricing varies widely. Barai noted that some third-party risk and ASM tools start at about $1,000 per organization per year, while $20,000 a year is common among the better tools. He used those two as the range.
| Annual price | Per assessment | Vendor’s delivery budget at 60% gross margin |
|---|---|---|
| $1,000 | 1.4 cents | about 0.5 cents |
| $20,000 | 27 cents | about 11 cents |
Vendors need margin. Assuming a 60 percent gross margin, before sales and admin costs, leaves about half a cent to 11 cents to actually test each asset each day.
Half a cent covers a banner grab, a port check and a version match against a CVE feed. It does not cover an authenticated crawl, an IDOR check or an attempt to chain two findings together. Those cost compute, and now they cost tokens.
The $10,000 alternative: depth from a manual pentest
The other option, Barai said, has always been a manual pentest. In the US, a consulting firm typically spends about 5 days on a single application, at a cost of around $10,000.
That money buys real depth. A good tester finds complex business logic flaws, IDOR and BOLA (Broken Object Level Authorization) issues, and multi-step attack chains. Traditional DAST tools and scanners miss these classes almost by design. The OWASP API Security Top 10 ranks broken object level authorization as the number one API risk.
So security teams have had two choices:
| Option | Cost per asset | Depth | Typical coverage |
|---|---|---|---|
| Traditional ASM / CTEM | about half a cent per assessment | Shallow: exposure and known-CVE matching | Broad, daily |
| Manual pentest | about $10,000 per application | Deep: logic flaws, IDOR, BOLA, chains | Narrow, about 20% of assets, once a year |
That was the market a year ago. By Barai’s estimate, it is still the market for about 90 percent of the industry today.
What changed: $5,000 a month of tokens reached the HackerOne top 3
To show what changed, Barai described an internal experiment. FireCompass capped token spend for its AI pentesting agents at $5,000 a month and pointed them at HackerOne. The question was how far that budget would go.
After 3 months, the agents ranked number 1, 2 or 3 across various HackerOne categories. On HackerOne, a researcher only scores by reporting a vulnerability before everyone else on the platform. The agents reported more than 200 vulnerabilities and held the number 2 reputation for reporting critical ones.
The finds included zero days, one of them a critical flaw in one of the world’s internet registries. They also covered the classes that used to need a human: IDOR, BOLA, business logic flaws and multi-stage attack chains.
Barai put the budget in context. $5,000 a month is about half the salary of an entry-level pentester in the US. “That’s a very scary thing, and that’s also a very promising thing,” he said, “depending on which side you are.”
He added a disclaimer. The result came from years of work and about $30 million raised to build the harness: FireCompass’s own small language models (SLMs), frontier models, and an agentic architecture on top. “You just can’t take AI and achieve this result,” he said. But sophisticated adversaries have built systems like this too.
Barai pointed to signals across the industry. Peers and customers report a sharp rise in bug bounty submissions. Microsoft’s September 2026 Patch Tuesday was its biggest ever, with 974 CVEs fixed. Anthropic gave Claude Mythos Preview first to operating system, browser and major software vendors through Project Glasswing, to give defenders a head start. But as Barai noted, comparable models, including strong open-weight ones, are already within an attacker’s reach.
What is the depth gap in exposure management?
The depth gap (or depth debt) is the difference between how deeply defenders test an asset and how deeply an AI-assisted attacker can now attack it. Traditional ASM and CTEM tools spend about half a cent per asset assessment. An attacker using AI agents can spend a few hundred dollars on the same asset and reach pentest-level depth.
Before AI, the attacker’s side of that equation cost about $10,000 per application, and few adversaries spent that on a single mid-value target. Barai’s point is that a few hundred dollars of tokens now buys the same sophistication, at scale across thousands of targets.
A half-cent scan cannot match a $300 AI pentest. That, in Barai’s view, is why traditional ASM and CTEM tools will fail in the next 12 months.
Why the depth gap compounds
The gap grows on three fronts at once:
- Attacker cost keeps falling. A deep test dropped from about $10,000 to a few hundred dollars of tokens. The half-cent defender budget has not moved.
- Exploitation keeps getting faster. New CVEs are now exploited in about 3 days. An annual pentest leaves roughly 365 days between checks.
- Untested surface stays untested. Most programs pentest about 20 percent of their assets. The other 80 percent only gets the half-cent treatment, and that is where AI-driven attackers will look first.
“But ASM was never meant to be a pentest”
Barai raised the obvious objection himself, and agreed with it. ASM was built for asset discovery, not penetration testing.
His point was about what buyers expect. Most teams treat ASM as their way to find risk, not just to list assets. Many CTEM products, he said, are ASM tools with a thin wrapper, branded as exposure management. CTEM with real AI-driven depth built in would cost far more than half a cent per asset.
The bigger problem sits behind the ASM tool. The applications that did get pentested got a $10,000 test, on a partial set of assets, once a year. Those same applications now face adversaries running AI attacks against them continuously.
A low-cost ASM or CTEM layer cannot close that pentest gap, because it was never designed to. Barai’s conclusion: the gap has to be closed directly, before anything else.
Flip the order: depth first, then breadth, then continuity
The industry has spent years pushing continuity first. Barai recalled being pleased a year ago when regulators said CART should be mandatory, since FireCompass coined the acronym. With AI in attackers’ hands, he now argues the order has to flip. “We need to flip the priority,” he said.
He framed it as a hierarchy of needs, like Maslow’s. Continuity does not help when the foundation was never tested deeply.
Close the depth gap before you add continuity
3
2
1
Depth-first testing order · 3 layers
- Close the depth gap. Identify Priority 1 (P1) applications and get a high-quality, deep pentest on each one, with business logic, IDOR, BOLA and attack-chain coverage.
- Test lateral movement. Look at the next tier of assets. Can an attacker get initial access there and move laterally to something that matters? Add deep red teaming where it counts.
- Extend breadth. Bring the rest of the estate into scope. Barai said the old model of pentesting about 20 percent of assets once a year will not hold.
- Add continuity. Only now does continuous, shallow monitoring earn its place, watching for drift and new exposure on a base that has already been tested deeply.
Why this order? Barai warned that if teams skip depth, someone else will find what they missed. It might be a bounty hunter, a vendor emailing management with a vulnerability report, or a breach. Any of those resets the priorities anyway.
Build the program by asset class and frequency
Once the depth gap is closed, Barai recommends building the program itself in this sequence: discover assets, label them, group them, then decide which class of testing each group gets and how often.
Not everything needs the same cadence. A deep pentest and a Day 1 CVE check are different jobs with different costs. The table below lays out the model he described.
| Asset group | Testing class | Frequency |
|---|---|---|
| P1 applications | Deep AI pentest: business logic, IDOR, BOLA, attack chains, PoC-validated | Every major release, or monthly, based on budget |
| Next-tier assets | Initial access and lateral movement testing, deep red teaming where it counts | Set by risk and budget |
| All assets | Day 1 CVE exposure check: is an affected version running? | Daily |
| Whole estate | Shallow continuous monitoring (ASM, CART) for drift and new exposure | Continuous, after depth and breadth are covered |
Why always-on deep testing will not be the default
Deep testing everything all the time sounds ideal, but Barai said continuous token costs make it very hard to afford. He expects testing to move to a consumption-based model, the way infrastructure moved to the cloud. Teams will spend deep-testing budget where depth matters, and cheap checks everywhere else.
In that model, a modern CTEM program decides which assets get which class of testing, at which frequency, for which budget.
Where FireCompass fits
FireCompass builds AI agents to close the depth gap at a price teams can run more than once a year. The agents cover discovery, web app and API pentesting, and multi-stage attack paths, including lateral movement from web into infrastructure. Every finding comes with a proof of concept.
The numbers behind it:
- Under 2 percent false positives, against 40 to 70 percent for scanners.
- 11x cheaper than manual testing. One Fortune 500 customer went from about $5,000 to under $1,000 per app.
- 10x faster. About 1 day of lead time, against 2 or more weeks for a manual engagement.
- 100 percent on public benchmarks. XBEN 104/104 and Acuart 12/12, PoC-validated.
At that price, deep testing every P1 application on every major release becomes something a team can budget for.
The takeaway
Half-cent tools will keep producing clean dashboards while attackers test the depth those dashboards never reach. Barai closed the talk with one line: “You cannot outscan an attacker who can now afford to out-think you.”
The practical first step is the one he put at the base of the hierarchy: deep-test the P1 applications and see what turns up before someone else finds it. Hack yourself before AI does.
