Cheap Competence, Hostile Frontier
Why AI makes excellent security people more valuable, not less — even as it hands everyone else a loaded gun
In September 2025, a hacking operation ran at a speed no human team could match. A group that Anthropic later attributed to the Chinese state pointed the company’s Claude models at roughly thirty targets — major technology firms, banks, chemical manufacturers, and government agencies across several countries — and let the machine do the work. It mapped networks, wrote its own exploit code, harvested credentials, and sorted the stolen data by intelligence value, issuing what Anthropic described as thousands of requests per second across six structured phases. Human operators stepped in only at a handful of decision gates — approving the jump from reconnaissance to live exploitation, signing off on exfiltration. By Anthropic’s estimate the model executed 80–90% of the campaign on its own. Four of the thirty organizations were breached.
The most unsettling detail isn’t the autonomy. It’s the banality of the toolkit. The operators didn’t deploy exotic custom malware; they leaned on ordinary, off-the-shelf open-source penetration-testing tools, and they slipped past the model’s safety guardrails by convincing it that it was performing authorized defensive testing. Elite nation-state tradecraft — the kind that used to require a scarce, expensive team of human operators — was reassembled out of cheap, scalable, commodity parts.
That campaign, which Anthropic disclosed in November 2025, is the most dramatic point in a pattern that formed across the whole year. In June 2025, an autonomous system called XBOW reached #1 on HackerOne’s U.S. bug-bounty leaderboard, filing roughly 1,060 vulnerability reports in about ninety days — 54 rated critical, 242 high. In August, at DEF CON 33, seven machines competing in DARPA’s AI Cyber Challenge found 77% of the vulnerabilities planted in real open-source code and auto-patched 61%, averaging 45 minutes each — and tripped over eighteen genuine, previously unknown zero-days along the way. And Google’s Big Sleep agent went from finding a novel flaw in SQLite in late 2024 to helping halt the in-the-wild exploitation of one in 2025.
For thirty years the security industry assumed this work was safe from automation — too contextual, too adversarial, too dependent on intuition. A scanner could flag a missing patch, but it took a human to know the patch didn’t matter because the box was air-gapped, or that the box everyone ignored was the one quietly bridging two networks. Judgment was the moat. The evidence of the past year says the moat is draining.
Competence in security — finding a bug, writing an exploit, triaging an alert, faking a face on a video call — is becoming cheap. Not free, not flawless, but cheap.
The interesting question is not whether the moat is draining. It’s what becomes scarce, and therefore valuable, once it does.
The scarce skill is no longer knowing how to find the vulnerability. It’s knowing which of the ten thousand things the machine just found actually matters for this organization, on this day, against this adversary.
My argument, in short: the commoditization of security competence does not make the excellent practitioner irrelevant. It makes them dramatically more valuable — while making the merely adequate one expendable, and the field as a whole more dangerous. The same forces that arm the defender arm the attacker, and the same forces that elevate the expert kick the bottom rung off the ladder that produced experts in the first place. That’s the frontier we’re walking into, and it is not a friendly one.
Offense: the exploit was the moat. The exploit is now a commodity.
Attackers have no compliance department slowing them down, so the change shows up there first. The traditional pen test was scarce because skilled testers were scarce: a good operator might cover a handful of targets a quarter. AI dissolves that constraint. An agentic pentester reasons about what it finds, chains tools together, adapts — and runs a thousand engagements in parallel.
XBOW is the loud proof point. The more sobering one is Google’s Big Sleep. In November 2024 the agent — built by DeepMind and Project Zero — found a previously unknown memory-safety flaw in SQLite, one of the most widely deployed pieces of software on Earth. By mid-2025 it had identified a vulnerability (CVE-2025-6965) known only to threat actors and helped halt its exploitation: what Google called the first documented case of an AI agent directly preventing a zero-day from being used in the wild.
For an offensive-security firm this cuts both ways. The commodity layer — recon, surface enumeration, the first pass of known-pattern exploitation, the boilerplate of the report — collapses toward zero marginal cost. But the same collapse is a gift to the operator who was never really selling hours. The expert sells the judgment to know which finding is a foothold and which is noise, the creativity to chain three “medium” issues into a critical kill chain, the business context to say this SSRF reaches your crown jewels and that one reaches a marketing microsite nobody cares about.
There’s already a tax on getting this wrong. When the curl project ended its bug-bounty program in early 2026, the complaint wasn’t that AI couldn’t find bugs — it was that roughly a fifth of submissions were AI slop and only about 5% of the year’s reports were genuine. Cheap competence without judgment produces a torrent of confident, well-formatted, hallucinated nonsense. The person who can tell signal from slop, instantly and at scale, is the one whose time is now worth defending.
Defense: the SOC drowns in alerts. AI bails water. Someone still has to steer.
The defender’s chronic disease is volume: too many alerts, too few analysts, a tier-1 queue that burns people out before they learn anything. Surveys routinely find a majority of SOC staff reporting burnout and naming triage as the place they most want help. The vendors are arriving there first, and the numbers they report are not subtle. CrowdStrike says its Charlotte AI agent autonomously triages endpoint detections with over 98% agreement with its human analysts — trained on millions of real analyst decisions and more than a trillion daily signals — saving teams an average of 40+ hours of manual work a week and cutting investigation steps by up to 85%. Microsoft reports that Security Copilot customers save up to 40% of their analysts’ time on foundational investigation, threat-hunting, and intelligence work. Whatever discount you apply to vendor math, the direction is unambiguous: the rote layer of defense is being automated out from under the people who used to do it.
This is the medical analogy almost exactly. When patients walk in armed with their own lab work and a plausible differential, the low-information visit disappears and the physician’s value migrates upward — to situated judgment about what to do next for this particular body. The SOC analyst is undergoing the same migration. When the machine has already triaged the obvious, the human is left with the residue: the ambiguous, the novel, the alert that’s technically benign but contextually alarming because of who it touches and when.
But defense reveals the catch that offense lets you ignore. A doctor who over-trusts an AI diagnosis harms one patient. A SOC that over-trusts an AI verdict can wave through the one alert that mattered — or be deliberately fed inputs designed to make the model misclassify. Defensive AI is itself an attack surface. Prompt injection, poisoned telemetry, and adversarial evasion are the natural next move once attackers know a model sits in the loop. Cheap competence raises the floor of defense and creates a brand-new ceiling to defend.
The arms race: every capability ships to both sides on the same day.
Here’s where the security version diverges sharply from the medical one. Cheap competence in medicine has essentially one user. Cheap competence in security has two, locked in direct opposition, receiving the upgrade simultaneously. The model that helps Big Sleep find a zero-day first is the same class of model that, pointed the other way, helps the bad guys find it first.
Return to the GTG-1002 campaign with that in mind, because it collapses a long-standing assumption: that advanced offensive operations require scarce human talent, that nation-state tradecraft is a labor bottleneck. A campaign that ran mostly on open-source tools, at machine speed, with humans touching only a handful of decision gates, says the bottleneck is dissolving. The barrier to sophisticated attack is dropping toward the cost of inference — and the same model that triages a defender’s alerts with 98% fidelity will, pointed the other way, generate the attacks those alerts are meant to catch.
Social engineering is the cleanest illustration, because it needs no exploit — just plausibility, manufactured cheaply. In early 2024 an Arup employee in Hong Kong joined a video call with what appeared to be the company’s CFO and several colleagues, all of them deepfakes, and wired out about $25.6 million across fifteen transactions. The raw materials were public meeting footage. The deepfake was the cheap competence; the scarce, expensive thing that failed was an organizational judgment process robust enough to doubt a face on a screen. And it’s not a stunt — it’s a curve, as the chart below shows.
Sumsub reported deepfake fraud attempts rose roughly 2,137% over three years, from 0.1% to 6.5% of all fraud attempts. The 2024 midpoint is interpolated.
In an adversarial system, automating one side forces the other to automate too, and the equilibrium settles at higher complexity rather than fewer people. The frontier moves. It does not close.
The ladder problem: we’re kicking out the rung that makes experts.
If the thesis is “excellent practitioners become more valuable,” the uncomfortable follow-up is: where do excellent practitioners come from? They come from being mediocre first — from years of grinding tier-1 alerts, running rote scans, and writing the boring sections of the report until pattern recognition becomes intuition. That apprenticeship is precisely the work AI is eating.
The 2025 ISC2 Workforce Study, drawing on more than 16,000 professionals, stopped publishing a single headcount “gap” figure altogether — the 4.7 million number it cited a year earlier — and reframed the crisis around skills: 95% reported at least one skills gap, and 59% named critical or significant skills needs, up from 44% the year before. Meanwhile the SANS 2026 Workforce report found that 61% of organizations had cut AI-related roles, with entry-level analyst positions taking the heaviest hit at 32%.
Entry-level positions are disappearing just as the training ground that produces senior talent erodes. We’re hiring people experienced enough to supervise the AI — and quietly defunding the only process that has ever produced them.
The medical version of this argument can end on a clean note, because med school and residency still exist as a deliberate, protected apprenticeship. Security has no equivalent. Its apprenticeship was always on the job, and the job is changing under it.
And there’s no clean market fix, because the incentives are broken at the root. A firm that cuts its junior tier captures the savings this year; the cost — a missing senior bench — is diffuse, shared across the whole industry, and lands five to ten years out. That’s a textbook externality, and left alone it resolves the way externalities always do: everyone defects, and the pipeline starves. The fix is unlikely to be something any single company does voluntarily.
Rebuilding the ladder: if the apprenticeship is gone, it has to be rebuilt on purpose.
The most promising move is to redesign the junior role around judgment rather than delete it. The old apprenticeship was hand-triaging ten thousand alerts until pattern recognition set in. The new one can be auditing the AI as it triages a hundred thousand — and learning to find where it’s wrong. That’s potentially faster exposure to more cases, not less, but only if someone builds the role that way. Most organizations aren’t; they’re simply removing the seat. The richest teaching moment in an AI-saturated SOC is the override: a junior watching an expert say the model is confident here and it’s wrong, and here’s how I know. That has to be designed in, not assumed.
Security can also borrow the move aviation made when real cockpit hours became scarce and expensive: simulators. Cyber ranges, purple-team exercises, capture-the-flag competitions, and — with some irony — AI-generated attack scenarios can manufacture the reps that once required grinding live production work. The same technology hollowing out the apprenticeship turns out to be unusually good at rebuilding it.
What neither fix addresses is who pays for training that isn’t immediately productive. Medicine solved that with residency: a protected, subsidized apprenticeship backed by professional bodies and public funding. Security has nothing like it. Given that this is a national-security domain, that’s probably where the answer has to come from — government, certification bodies, or industry consortia treating the talent pipeline as shared infrastructure rather than a line item any one employer is foolish enough to fund alone.
The frontier: what stays scarce.
So what is actually scarce on this hostile frontier? The honest answer has to survive an objection, because “judgment is the moat” is exactly what the field said about the last rung, and the rung before that. Every time AI climbs one, we relocate the moat to the next floor up and christen that one uniquely human. A retreating frontier should make you suspicious that you’re describing a temporary position as a permanent one. There’s no law of nature that reserves contextual judgment for people. If the trend holds, situated judgment is just where the water sits today — not where it stops.
Which means the durable thing probably isn’t a skill at all. It’s a role, and roles can be sticky for reasons that have nothing to do with who’s more capable. Three of them hold up the floor.
Accountability is structural, not technical: when twenty-five million dollars moves or a hospital goes dark, some person or licensed entity has to own that call to a board, a regulator, a court — and that requirement outlasts even a superhuman model, because it’s a fact about liability, not competence.
Adversarial novelty resists automation in a way a chest X-ray never will: the moment a defense is fully automated and therefore predictable, the attacker optimizes against that exact system, creating permanent demand for unpredictability.
And intent-setting — deciding what’s even worth protecting and what risk is acceptable — is a question of values and business, the last thing any sane organization hands to the system it’s trying to constrain.
Notice what’s missing from that list: the ability to perform a security task. Finding the bug, writing the exploit, triaging the alert — all of it is migrating to the machine. What’s left to humans compresses toward the top of the stack: defining the goal, owning the consequence, and supplying the novelty an adversary can’t pre-compute. That compression has further to run than most practitioners want to admit, and its floor is a policy choice as much as a technical fact. How much we keep humans meaningfully in the loop — through liability regimes, regulation, and what we decide to let machines be accountable for — is something we choose, not something physics guarantees.
This is Dan Shipper’s point with the safety filed off. Automation doesn’t end human work; it hands humans a new frame to give the models, and in security that frame is the threat itself — intelligent, funded, and now holding the same cheap competence you are. The work doesn’t disappear. It moves up the stack: from doing to deciding, from coverage to judgment, from finding the bug to owning the call. But “humanity still has work” and “your job survived” are different claims, and the gap between them is exactly the pipeline we’re letting starve.
I’m optimistic, with a sharper edge than the medical version warrants — and fewer guarantees. AI does not make security practitioners irrelevant. It makes excellent ones more valuable, makes adequate ones replaceable, arms every adversary on the same schedule, threatens to starve the pipeline that produces the excellence it rewards, and offers no promise that today’s expert judgment is safe tomorrow.
The competence is getting cheap. The judgment is getting expensive. The floor under the human role is real — but it’s held up by law, by adversaries, and by our own choices, not by any assurance that the machine won’t get that good. The frontier is wide open, and there’s someone on the other side of it using the exact same tools you are.
Sources:
· Anthropic AI espionage disclosure and full report (PDF)
Framing adapted from Every’s “Cheap Competence, New Frontier.”


