I pasted a root credential into a chat window and pressed enter, and the thing I felt was not confidence. It was the particular calm you get when you have run out of better ideas.
The server belonged to a small nonprofit, and someone had been inside it for a while — the website serving things it had never been asked to serve, admin accounts no human there had created, mail leaving the building in volumes nobody could explain. The people who ran it were not careless. They were four people doing the work of eleven, and their technology budget for the year would not have covered two days of a competent incident-response firm.
So I gave an AI agent shell access to a compromised production server and told it to find out what happened.
Under fifty minutes later it had walked the logs into a timeline, written the whole thing up in a form the police could actually accept, removed the malicious files, killed the accounts that should not have existed, rotated every secret on the box, closed the door that had been left open, and brought the stack up to a version that was no longer trivially exploitable.
That story is a story: one incident, one organization, one operator watching, and I am the least neutral witness available. Do not update your view of anything based on my Tuesday. What I want to argue survives if you discount my afternoon entirely.
The ledger only has one column
There is a public register of what artificial intelligence does in the world. The AI Incident Database is genuinely good — careful, well-organized, run by people who take it seriously. It describes itself as indexing the collective history of harms or near harms realized in the real world by deployed AI systems.
Harms. That is the whole scope, and it is deliberate.
Want the Replit coding agent reported to have deleted a live production database during an active code freeze, despite repeated instructions to touch nothing? Incident 1152, filed and indexed and searchable — and hedged, four times over, with a "reportedly." Good. It should be there, it was bad, and pretending otherwise helps nobody.
Now find the register of the other thing — the one where an agent shortened somebody's worst week. Nobody keeps one. Not small, not underfunded: not a category anyone maintains.
And here is the part I did not expect. The database is explicit about where it got the idea: its About page says that much like the transportation sector before it, and computer security after that, intelligent systems need a repository of problems experienced in the real world. Aviation. Hold onto that — we are coming back to it, and the borrowing stopped halfway.
First, the part I am not going to argue with
The rogue-agent coverage is not hype, and I am not here to tell you the press invented it.
Late last year Anthropic disclosed that a state-sponsored group had turned an agentic coding tool on roughly thirty global targets, succeeding in a small number of cases, and — the detail that should stop you — that the AI performed 80 to 90 percent of the campaign, with humans stepping in at perhaps four to six decision points. That is not a chatbot writing a phishing email. That is an intrusion campaign where the machine is the operator and the person merely the manager.
The same company later mapped a year's worth of banned accounts — 832 of the accounts it banned for malicious cyber activity between March 2025 and March 2026, the subset detailed enough to assess — onto the standard industry taxonomy of attacker behavior, and found the taxonomy had run out of room: there is, in their words, no ATT&CK ID for this type of agentic orchestration. We are watching a category of behavior that our filing system has no drawer for.
(Yes — both of those come from a company that sells the thing. I cite them because they are the most detailed public accounts available and because they cut against the vendor's own interest, but weigh them knowing who wrote them. Better I say it than you find it.)
Verizon's annual industry-wide breach report, meanwhile, found the exploitation window shrinking from months to mere hours as attackers use AI to move faster, with vulnerability exploitation overtaking stolen credentials as the leading way in.
So: real, documented, getting worse. All of it deserves the coverage it got.
The other half, which is equally documented and got almost none
Now the same period, from the defensive side.
DARPA ran a two-year competition to see whether AI systems could find and fix vulnerabilities in the open-source code critical infrastructure runs on. In the final round, competitors' systems found fifty-four planted flaws and patched forty-three — and then the part that matters more, turned up eighteen real, previously unknown vulnerabilities in live software and supplied eleven patches. Average time per patch: forty-five minutes. Average cost per task: about $152. All seven finalist systems were committed to open-source release under an OSI-approved license, four of them public that same week — so the defensive capability is not being locked away in a vendor's vault.
Separately, Google's vulnerability-hunting agent found a flaw the company says was known only to threat actors and at risk of being exploited, and got it closed first — the first time, Google claims, that an AI agent has directly foiled an attempt to exploit a vulnerability in the wild. A claim, made by the party with the most to gain from it.
And at the boring, unglamorous end where most security actually lives: this year's IBM breach-cost research found the mean time to identify and contain a breach rose to 247 days, reversing five straight years of improvement — while organizations running AI and automation across prevention, detection, investigation and response closed breaches roughly two months faster and paid close to two million dollars less than the ones running none. Half of breached organizations now have AI agents somewhere inside their security operations center.
Let me disarm my own best line before someone else does. It is tempting to set my fifty minutes beside that 247 days and let you do the arithmetic. Don't — they are not the same measurement. The 247-day mean includes the months organizations spend not knowing they are breached at all; my nonprofit already knew. The honest comparison is against the containment tail of it, at a fraction of the scale, and even then it is one anecdote against a mean. The number to keep is forty-five minutes: DARPA's average, measured under competition conditions, published by the government, with the code there for you to check.
Every result above is a documented case of an AI agent doing the thing we say we want. None of them sits in a register of AI agents doing the thing we say we want, because there is no such register.
Now the strongest case against everything I have just told you
Here is where I stop defending myself.
On May 1, 2026, CISA and a set of international partners published a joint guide on adopting agentic AI. Two of its plainest recommendations are: avoid granting broad or unrestricted access, especially to sensitive data or critical systems; and begin with agentic use cases that are low-risk and non-sensitive.
Read that against what I did. Broad, essentially unrestricted access, to the most sensitive system in the building, on the worst day in that organization's history, as my opening move. To construct the scenario the guidance warns about, you would construct mine.
It gets more specific. The risks that guide names include privilege creep, behavioral misalignment and — the one that should worry anybody who does what I did — obscure event records. I can tell you what the agent reported. I would struggle to reconstruct, to an evidentiary standard, everything it actually ran.
Then there is the attack I walked straight into the mouth of. A compromised server is not a neutral environment; it is full of text an attacker wrote. And a preprint published in April 2026 describes precisely this: debugging agents that consume logs and then execute remediation commands are vulnerable to indirect prompt injection through log content — an attacker plants instructions in a log line, the agent reads them as instructions, and the cleanup crew becomes the second intruder. The same paper reports that commercial guardrails from major providers largely fail to catch these, and that some models complied at strikingly high rates. My agent read the attacker's logs. Nothing about my competence prevented that attack; its absence is the only reason we are not talking about a worse story.
And the oldest objection is the best one. Incident response is a discipline with a whole federal publication behind it, one that tells you to collect and retain evidence from an incident — because a compromised machine is evidence before it is a problem. Every rogue account I deleted was an artifact; every rotated key changed the state of a scene. I optimized for getting a nonprofit back on its feet and made somebody's future forensic case harder doing it. Being right about the outcome does not make me right about the order of operations.
So no — I am not telling you to do what I did. I am telling you I did it, that it worked, and that both of those facts belong in the same record.
Aviation keeps two files. We copied one.
Which brings us back to the borrowing.
On December 1, 1974, TWA Flight 514 flew into a mountainside in Virginia after a misunderstanding between air traffic control and the flight crew. The devastating detail, in NASA's own account of it, is what had happened six weeks earlier: a United Airlines flight had narrowly escaped a similar fate. United knew. United had discussed it internally. And in 1974, NASA writes, there was simply no way for that information to reach the wider aviation community.
A near-miss sat inside one company, and ninety-two people died for the silence.
The answer, in April 1976, was the Aviation Safety Reporting System — and the design of it is the whole argument I am making. It is a confidential, voluntary, non-punitive reporting system taking reports from pilots, controllers, dispatchers, cabin crew, ground ops, maintenance technicians and drone operators. It does not wait for wreckage. It explicitly welcomes reports describing close calls and hazards — the flights that landed fine, where something went wrong and somebody caught it in time.
And how do you get someone to file a report about their own worst moment? You pay them in the only currency that works: protection. Under the enabling regulation the FAA will not use a report filed with NASA in any enforcement action, except where accidents or criminal offenses are involved — and that protection is deliberately narrow, requiring the violation to have been inadvertent and filed within ten days. A designed bargain, not a blanket amnesty. The confidentiality is not a promise either, it is a practice: names stripped, dates generalized. Fifty years in, it had logged over 2.1 million reports by 2024, takes more than 120,000 a year, and in all that time no reporter's identity has ever been breached.
Aviation keeps two files: the one about the crashes, and the much larger one about the times something went wrong and did not become a crash. We built the AI version of the first file, said out loud that we were copying aviation, and never built the second.
Just imagine the next few years of a one-sided file
Play it forward, because a record is not a museum — it is an input. So what does a one-column file build?
Insurers price agentic AI off the harm file, the only loss data there is — a corpus with no denominator. So the premium for letting an agent near anything that matters climbs until only large enterprises can pay it. Procurement officers, being reasonable people, read the same file and write the same rule: no agent gets privileged access, anywhere. Legislatures write it into statute.
And the nonprofit with four staff and no security budget gets what follows, which is nothing. No agent, no firm it can afford, no help — it just stays breached a while longer, except now that is a policy outcome rather than an accident, reached by a chain of sensible people reading the same half-empty ledger.
Run the other branch. A confidential, non-punitive place to file: here is what an agent did during an incident, here is the access I gave it, here is what it got right and here is what it nearly broke. De-identified and protected, like ASRS. Within three or four years somebody could answer the question nobody can answer today — how often does this go well, under what supervision, and where is the edge? Every answer to that is currently a vibe, including mine.
What the people who study this actually say
Bruce Schneier, unsentimental about this technology for longer than most people have been paying attention, put the balance plainly last October: AI-based hacking benefits defenders as well as attackers. Part of his reason is distribution — AI, he writes, gives far more people the ability to perform previously complex tasks. Which describes, fairly precisely, a nonprofit that could not otherwise buy an hour of forensics.
From a different corner, Carnegie's July 2026 analysis of autonomous cyber operations argues that defenders also need AI-enabled capabilities to identify anomalous behavior earlier and contain intrusions faster — while noting that Europe's framework was built for human operators and static software, not for autonomous systems moving at machine speed. And from the market-oriented side of the aisle, the R Street Institute, reading the administration's AI Action Plan, flags approvingly its call for federal cyber playbooks to be updated to account for AI-specific risks.
Notice that none of them are arguing about whether agents will be in the loop. They are arguing about the rules — and the rules are being written now, from a file with one column in it.
The tilt is not only institutional, either. Stuart Soroka's cross-national work, run across seventeen countries, found that people give more weight to negative information than to positive information. That is a fact about our nervous systems, not a conspiracy in a newsroom. But it means a register that collects only harm will always feel like the complete picture, and never be one.
So what does this mean for you?
You are probably not handing an agent a root credential this week. Here is the version that applies anyway.
If you run or volunteer for a small organization, do the boring half now. Offline backups you have actually restored from once, multi-factor authentication on email and the site admin, and a written list of every account with elevated access. In a 2023 study of Geneva-based nonprofits, 56 percent had no budget allocated to cybersecurity and 70 percent doubted they had the skills to respond to an attack; 41 percent had already been hit. Nonprofits are also largely missing from the responsible-disclosure ecosystem, so nobody is quietly finding your bugs for you.
If you ever do reach for an agent mid-incident, image the machine first. Snapshot, then remediate. It costs minutes and preserves what you cannot recreate. I did it in the wrong order and got away with it; that is not a method.
Assume the compromised system is talking to your agent. Logs, filenames and error strings on a breached box were written by someone with an interest in what your tools do next. Read the proposed commands before they run — especially the ones that came without an explanation.
Keep your own flight recorder. Full transcript, command history, timestamps, and exactly what access you granted and when you revoked it. That is what "obscure event records" means in practice, and it is the only way your good outcome becomes evidence for anybody else.
Report the near-miss, to somebody. Your ISAC, your funder, your professional community, your own blog. The one-column file stays that way until people under no obligation to talk decide to.
At the next rogue-agent headline, ask what the denominator is. Not to dismiss the story — it is usually true — but because "this happened" and "this is what happens" are different sentences, and only one of them is supported.
The lesson, as I see it
I did something the best available guidance tells you not to do, it worked, and both halves of that sentence are true at once. The temptation is to resolve the tension — decide I was reckless, or decide the guidance is behind the times. Neither, I think.
What I think is that we are trying to govern a technology using a record that structurally cannot hold its saves, then acting surprised when the resulting rules feel out of step with what practitioners see. Aviation worked this out in 1976, after the silence between a near-miss and a mountainside. It did not solve it by trusting pilots more. It solved it by building somewhere safe to file the story that stopped short of disaster — and then reading those files for fifty years.
Credit, in the end, is not a feeling anyone owes a machine. Agents do not need credit. It is a record — and the record is the thing that decides, quietly and years in advance, what the rest of us are allowed to do.
Four people, no security budget, and a very bad week — that is the story that never gets filed anywhere. Chasing down the half of the record nobody keeps is most of what the HAIA Foundation does, and it turns up here every week. If you know someone still running the whole organization off one admin password, send this to them before Monday.





