FR
live

Greg Kroah-Hartman audits the 79 AI-reported Linux flaws: only ten were real

At Kernel Recipes 2026, Greg Kroah-Hartman walks through his audit of the 79 vulnerabilities Anthropic’s AI tool claimed to have found in Linux: ten real fixes, the rest being noise, duplicates or invented data. The graver signal is elsewhere — the mean time from disclosure to exploitation has dropped to minus seven days, and half of AI-generated patches are wrong.

A single tall stack of printed bug reports on a dark desk, one page carrying a lone amber highlighter mark among the crossed-out ones.

Late September 2026. Greg Kroah-Hartman, maintainer of the stable Linux kernel releases, takes the stage at Kernel Recipes in Paris for a talk titled “Security in the LLM age.” 79. The number of vulnerabilities Anthropic’s AI tool claimed to have found in Linux. 10. The number of real fixes Kroah-Hartman keeps after going through the report line by line. Why it matters: the kernel now issues 33 CVEs a day, up from 50 a week a year ago, and the mean time from disclosure to exploitation has gone negative — the adversary is now faster than the patch.

79 claimed flaws, ten real fixes

The core of the talk is bookkeeping applied to marketing. Anthropic had announced that its tool discovered 79 vulnerabilities in Linux. Kroah-Hartman obtained the full report and sorted it category by category.

The tally is damning. 24 reports had “no detail” beyond “something crashed.” 14 were not bugs. 3 were made-up data — he avoids the word “hallucination,” which implies an entity behind the output. 15 were already fixed in the latest release, 11 of them by other people before the report even came out: the tool had scraped the mailing lists. “These tools want to please you,” he says. “Ask for a bug and it will work very hard to give you one — including somebody else’s.”

Then come the technically real cases: 7 assumed a malicious filesystem image mounted by root — the textbook non-security bug, for which the kernel security team has a canned reply. 2 assumed packet injection mid-stack. 2 were NOMMU bugs, which he fixed while noting that nobody runs io_uring on a no-MMU system. 6 were SCTP issues on authenticated telecom networks, 2 minor IPv6 issues, and 1 required a local malicious user with GPU access.

The final count: ten real fixes. Since the kernel merges about 10.5 patches an hour, the entire marketing event amounts to “one hour of kernel development.” Kroah-Hartman does credit the framework around the model, which built him a NOMMU RISC-V VM, a test case and a Perl script that reproduced the io_uring bug.

The real signal: a negative time-to-exploit

The slide that should worry people is not about bug counts. The mean time from disclosure to exploitation was 63 days in 2018, 32 in 2020, 5 in 2024 — and is now minus seven days. The discover-disclose-patch-deploy cycle “was designed for a slower adversary, and that adversary no longer exists.”

LLMs are “dumb but persistent,” and persistence lets them chain several minor issues into real access. So the minor fixes matter, and you have to take the stable updates. “The bill is finally coming due,” he says — while noting he sees banks committing to update their software, which kernel developers have been asking for for fifteen years.

He remains optimistic about the outcome. The fuzzer wave six or seven years ago produced the same doom talks, and the answer was to sit down and grind through the bugs one at a time. Andrew Tridgell did exactly that with rsync using these tools, and the latest release now scans clean.

Half of AI patches are wrong

The other half of the problem is the quality of the patches AI generates. This summer, Kroah-Hartman had six graduate students from VU Amsterdam review a large set of LLM-generated security patches. About half were “flat-out wrong”: they did not apply, did not fix anything, targeted a nonexistent problem or an unreachable path, or were not security issues even under the documented threat model. One fooled him.

The failure patterns repeat. The most common swaps mutex_unlock() for mutex_destroy(), apparently because “destroy” sounds more thorough — and that would break things badly. Generated patches arrive with enormous changelogs and seven lines of comments for two lines of code. Trained on decades of archives, they reproduce old coding styles and old insecure patterns, and one bot even cursed. The students’ fix was to delete the changelog entirely and judge the code on its own.

He also warns that the bots leak: anything uploaded will reach someone else. And he recalls Coverity’s history — developers never tolerated false-positive rates above 20%, and a 50% rate will not sell to developers.

What the kernel changed

The kernel has drawn concrete conclusions. It documented its threat model per subsystem — the people running the bots sometimes read it — asks reporters to CC maintainers directly to spread the load, and prefers a patch over a report, which cuts noise while appealing to the reporter’s wish for credit. OpenSSF and Alpha-Omega now fund a full-time developer on security tooling at kernel.org.

His list for developers fits in five points: push back on anything that feels wrong, asking how it was tested; ignore the doom marketing; run local models; never upload non-public information; fix today’s bugs rather than waiting for tomorrow’s model. In Q&A, he adds that he has banned LLM patches from drivers/staging unless the author owns the hardware and tested it — staging exists to teach kernel development, not to outsource it. His estimate for how long the flood lasts: “about eighteen months, maybe twelve.”

Verdict

If you maintain an open-source project, keep the two numbers: an AI report contains on average under 13% real fixes, and an AI patch is wrong half the time — human review remains the bottleneck, and the useful reflex is to ask “how did you test this?” before reading the code. If you run Linux systems in production, the lesson is more urgent and does not depend on the bots: time-to-exploit is negative, so the window between “patch published” and “patch deployed” is your only margin — take stable updates the moment they ship, without hand-sorting them. In both cases, do not buy the “model that finds flaws” narrative without the tally behind it: the value of these tools is real, but it is measured in one hour of maintainer time, not in marketing.

References

The cyber brief, every Tuesday

The flaws that matter and the patches to apply, in a ten-minute read.

No spam. One-click unsubscribe.
read next

On the same topic

antiX 26.1 keeps a systemd-free Debian 13 alive, with five init systems and 32-bit

antiX 26.1, released in late September 2026, updates the lightweight Debian 13 “Trixie”-based distribution with no systemd or elogind, five init systems to choose from, and still-maintained 32-bit images. If you are reviving old PCs or want a minimal base whose init you control, antiX is a serious option; otherwise, stay on standard Debian.

MGLRU-FG speeds up Linux memory reclamation by up to 40% in early tests

Kairui Song’s MGLRU-FG patches add frequency-guided promotion to Linux’s Multi-Gen LRU and deliver 10–40% gains depending on workload, peaking at 76% under zRAM. The series is still RFC, but the results on MongoDB, Chromium and kernel builds make it worth tracking closely.

← Back to the feed

Type at least two characters.

↑ ↓ navigate ↵ open esc dismiss