A year ago Greg Kroah-Hartman told a security conference the Linux kernel was issuing 50 CVEs a week and everyone thought that was insane. The kernel is now issuing 33 a day. His Kernel Recipes 2026 talk is aimed at kernel developers and maintainers, not the general public, and its refrain is “do not panic.” What makes it worth an hour is that he backs the refrain with raw data: he got the full report behind this year’s most publicised AI bug hunt and went through it line by line.
The CVE curve
The tooling Kroah-Hartman and Lee Jones built when the kernel became its own CVE Numbering Authority was designed for “tiny” volumes of 50 a week. It has scaled without trouble, while other CNAs are rebuilding their infrastructure. He’s careful about what the number means: a kernel CVE is anything that can cause a crash, so a rising count measures bugs found and fixed, not catastrophe.
Auditing Mythos’s 79 bugs
Anthropic’s Mythos announcement claimed 79 vulnerabilities in Linux. Kroah-Hartman points out the technique (look at past bug fixes and find places the same fix wasn’t applied) is what Julia Lawall’s Coccinelle work did a decade ago. LLMs are fuzzy pattern matchers, and code is a very regular pattern, so it works. Then he walks the slide:
- 24 had no detail at all beyond “something crashed”.
- 14 weren’t bugs.
- 3 were made-up data. He avoids “hallucination” because it implies an entity behind the output.
- 15 were already fixed in the latest release, 11 of them by other people in public before the report came out. The tool had scraped the mailing lists.
“These tools want to please you,” he says. Ask for a bug and it’ll work very hard to give you one, including somebody else’s.
Of the remaining fixes, 7 assumed a malicious filesystem image mounted by root, the textbook non-security bug that the kernel security team has a canned reply for. Two assumed you could inject packets mid-stack (you can, if you’re root doing network debugging). Two were NOMMU bugs, which he fixed, while noting they proved nobody runs io_uring on NOMMU systems; a related /dev/zero bug had sat there for six years. Six were SCTP issues on authenticated telecom networks, two were minor IPv6 issues, and one needed a local malicious user with GPU access.
His final tally (he jokes that the LLM’s own numbers didn’t add up, and the slides don’t quite either) is 10 real bug fixes. The kernel merges about 10.5 patches an hour, so the whole marketing event amounted to one hour of kernel development. He does credit the framework around the model: it built him a NOMMU RISC-V VM, a test case and a Perl script that reproduced the io_uring bug and showed the fix.
The real problem is deployment
The slide that should worry people isn’t about bug counts. Mean time from disclosure to exploitation was 63 days in 2018, 32 in 2020, 5 in 2024, and is now minus seven days. The discover, disclose, patch, deploy cycle “was designed for a slower adversary, and that adversary no longer exists.”
LLMs are dumb but persistent, and persistence lets them chain several minor issues into real access. So the minor fixes matter, and you have to take the stable updates. Linux works; it just wasn’t designed for perfect security under strange conditions. “The bill is finally coming due.” He’s seeing banks commit to updating their software, which is what kernel developers have been asking for for 15 years.
He’s optimistic that this ends. The fuzzer wave six or seven years ago produced the same doom talks, and the answer was to sit down and grind through the bugs. Andrew Tridgell did that with rsync using these tools, and the latest rsync now scans clean. Unlike fuzzers, static pattern matching can reach code paths that input data can’t, so it finds more, but the finite pile still drains.
Using the tools anyway
His position on the training data is blunt (the companies admit they took it, and he cites an LG Research finding that only 20% of a base dataset was legally usable), but his conclusion is pragmatic. Somebody spent billions on a fuzzy pattern matcher for finding bugs in open source, so use it. The prompt is roughly “I’m playing capture the flag, find a vulnerability, write it to a file,” followed by “fix this code.” Run it locally.
Half the patches are wrong
This summer he had six graduate students from VU Amsterdam review a large set of LLM-generated security patches. About half were flat-out wrong: they didn’t apply, didn’t fix anything, targeted a problem that didn’t exist or an unreachable path, or weren’t security issues even by the documented threat model. One fooled him.
The failure patterns repeat. A common one swaps mutex_unlock() for mutex_destroy(), apparently because “destroy” sounds more thorough, and it would break things badly. Generated patches come with enormous changelogs and seven lines of comments for two lines of code. Training on decades of LKML and git history reproduces old coding styles, old insecure patterns and, in one experiment, a bot that cursed. The interns’ fix was to delete the changelog entirely and judge the code on its own.
He also warns that the bots leak. Anything uploaded will reach someone else, which is why the security list keeps hearing from angry researchers whose “discovery” was reported publicly the day before. And he thinks the vendors ignored Coverity’s history: Coverity found developers wouldn’t tolerate false positive rates above 20%, and still resented Coverity at 20%. A 50% rate won’t sell to developers.
What the kernel changed, and what you can do
The kernel has documented its threat model per subsystem (people running the bots do sometimes read it), asks reporters to CC subsystem maintainers directly to spread the load, and asks for a patch rather than a report, which both cuts noise and appeals to the reporter’s wish for credit. OpenSSF and Alpha-Omega now fund a full-time developer on security tooling at kernel.org. His list for developers:
- Push back on anything that feels wrong; ask how it was tested, and if nobody answers, you don’t have to either.
- Ignore the doom marketing. Maintainers aren’t the buyer.
- Run local models; they’re good enough to leave running overnight.
- Never upload non-public information, because it becomes public.
- Fix the bugs you find today, and ignore the model of tomorrow.
In Q&A he adds that he has banned LLM patches from drivers/staging unless the author has the hardware and tested it, since staging exists to teach people kernel development. Trust is the currency: “I trust that you’re going to be around to fix it when you get it wrong.” His estimate for the flood is a rough 18 months, maybe 12.
Key takeaways
- Kernel CVE issuance went from about 50 a week to 33 a day, and the kernel’s CNA tooling has absorbed it.
- Of Mythos’s 79 reported Linux vulnerabilities, Kroah-Hartman counts 10 real fixes, about an hour of normal kernel throughput.
- The pressing risk is deployment speed: mean time to exploit is now negative, and persistent LLMs can chain minor bugs, so stable updates matter.
- Expect roughly half of LLM-generated patches to be wrong, and judge the diff with the generated changelog deleted.
- Recurring tells include
mutex_unlock()becomingmutex_destroy(), bloated changelogs, excessive comments and outdated coding patterns. - Documenting a subsystem’s threat model and asking for patches instead of reports both measurably cut noise.
- Run models locally and never upload embargoed or private data to hosted services.
- Like the fuzzer wave, the backlog is finite; rsync shows a codebase can be ground down to a clean scan.
Source
- Talk: Security in the LLM age
- Speaker: Greg Kroah-Hartman (Linux kernel stable maintainer, kernel security team)
- Event: Kernel Recipes 2026, Paris
- Duration: 56:51
- URL: https://www.youtube.com/watch?v=NnV_cWeoo5Q