Skip to content

AI: The World’s Greatest Pentester — And Why It Has to Be This Way

Anthropic's new model Mythos is said to be exceptionally good at finding bugs and vulnerabilities in software. How exploitable many of them are remains unclear, at least for humans — but AI does not think like a human, so they must still be fixed.

Graphic "AI Doesn't Ignore Weaknesses – It Connects Them": a man and an AI figure linked by a path through APIs
In this article

What happens when software is no longer attacked only by humans, but by synthetic actors that think differently than we do?

Anthropic's Mythos: how exploitable are its findings?

Claude developer Anthropic made headlines last week with the internal release of a new model called Mythos. It is said to be exceptionally good at finding bugs and vulnerabilities in software. Due to these capabilities, Anthropic is refraining from a public release for now and instead aims to work with large tech companies and governments to prevent misuse.

It remains unclear how realistic and actually exploitable many of these vulnerabilities are… at least for humans.

Anthropic itself classified a 16-year-old FFmpeg vulnerability as non-critical and considered a working exploit difficult. Potential exploits found in the Linux kernel could not be leveraged due to its layered security mechanisms, and some appear to have already been patched.

Anthropic also admits that the thousands of reported issues are not fully verified, but rather based on extrapolation. The basis: around 90% agreement in 198 manually reviewed cases. [1]

I am very grateful that many of these discovered bugs are unlikely to cause harm to users—whether because patches already exist, layered security concepts mitigate them at higher levels, or other safeguards are in place.

Why these vulnerabilities must be fixed anyway

Nevertheless, these vulnerabilities must be fixed. Why? Because with AI, we have introduced an intelligent actor into our world that does not think like a human. Alongside human logic, there is now (brilliant) AI logic. We already saw this over 10 years ago with Move 37 from AlphaGo.[2]

Our IT systems no longer need to be hardened “only” against human hackers, but also against AI hackers. And AI hackers are far more capable than we are of turning even the most complex and convoluted vulnerabilities into working exploits.

The real risk: a new way of thinking

Conclusion: The real risk does not lie in the individual vulnerability, but in the new way of thinking that can exploit it. Many of these bugs may seem harmless today—but that is a human assessment. AI may arrive at a very different conclusion.

Security used to be a game against human creativity. Now it is a game against something that systematically scales that creativity.

An in-depth AI penetration test for our platform

This month, we will be conducting an in-depth penetration test for our cybersecurity awareness platform and the on-premises versions. I will definitely ask the team to perform an in-depth AI penetration test.

[1]https://www.reddit.com/r/theprimeagen/comments/1siv34l/anthropics_claude_mythos_isnt_a_sentient/

[2]https://www.google.com/search?q=AlphaGo%20Move%2037

What is Anthropic's Mythos model?
Mythos is a new model that Claude developer Anthropic released internally. It is said to be exceptionally good at finding bugs and vulnerabilities in software.
Why is Anthropic not releasing Mythos publicly?
Due to its capabilities, Anthropic is refraining from a public release for now and instead aims to work with large tech companies and governments to prevent misuse.
Are the vulnerabilities Mythos found actually exploitable?
It remains unclear how realistic and actually exploitable many of them are, at least for humans. Anthropic itself classified a 16-year-old FFmpeg vulnerability as non-critical and considered a working exploit difficult, and potential exploits found in the Linux kernel could not be leveraged due to its layered security mechanisms.
How reliable are the thousands of reported issues?
Anthropic admits they are not fully verified, but rather based on extrapolation. The basis is around 90% agreement in 198 manually reviewed cases.
If many of the bugs are unlikely to cause harm, why fix them?
Because with AI, an intelligent actor has been introduced that does not think like a human. Many of these bugs may seem harmless today, but that is a human assessment — AI may arrive at a very different conclusion.
Book a DemoDownload

Sources

  1. AI: The World’s Greatest Pentester — And Why It Has to Be This WayCyberdise

Written by

Palo Stacho

Founder and Managing Director

Founder and Managing Director of Cyberdise AG in Zug, Switzerland. He writes about the state of the awareness industry and why behavior, not knowledge, decides whether an attack succeeds.