Anthropic Revealed Cybersecurity Details of Fable 5 and Proposed a Jailbreak Severity Scale
On July 2, 2026, Anthropic published details about how Fable 5's safety classifiers work: the model categorizes all cyber requests into four categories—from prohibited to harmless—and uses an extended 'safety margin.' In parallel, Anthropic and partner Glasswing unveiled the first draft of an industry jailbreak severity scale and launched a HackerOne program for researchers.
AI-processed from Anthropic Blog; edited by Hamidun News
On July 2, 2026, Anthropic described in detail the principles of how Fable 5 security classifiers work and released the first draft of an industry framework for evaluating the criticality of jailbreaks.
How Fable 5 security classifiers are organized
Cybersecurity is fundamentally a dual-use field: the same capabilities can serve both defense and attack. This is precisely why Anthropic does not seek to block everything related to it. Instead, Fable 5 classifiers evaluate each request across four categories:
- Prohibited use—actions that in most cases are capable of causing significant harm and lack defensive value. Blocked unconditionally.
- High-risk dual use—tools widely used by attackers but also have legitimate applications. Also blocked.
- Low-risk dual use—capabilities that are primarily defensive in nature, theoretically useful and for attackers. Monitoring; sometimes blocking as a precaution.
- Harmless use—legitimate tasks with no potential for harm. Allowed with monitoring.
The key concept of the system is "safety margin": an intentionally expanded zone in which the classifier blocks requests out of caution, even if they look potentially harmless. A request must look clearly safe to guarantee passage through verification.
For Fable 5, this margin was intentionally made wider than in Anthropic's previous models.
Why is a unified scale of jailbreak criticality needed?
A jailbreak is an unconventional way to make an AI model bypass its own limitations. It can be almost harmless (removes only one minor restriction) or critically dangerous (opens a wide range of harmful capabilities, making the model fundamentally more dangerous).
At the same time, the industry lacks a unified terminology to assess these risks, which seriously complicates dialogue between AI companies and regulators. Anthropic, together with Glasswing partners, prepared the first draft of such a framework. The company views it as a starting point for broad discussion involving the academic community, industry, civil society, and government bodies.
Proposals are accepted at [email protected].
"We believe: working together, we will be able to develop a standard that will allow us to use this technology for defensive purposes, while simultaneously preventing abuse," said in
Anthropic's official statement.
In parallel, a program has been launched on the HackerOne platform: security researchers can officially send discovered Fable 5 jailbreaks for verification by the Anthropic team.
What this means
Anthropic openly acknowledges the fundamental complexity of cybersecurity as a field of dual use and chooses a calibrated approach instead of total blocks. The framework for jailbreak criticality is the first serious industry attempt to create a common language for dialogue between AI laboratories and regulators.
If the standard takes root, companies will be able to describe threats of security bypasses in agreed terms, and governments will be able to more accurately assess the risks of new models.
Frequently asked questions
What is "safety margin" in Fable 5?
"Safety margin" is a zone in which the classifier blocks requests out of caution, even if they do not look explicitly harmful. This reduces the risk of accidentally allowing dangerous tasks at the cost of some false positives. In Fable 5, Anthropic intentionally expanded this zone compared to previous models.
How can researchers report a Fable 5 jailbreak they found?
Anthropic launched a special program on HackerOne: security researchers can send discovered Fable 5 jailbreaks for official verification. Conceptual proposals on the criticality framework are accepted at [email protected].
Need AI working inside your business — not just in your newsfeed?
I build production AI for companies — custom CRM, internal tools, autonomous agents, workflow automation. Owned by you, shaped to your process, no per-seat tax. Built by Zhemal Khamidun, CPO of AlpinaGPT (AI platform, 6,000+ users).
The AI world, distilled — once a week
Seven stories that actually mattered, hand-picked. No noise, no reposts, no press releases.
Done! Check your inbox for a confirmation.