Anthropic restores Fable 5 access with new guardrails & jailbreak metrics after exploit. Industry-wide vulnerability led to CJS scale & enhanced safety.

Copy, download or open this article in ChatGPT or Claude
Anthropic has restored global access to its Fable 5 model following the U.S. government's decision to lift export controls on June 30, 2026. The freeze, which began on June 12, temporarily sidelined both Fable 5 and Mythos 5 after Amazon researchers discovered a critical exploit. This vulnerability allowed the model to find software bugs and write working exploit code. To get back online, Anthropic teamed up with federal agencies and industry partners to build stronger safety systems and create a standard way to measure these risks.
During its investigation, Anthropic tested other leading models to see if the vulnerability was unique to Fable 5. It was not. The same exploit bypassed safety filters on Anthropic's own Opus 4.8, OpenAI's GPT-5.5, and Moonshot's Kimi K2.7, producing the same dangerous code. This confirmed the issue was an industry-wide alignment challenge rather than a flaw in one specific architecture.
To patch the gap, Anthropic introduced a new safety classifier designed to block this specific exploit vector in over 99% of cases. To keep operations running smoothly for enterprise clients, any prompt flagged by this system is automatically routed to the older, more stable Opus 4.8 model.
The new classification engine splits security-related prompts into four distinct categories:
This conservative approach has a downside: it flags a high number of false positives, which can occasionally disrupt normal software development and debugging.
Because the industry lacks a unified way to measure jailbreak severity, Anthropic teamed up with Amazon, Microsoft, and Google under a new initiative called Project Glasswing. The group has proposed the Cyber Jailbreak Severity (CJS) scale. This metric rates risks from 0 to 4 based on four main factors:
Using these factors, the CJS scale rates risks across five levels:
To make its models more resilient, Anthropic launched a HackerOne bug bounty program to crowdsource new vulnerability discoveries for Fable 5. Meanwhile, under the Project Glasswing framework, a vetted group of U.S. organizations has been granted restricted access to Mythos 5 for defensive security work. Anthropic has also committed to giving the U.S. government early access to any future models with national security implications.
Microsoft unveils its Humanist AI Code of Conduct draft, setting ethical & safety rules for future AI, prioritizing human welfare & preventing misuse.
BenchMIRT leverages psychometrics for precise AI evaluation, analyzing individual prompts for capability, safety, and bias, cutting costs and improving transparency.
The first advocate-general at Belgium's Court of Cassation devoted this year's opening address to AI, and to the tools the judiciary does not have.