Last updated:
Project Glasswing is a cybersecurity consortium led by the AI safety and research company Anthropic, announced on April 7, 2026.[1] The project brings together major technology and finance companies to use a powerful, unreleased AI model, Claude Mythos Preview, in a controlled environment for defensive cybersecurity purposes.[5][6] Within the first month after launch, Glasswing partners collectively found more than 10,000 high
Project Glasswing was established as a direct response to the powerful, dual-use capabilities observed in Anthropic's frontier AI model, Claude Mythos Preview. After internal testing showed that the model could autonomously find and exploit software flaws, Anthropic decided against a public release, deeming the model too potent for general availability.[7][3]
In this context, the core mission of the consortium is to leverage Claude Mythos Preview in a controlled environment to fortify digital defenses. By providing private, early access to the model, Project Glasswing allows partners to discover and fix vulnerabilities in their own systems and in widely used open-source software. Within the first month of operations, partners reported finding more than 10,000 high
More broadly, the initiative addresses the concern that AI models with advanced cyber-offense capabilities will soon become widespread, potentially empowering state and non-state actors to conduct more frequent and sophisticated attacks. By harnessing the same technology for defense, the project seeks to rebalance the security landscape and establish new best practices for vulnerability management in an era of AI-accelerated threats.[1]
The impetus for Project Glasswing arose from Anthropic's internal development and red-teaming of its frontier AI models. During this work, the company discovered that its latest model, Claude Mythos Preview, possessed emergent capabilities for autonomously identifying software vulnerabilities and creating functional exploits with minimal human guidance. This discovery led Anthropic to launch what it described as an "urgent attempt" to prioritize a defensive application for this powerful technology.[1]
The project's rationale is rooted in managing the dual-use nature of advanced AI. While the model is a powerful tool for defense, its capabilities make it an equally potent instrument for malicious cyberattacks. Anthropic executives stated that they anticipate adversaries will develop AI with similar capabilities in "months, not years," framing the project as a "critical race to secure infrastructure." This sense of urgency was echoed by Logan Graham, Anthropic's Frontier Red Team Lead, who stated, "We need to prepare now for a world where these capabilities are broadly available in 6, 12, 24 months. Many of the assumptions that we’ve built the modern security paradigms on might break."[3][7]
Project Glasswing was officially announced on April 7, 2026.[3] Along with the launch, Anthropic publicly committed to reporting on the project's progress, including specific findings and fixed vulnerabilities, within 90 days. In addition to the technical work, the project was framed as a collaborative effort to engage with government and security organizations to inform national security strategies and evolve industry-wide security practices.[1]
The central technology enabling Project Glasswing is Claude Mythos Preview, a proprietary frontier AI model developed by Anthropic. It is not slated for public release because of its potent and potentially dangerous dual-use capabilities in cybersecurity.[1]
In practice, the model's advanced cybersecurity skills are an emergent property of its general advanced coding and reasoning abilities, rather than the result of specific training for cyber tasks. Anthropic has assessed the model's capabilities as comparable to a "senior security researcher."[7]
Internal testing and initial use within the consortium have shown that Claude Mythos Preview is capable of a range of advanced security tasks, including:
More broadly, benchmark and evaluation results reported by Anthropic indicate substantially higher performance than previous Anthropic models, which the company cites as a reason for restricting access to a controlled environment.[7][3]
In its pre-launch and initial phases, the model was used to identify thousands of previously unknown (zero-day), high-severity vulnerabilities across major software projects.[4] Notable examples of its discoveries include:
In large-scale open-source scanning, Mythos Preview has been used to analyze more than 1,000 open-source projects, surfacing an estimated 23,019 potential security issues, of which 6,202 were classified as high
All vulnerabilities identified during the initial testing phase were reportedly patched in coordination with the respective software maintainers before the project's public announcement.[4]
Claude Mythos Preview demonstrates a notable performance increase in cybersecurity, coding, and reasoning benchmarks compared to Anthropic's next-best publicly available model at the time of launch, Claude Opus 4.6.[1][3]
| Benchmark | Claude Mythos Preview | Claude Opus 4.6 | Description |
|---|---|---|---|
| CyberGym | 83.1% | 66.6% | Measures performance in cybersecurity tasks. |
| SWE-bench Pro | 77.8% | 53.4% | An advanced benchmark for fixing real-world bugs in GitHub repositories. |
| SWE-bench Verified | 93.9% | 80.8% | A benchmark for fixing real-world bugs in GitHub repositories. |
| Terminal-Bench 2.0 | 82.0% | 65.4% | Measures performance in agentic, terminal-based tasks. |
| GPQA Diamond | 94.6% | 91.3% | A benchmark measuring advanced reasoning capabilities. |
| OSWorld-Verified | 79.6% | 72.7% | Measures performance in agentic tasks within an operating system environment. |
Taken together, the data from these benchmarks illustrates the increase in capability that prompted the creation of Project Glasswing.[1][3]
Project Glasswing operates with several key goals:
To extend these capabilities beyond the consortium, Anthropic has introduced Claude Security, a public beta feature for Claude Enterprise that allows organizations to use public Claude models to scan their own codebases and generate proposed patches. In parallel, the company has launched a Cyber Verification Program that enables vetted security professionals to access its models with relaxed cyber-misuse safeguards so they can more effectively validate and remediate vulnerabilities.[5][8]
Together, these goals aim to create a more resilient digital ecosystem prepared for the next generation of cyber threats.[1][4]
To manage its powerful findings responsibly, Project Glasswing follows a structured, multi-step process for vulnerability disclosure:
Overall, this methodology is designed to maximize the defensive benefits of the AI model while minimizing the risk of discovered flaws being exploited.[3]
Project Glasswing was launched as a consortium of over 45 organizations, led by Anthropic.[7]
The founding members of the consortium include major companies from the technology, finance, and cybersecurity sectors:
In addition to the core founding group, access to Claude Mythos Preview has been granted to over 40 other organizations responsible for maintaining critical software infrastructure.[1][3][4]
Building on this initial consortium, on June 2, 2026, Anthropic announced that Project Glasswing would expand from roughly 50 initial partners to about 150 organizations in more than 15 countries. The broader group includes critical infrastructure operators in sectors such as power, water, healthcare, and communications, as well as hardware vendors and nonprofit software maintainers, and Anthropic has estimated that a major attack on many partners’ codebases could affect more than 100 million people.[6][8]
Representatives from partner companies have publicly supported the initiative. Heather Adkins, Google's Vice President of Security Engineering, stated, "Google is pleased to see this cross-industry cybersecurity initiative coming together. We have long believed that AI poses new challenges and opens new opportunities in cyber defense." Similarly, Igor Tsyganskiy, Microsoft's Global CISO, noted, "Joining Project Glasswing, with access to Claude Mythos Preview, allows us to identify and mitigate risk early and augment our security and development solutions so we can better protect customers and Microsoft.”[7]
Anthropic’s initial update on Project Glasswing reported that consortium partners used Claude Mythos Preview to discover more than 10,000 high
Anthropic has committed significant resources to support the project and the broader open-source security ecosystem:
During the initial research preview phase, participants can use the model with costs largely covered by Anthropic's commitment of usage credits. Following this phase, approved participants can access the model at a rate of $125 per million output tokens. The model is made available to participants through multiple platforms, including the Claude API, Amazon Bedrock, Google Cloud's Vertex AI, and Microsoft Foundry.[3]
The official announcement of Project Glasswing on April 7, 2026, was preceded by two unrelated security incidents at Anthropic in late March 2026 that drew public attention.
In response to these concerns, Anthropic officially characterized the incidents as "human errors in publishing tooling, not breaches of our security architecture" and stated that it had implemented improved processes to prevent such errors in the future.[3]
The project is named after the glasswing butterfly (Greta oto). The name serves as a dual metaphor for the project's mission:
Taken together, this name was chosen to reflect both the challenge of finding hidden flaws and the collaborative nature of the solution.[1]
On August 28, 2026. 21:40 UTC
Edit summary:
Expand Project Glasswing summary and timeline (+407 words)