2 US Agencies Now Using Anthropic's Mythos to Hunt Software Flaws
CISA has joined the NSA in deploying Anthropic's offensive-security AI model Mythos to scan government code for vulnerabilities. Early results point to a large number of flaws, accelerating the shift toward AI-driven vulnerability management in critical infrastructure.
Key Takeaways
- CISA has joined the NSA in deploying Anthropic's offensive-security AI model Mythos to scan government code for vulnerabilities.
- Early results point to a large number of flaws, accelerating the shift toward AI-driven vulnerability management in critical infrastructure.
Mentioned
Key Intelligence
Key Facts
- 1CISA's Attack Surface Evaluation team is using Anthropic's AI model Mythos to scan government code repositories for vulnerabilities, according to three anonymous sources.
- 2The NSA has also been using Mythos since at least April 2026, despite the earlier Pentagon supply-chain risk designation against Anthropic.
- 3Two sources said the CISA audits have already uncovered a 'large number' of vulnerabilities, though no details on the nature or severity have been disclosed.
- 4The Pentagon designated Anthropic a supply-chain risk in February 2026 after the company refused to remove AI safeguards for autonomous weapons and domestic surveillance; a federal judge blocked the designation in March 2026.
- 5Anthropic has confidentially filed for a U.S. IPO, adding financial stakes to the success of its government collaborations.
- 6The supply-chain risk classification had previously been used for foreign companies suspected of facilitating espionage, underscoring the severity of the initial government stance.
Analysis
- Scans vast codebases faster than any human team, shrinking window of exposure
- Designed to exploit flaws, increasing likelihood of catching logic bugs traditional tools miss
- Can simulate advanced threat actor techniques, improving defensive readiness
- Offensive capability creates dual-use risk—malicious compromise could weaponize the model
- Black-box transparency: no public benchmarks or disclosure on error rates and coverage
- Over-reliance may lead to validation bias, where human auditors defer too readily to AI findings
Analysis
For cybersecurity leaders, the news that CISA is wielding Anthropic's Mythos signals more than another vendor win—it’s the emergence of AI-powered adversarial testing as a standard federal practice. With two agencies now using a model built to exploit weaknesses, not just catalog them, the community must grapple with both the speed of discovery and the unresolved risks of AI dual-use in national security.
The U.S. Cybersecurity and Infrastructure Security Agency (CISA) has begun using Anthropic's AI model Mythos to scan federal code repositories for security flaws, three sources familiar with the matter told Reuters, marking a significant step in the government's adoption of offensive AI for defensive purposes. The initiative, carried out by CISA's Attack Surface Evaluation team, has already uncovered a 'large number' of vulnerabilities, though the nature, severity, and scale of the flaws remain undisclosed. This development comes as Anthropic navigates a turbulent relationship with the U.S. government, having only months earlier been branded a supply-chain risk by the Pentagon after refusing to strip safety guardrails from its models. The early success of Mythos within both CISA and the National Security Agency (NSA) suggests a potential turning point, not only for Anthropic's government standing but also for the future of AI-driven vulnerability management in critical infrastructure.
For cybersecurity leaders, the news that CISA is wielding Anthropic's Mythos signals more than another vendor win—it’s the emergence of AI-powered adversarial testing as a standard federal practice.
The backstory is essential. In February 2026, the Pentagon designated Anthropic a supply-chain risk—a classification typically reserved for foreign firms suspected of espionage—after the company declined to modify its AI systems for autonomous weapons or domestic surveillance. A federal judge blocked that designation in March 2026, and tensions began to ease when Anthropic privately released Mythos, a specialized model designed to identify and exploit cybersecurity vulnerabilities. By April 2026, the NSA was already using Mythos, despite the earlier blacklist. CISA's confirmed usage now means at least two federal agencies are leveraging the same AI tool to proactively hunt for flaws that could be exploited by adversaries. This dual-agency adoption is noteworthy given how rapidly it followed the lifting of the legal cloud.
For cybersecurity practitioners, the CISA-Mythos project underscores a critical shift: AI is moving beyond alert triage and log analysis into full-fledged vulnerability discovery at scale. Traditional static and dynamic application security testing tools remain vital, but they often miss the nuanced, context-dependent logic flaws that an advanced language model might detect. Mythos' design to exploit vulnerabilities, not just identify them, suggests an adversary-simulation capability that could root out weaknesses human red teams overlook. However, this power cuts both ways. If such a model—or its training data, prompts, or weights—were compromised or repurposed by malicious actors, the same capability could supercharge zero-day hunting by nation-states or criminal groups. The lack of transparency around the CISA audits—no details on the types of vulnerabilities found, the volume of code reviewed, or the model's performance metrics—leaves the cybersecurity community with more questions than answers about the reliability and broader implications of AI-driven government security assessments.
What to Watch
Another layer is the commercial backdrop. Anthropic has confidentially filed for an initial public offering, and government contracts are often a bellwether for enterprise trust. A successful, demonstrable impact at CISA could accelerate adoption across other federal departments and critical infrastructure operators, potentially opening a massive new revenue stream. Conversely, any publicized failure—such as a critical flaw missed by the AI or a breach traced back to model misuse—would not only damage Anthropic's reputation but also reinforce calls for stringent AI regulation in cybersecurity. The company's earlier stance on safety, while principled, highlights the friction between its founding mission of responsible AI development and the operational realities of defense work. The Mythos narrative thus becomes a delicate balance: proving the technology's value without compromising ethical boundaries or enabling dual-use harm.
Looking ahead, the CISA initiative is likely a pilot that will shape broader policy. If the vulnerability yield reported by anonymous sources can be substantiated, expect a push to embed AI-based code review into the software supply chain security mandates enforced by Executive Order 14028 and related OMB guidance. But without public benchmarks, independent audits, and clear rules of engagement, the program could become a black box that erodes trust among developers and privacy advocates. The coming months may reveal whether Mythos evolves from a quiet government tool into a standard component of the national cybersecurity apparatus, setting a precedent for how AI safety and national security can coexist in the defense of digital infrastructure.
Timeline
Timeline
Pentagon designates Anthropic as supply-chain risk
After Anthropic refused to remove AI safeguards for autonomous weapons and domestic surveillance, the Pentagon imposes a classification typically used for foreign espionage suspects.
Federal judge blocks Pentagon designation
A judge issues an injunction, temporarily halting the supply-chain risk label against Anthropic.
NSA begins using Mythos
The National Security Agency starts using Anthropic's Mythos for vulnerability assessments, despite the recent blacklist.
CISA project with Mythos reported
Sources reveal that CISA's Attack Surface Evaluation team is actively using Mythos to scan government code repositories, with multiple vulnerabilities already discovered.
Sources
Sources
Based on 2 source articles- azerbaijannews.netCISA taps Anthropic Mythos to find software vulnerabilitiesJul 9, 2026
- 2lt.com.auCISA taps Anthropic Mythos to find software vulnerabilitiesJul 9, 2026
Cite This Page
"2 US Agencies Now Using Anthropic's Mythos to Hunt Software Flaws." Cyber Intelligence Brief, July 23, 2026. https://getcyberbrief.com/story/cisa-nsa-anthropic-mythos-vulnerability-scanning-ai
How we covered this story
Every story in our cybersecurity coverage is assembled from multiple primary sources, cross-referenced for factual consistency, and scored along three independent dimensions: sentiment, operational impact, and source-cluster confidence. Single-source rumors and unverifiable claims do not pass our editorial gate. When a story shows "Verified by N sources" with N≥2, the development is independently corroborated; when N=1, we mark it explicitly so readers can weigh the signal accordingly.
Impact scoring uses a 1-10 scale weighted toward regulatory, financial, and operational consequence rather than coverage volume. A topic that runs in every outlet but moves no real decisions ranks lower than a niche regulatory filing that reshapes how operators in the cybersecurity space have to behave. Read our full methodology for the scoring rubric, our glossary for term definitions, and our trends index for the longitudinal view across the beat.
Sources are only linked to a story once they clear our classification pipeline at a minimum 35 percent relevance threshold. According to that methodology, reviewed July 2026, this follows multi-source corroboration standards recommended by journalism research bodies such as the Reuters Institute for the Study of Journalism.
See something wrong in this story — a wrong fact, a broken source link, a misattributed entity? Report a data issue.
| Signal on this page | What it tells you |
|---|---|
| Verified by N sources | Independent corroboration count. N≥2 is our confidence floor; N=1 is marked explicitly. |
| Impact score (1-10) | Regulatory + financial + operational weight. 8+ signals an experienced-operator action item. |
| Sentiment | Five-tier classification trained on labeled cybersecurity-specific corpora. |
| Timeline | Where applicable, the related-events sequence that contextualizes today's development. |