Anthropic OSS-Fuzz Any-Crash Rate (Anthropic OSS-Fuzz crash)
Share of evaluated OSS-Fuzz entry points where the model produced at least a crash.
Benchmark score on Anthropic OSS-Fuzz crash — August 13, 2026
BenchLM mirrors the published score view for Anthropic OSS-Fuzz crash. Claude Mythos 5 leads the public snapshot at 80.0%. We do not use these results to rank models overall.
Benchmark score table (1 model)
ScoreAbout Anthropic OSS-Fuzz crash
Year
2026
Tasks
Approximately 830 OSS-Fuzz entry points
Format
Any-crash rate
Difficulty
Vulnerability discovery
This is Anthropic's internal OSS-Fuzz-derived protocol over roughly 830 entry points, not a public OSS-Fuzz leaderboard. The row is display only and preserves the provider-specific protocol name.
BenchLM freshness & provenance
Version
Anthropic OSS-Fuzz crash 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does Anthropic OSS-Fuzz crash measure?
Share of evaluated OSS-Fuzz entry points where the model produced at least a crash.
Which model scores highest on Anthropic OSS-Fuzz crash?
Claude Mythos 5 by Anthropic currently leads with a score of 80.0% on Anthropic OSS-Fuzz crash.
How many models are evaluated on Anthropic OSS-Fuzz crash?
1 AI models have been evaluated on Anthropic OSS-Fuzz crash on BenchLM.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.