# FrontierCyber

> Independent evaluation of AI agents against vulnerable real-world systems in dynamic environments.

Canonical page: https://benchlm.ai/benchmarks/frontiercyber

- Category: [Agentic](/agentic)
- Last updated: September 27, 2026

## About FrontierCyber

- Year: 2026
- Tasks: 197 dynamic cyber tasks
- Format: Tasks solved
- Difficulty: Easy through elite cyber operations
- Paper: [FrontierCyber](https://www.irregular.com/research/frontiercyber)

FrontierCyber tests autonomous cyber performance on live, realistic systems and reports solved tasks across difficulty bands. BenchLM stores the independent snapshot as display-only external evidence.

FrontierCyber is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (2 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [GPT-6 Astra](/models/gpt-6-astra) | OpenAI | 38.1% |
| 2 | [GPT-5.6 Sol](/models/gpt-5-6-sol) | OpenAI | 9.6% |

## FAQ

### What does FrontierCyber measure?

Independent evaluation of AI agents against vulnerable real-world systems in dynamic environments.

### Which model scores highest on FrontierCyber?

GPT-6 Astra by OpenAI currently leads with a score of 38.1% on FrontierCyber.

### How many models are evaluated on FrontierCyber?

2 AI models have been evaluated on FrontierCyber on BenchLM.

### Does FrontierCyber affect BenchLM's overall score?

Not directly. FrontierCyber is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Compare Top Models on FrontierCyber

- [GPT-6 Astra vs GPT-5.6 Sol](/compare/gpt-5-6-sol-vs-gpt-6-astra)
