# FrontierCode 1.1 Main

> Cognition's 100-task software-engineering benchmark for whether coding agents produce mergeable, production-quality pull requests, scored for correctness, tests, scope, style, and maintainability through maintainer-authored rubrics.

Canonical page: https://benchlm.ai/benchmarks/frontiercode

- Category: [Coding](/coding)
- Last updated: September 4, 2026 snapshot

## About FrontierCode 1.1 Main

- Year: 2026
- Tasks: 100 private Main tasks (150 in Extended)
- Format: Repository task completion with maintainer rubrics
- Difficulty: Frontier coding-agent quality
- Paper: [FrontierCode leaderboard](https://cognition.com/frontiercode)

FrontierCode 1.1 Main uses 100 of the benchmark's 150 private software-engineering tasks. The leaderboard reports the best-performing published reasoning effort for each model-agent row. We keep the results display-only because each row combines a model with an agent harness and the private tasks cannot be independently rerun from the public artifact.

FrontierCode 1.1 Main is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (8 models)

| Rank | Model | Configuration | Creator | Score |
|------|-------|---------------|---------|-------|
| 1 | [Claude Fable 5](/models/claude-fable) | claude-code | Anthropic | 53.5% |
| 2 | [Claude Opus 4.8](/models/claude-opus-4-8) | claude-code | Anthropic | 46.5% |
| 3 | [GPT-5.5](/models/gpt-5-5) | codex | OpenAI | 43.0% |
| 4 | [Claude Sonnet 5](/models/claude-sonnet-5) | claude-code | Anthropic | 42.7% |
| 5 | [Claude Opus 4.7](/models/claude-opus-4-7) | claude-code | Anthropic | 38.5% |
| 6 | [GPT-5.4 mini](/models/gpt-5-4-mini) | codex | OpenAI | 27.0% |
| 7 | [Claude Opus 4.6](/models/claude-opus-4-6) | claude-code | Anthropic | 26.9% |
| 8 | [Claude Sonnet 4.6](/models/claude-sonnet-4-6) | claude-code | Anthropic | 24.3% |

## FAQ

### What does FrontierCode 1.1 Main measure?

Cognition's 100-task software-engineering benchmark for whether coding agents produce mergeable, production-quality pull requests, scored for correctness, tests, scope, style, and maintainability through maintainer-authored rubrics.

### Which model leads the published FrontierCode 1.1 Main snapshot?

Claude Fable 5 currently leads the published FrontierCode 1.1 Main snapshot with a score of 53.5%.

### How many models are evaluated on FrontierCode 1.1 Main?

The September 4, 2026 snapshot contains 8 AI models.

### Does FrontierCode 1.1 Main affect BenchLM's overall score?

Not directly. FrontierCode 1.1 Main is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
