Skip to main content
Radar

Five or fewer confirmed AI changes, with original sources, on mornings when something changed.A free source-linked morning brief.

Start free brief
Weekly brief and archive

Weekly LLM benchmark digest

Daybreak Red does not replace Blue

OpenAI split Daybreak into two access tiers on August 10. Blue exposes GPT-5.6 Sol with safeguards tuned for defensive work. Red provides access to the separate GPT-5.6 Cyber model for advanced, authorized research. GPT-5.6 Cyber completed 95% of OpenAI’s internal advanced-cyber requests, against 2% for Sol with Blue access. That measures completion and refusal behavior, not general cyber capability. Red led on ExploitGym and OpenAI’s internal zero-day evaluation. Blue performed best in the standard 300-turn ExploitBench setting and wrote stronger vulnerability reports.

Blue is the default; Red is the exception

The limit

OpenAI had not published separate token pricing or a context-window specification for GPT-5.6 Cyber when this issue was sent. We kept its benchmark rows display-only because the public results mixed internal harnesses, budgets, and task definitions.

Codex Cloud makes the scan persistent

Model the attack surface

Codex Security builds an editable threat model so entry points, trust boundaries, sensitive data, and high-impact paths stay visible.

Reproduce before escalating

An isolated validator tests candidate vulnerabilities and records execution details and proof-of-concept artifacts.

Hand the patch to a person

Codex proposes a focused change for review; it does not modify the repository automatically.

Analysis worth opening

New on BenchLM

Archive copy reflects the rankings, prices, and availability stated when this issue was sent. Current pages may show newer evidence.

Want the next issue?

The signup form and another real sample are on the weekly brief page.

Subscribe to the weekly brief