Weekly LLM benchmark digest
Daybreak Red does not replace Blue
OpenAI split Daybreak into two access tiers on August 10. Blue exposes GPT-5.6 Sol with safeguards tuned for defensive work. Red provides access to the separate GPT-5.6 Cyber model for advanced, authorized research. GPT-5.6 Cyber completed 95% of OpenAI’s internal advanced-cyber requests, against 2% for Sol with Blue access. That measures completion and refusal behavior, not general cyber capability. Red led on ExploitGym and OpenAI’s internal zero-day evaluation. Blue performed best in the standard 300-turn ExploitBench setting and wrote stronger vulnerability reports.
Blue is the default; Red is the exception
The limit
OpenAI had not published separate token pricing or a context-window specification for GPT-5.6 Cyber when this issue was sent. We kept its benchmark rows display-only because the public results mixed internal harnesses, budgets, and task definitions.
Codex Cloud makes the scan persistent
Model the attack surface
Codex Security builds an editable threat model so entry points, trust boundaries, sensitive data, and high-impact paths stay visible.
Reproduce before escalating
An isolated validator tests candidate vulnerabilities and records execution details and proof-of-concept artifacts.
Hand the patch to a person
Codex proposes a focused change for review; it does not modify the repository automatically.