# Daybreak Red does not replace Blue

Sent August 13, 2026

OpenAI split Daybreak into two access tiers on August 10. Blue exposes GPT-5.6 Sol with safeguards tuned for defensive work. Red provides access to the separate GPT-5.6 Cyber model for advanced, authorized research. GPT-5.6 Cyber completed 95% of OpenAI’s internal advanced-cyber requests, against 2% for Sol with Blue access. That measures completion and refusal behavior, not general cyber capability. Red led on ExploitGym and OpenAI’s internal zero-day evaluation. Blue performed best in the standard 300-turn ExploitBench setting and wrote stronger vulnerability reports.

## Blue is the default; Red is the exception

- [Daybreak Blue](https://benchlm.ai/cybersecurity): GPT-5.6 Sol is the starting point for vulnerability discovery, secure code review, malware analysis, incident response, and patch validation. It loses when authorized work needs exploit development or fewer refusals on high-risk dual-use tasks.
- [Daybreak Red](https://benchlm.ai/cybersecurity): GPT-5.6 Cyber is for authorized vulnerability research, exploit validation, penetration testing, and red teaming. It uses more tokens, does not win every cyber task, and requires verification, monitoring, scoped controls, and review.

## The limit

OpenAI had not published separate token pricing or a context-window specification for GPT-5.6 Cyber when this issue was sent. We kept its benchmark rows display-only because the public results mixed internal harnesses, budgets, and task definitions.

## Codex Cloud makes the scan persistent

- Model the attack surface: Codex Security builds an editable threat model so entry points, trust boundaries, sensitive data, and high-impact paths stay visible.
- Reproduce before escalating: An isolated validator tests candidate vulnerabilities and records execution details and proof-of-concept artifacts.
- Hand the patch to a person: Codex proposes a focused change for review; it does not modify the repository automatically.

## Analysis worth opening

- [We measure the session. They sell the week.](https://benchlm.ai/blog/posts/we-measure-the-session-they-sell-the-week): Why persistent agent systems need a fixture that counts interventions, silent wrong writes, and credential scope across a week.
- [How to monitor OpenAI API changes](https://benchlm.ai/blog/posts/monitor-openai-api-changes): The smallest monitoring stack for announcements, deprecations, pricing, incidents, callable IDs, and detailed changelog entries.

## New on BenchLM

- [Pay for the monitoring, not another tab](https://benchlm.ai/radar): Free Radar Brief sends confirmed source-linked changes on qualifying mornings. Radar Pro adds the complete supported event stream, focused notifications, the Event API, and remote MCP. Current pricing and trial eligibility appear before checkout.

Archive copy reflects the rankings, prices, and availability stated when this issue was sent. Current pages may show newer evidence.

Canonical page: https://benchlm.ai/newsletter/issues/2026-08-13
