Skip to main content
Radar

Five or fewer confirmed AI changes, with original sources, on mornings when something changed.A free source-linked morning brief.

Start free brief

Sent every week

The model changes worth your attention

Rank moves, price changes, launches, and one clear decision note. One email every week.

Get the next weekly brief

Read the sample below first. Subscribe when the format earns the space in your inbox.

Join 2,000+ readers.

What arrives

Useful change, not a news firehose

Each issue is built for people choosing, buying, or operating language models.

Rank moves with the evidence attached

See who moved, what changed, and whether the public benchmark coverage supports the move.

Material API price changes

Track pricing that could change a production shortlist, without filling your inbox with every minor update.

Release and availability context

Separate a launch headline from the access limits, evidence gaps, and reasons to wait.

One decision note

Leave with a model to consider, a lower-cost alternative, or a clear reason not to switch yet.

Real sample issue

See the whole format

This is the editorial content of a sent issue, adapted from email to the web.

Weekly LLM benchmark digest

Daybreak Red does not replace Blue

OpenAI split Daybreak into two access tiers on August 10. Blue exposes GPT-5.6 Sol with safeguards tuned for defensive work. Red provides access to the separate GPT-5.6 Cyber model for advanced, authorized research. GPT-5.6 Cyber completed 95% of OpenAI’s internal advanced-cyber requests, against 2% for Sol with Blue access. That measures completion and refusal behavior, not general cyber capability. Red led on ExploitGym and OpenAI’s internal zero-day evaluation. Blue performed best in the standard 300-turn ExploitBench setting and wrote stronger vulnerability reports.

Blue is the default; Red is the exception

The limit

OpenAI had not published separate token pricing or a context-window specification for GPT-5.6 Cyber when this issue was sent. We kept its benchmark rows display-only because the public results mixed internal harnesses, budgets, and task definitions.

Codex Cloud makes the scan persistent

Model the attack surface

Codex Security builds an editable threat model so entry points, trust boundaries, sensitive data, and high-impact paths stay visible.

Reproduce before escalating

An isolated validator tests candidate vulnerabilities and records execution details and proof-of-concept artifacts.

Hand the patch to a person

Codex proposes a focused change for review; it does not modify the repository automatically.

Analysis worth opening

New on BenchLM

Archive copy reflects the rankings, prices, and availability stated when this issue was sent. Current pages may show newer evidence.

Past issues

Read before you subscribe

Every archived issue is dated. Its claims stay tied to the evidence available on that send date.

Sponsored editions

Paid partner briefs are labeled before the first partner link. Provider-run claims stay attributed to the sponsor.

Subscribe to the weekly brief