Skip to main content
BenchLM
Data

BenchLM Data

AI benchmarks.
12+ months of history.

Explore 649 benchmark pages and 784 model profiles. Query current rankings and eligible historical records through REST or MCP.

Arena overall snapshots date back to August 2024. Coverage varies by source, model, and date.

Check source and date coverage free before subscribing.

649
benchmark pages
784
model profiles
11
capability categories
Aug 2024
earliest retained Arena overall snapshots

Catalog totals describe the website. API results and historical coverage vary by source.

Follow model performance over time.

A current ranking helps you choose today. Historical records let you inspect how the published results changed.

Compare published results across dates.

Follow Arena overall ratings, ranks, and confidence intervals across retained snapshots. Source-reported benchmark runs date back to January 2025. Research includes eligible records beyond 12 months.

Check the dates available

Keep the evidence with the score.

Inspect the exact source model, original units, and reported dates. Saved BenchAlign rankings retain their scoring version, so a methodology change stays visible in your analysis.

Read the scoring methodology

Put the records into your workflow.

Ask an MCP client how a model's rank changed, or query REST from a notebook or application. One API key works with both. Resolve model IDs and check coverage before requesting history.

See how to connect

Choose the history your research needs.

All plans use the same API key for REST and MCP.

Current data

Free

$0 USD/month

Look up models and benchmarks. Read current BenchAlign rankings and licensed benchmark results.

Create a free account
Monthly reads
1,000
Request rate
10/minute
  • Current final BenchAlign scores and ranks
  • All admitted current benchmark results
  • Model and benchmark catalogs, with search and category filters
  • Usage, reset dates, and historical coverage checks

No payment details required. Historical queries require Pro or Research.

Current + entire available archive

Research

$249 USD/month

Investigate changes beyond the last three months. Get the entire eligible archive and 5× the monthly reads of Pro.

Get Research access
Monthly reads
500,000
Request rate
60/minute
  • Everything in Data Pro, with no history age cap
  • Older records, including those beyond 12 months where retained
  • Eligible undated source records
  • Published snapshots, evaluation runs, and saved ranking history

Coverage varies by source, model, and date. Check it free in your account before subscribing. Checkout opens when available.

Current + three months of history

Data Pro

$49 USD/month

Query retained rankings, publisher snapshots, and dated runs from the last three rolling calendar months.

Get Data Pro
Monthly reads
100,000
Request rate
60/minute
  • Everything in Free, with 100× the monthly reads
  • Three rolling calendar months, measured by UTC date
  • Published Arena ratings and ranks over time
  • Source-reported evaluation runs for an exact model profile
  • Ranking snapshots at a date
  • Covered benchmark measurements and catalog changes

Need records older than three months? Choose Research. Check coverage before subscribing.

Access plans cover queries and internal analysis. Commercial display and redistribution need a separate license. See usage terms.

Also want model-change alerts?

Radar + Data includes Radar and Pro’s three-month history window.

$59 USD/month

See the bundle

Radar and Data remain separate products, with separate API keys and allowances. Existing monthly, annual, and legacy Radar subscriptions keep their terms until period end before bundle checkout opens.

Ask a question. Get the data behind it.

Use an MCP client for natural-language questions, or call the REST API from your own application.

All plans

“Which models lead the current coding ranking?”

Read final BenchAlign scores and ranks for a published surface. Search the catalogs and filter benchmarks by capability, including coding, reasoning, or math.

benchlm_list_current_rankings

Pro · Research

“How has this model’s BenchAlign rank changed?”

Read recorded scores and ranks over time, with the saved methodology version. Retrieve a ranking snapshot to compare the models recorded at a date.

benchlm_rank_history

Pro · Research

“When was this benchmark first recorded?”

Find catalog additions, removals, and reintroductions in the archive. Dates describe saved repository states, rather than a publisher’s release date.

benchlm_benchmark_events
See all available queries
Data coverage, plans, and MCP tools
QueryFields and coveragePlanMCP tool
Account usageCurrent plan, successful and remaining monthly reads, request-rate usage, and both reset times. Usage checks do not consume a monthly read.Freebenchlm_get_usage
Current BenchAlign rankingsFinal scores, published ranks, evidence labels, interval bounds, scoring date, and methodology version for each available surface.Freebenchlm_list_current_rankings
Current model rankingsOne model’s final score and rank across every available BenchAlign surface, with explicit gaps where it is unranked.Freebenchlm_get_current_model_rankings
Model catalogNames, creators, canonical IDs, source links, and sourced release dates with their date precision. Search by name or filter by creator.Freebenchlm_list_models / benchlm_get_model
Benchmark catalogNames, catalog IDs, source references, and curator categories and tags. Search by name or filter by category, tag, and relationship; discovery labels do not establish result coverage.Freebenchlm_list_benchmarks
Current benchmark resultsAll admitted current source results, with native units and exact profiles. Examples include Epoch AI, BullshitBench v2, and Arena overall. Search by category or model name.Freebenchlm_list_current_benchmark_results
Benchmark result coverageAvailable source boards, row counts, metrics, units, and source licenses. Coverage checks do not use a read.Freebenchlm_benchmark_result_coverage
Published evaluation runsSource-reported benchmark runs for an exact board and model profile. Run dates are distinct from archive commit dates; undated records are identified.Pro · Researchbenchlm_benchmark_result_history
Published benchmark snapshotsFollow an exact source model’s published rating, rank, and confidence interval over time. Arena overall history preserves three scoring methods and publisher dates.Pro · Researchbenchlm_benchmark_result_snapshot_history
Historical coverageAvailable methods, retained date windows, and delivery coverage. Check this before requesting history.Freebenchlm_history_coverage
BenchAlign historySaved BenchLM scores, ranks, evidence labels, ranking surfaces, and methodology versions for an exact model.Pro · Researchbenchlm_rank_history
Ranking at a dateStored BenchAlign ranking rows as of a UTC date. The response identifies incomplete coverage and preserves the saved scoring version.Pro · Researchbenchlm_ranking_snapshot
Benchmark historyCleared measurements for an exact model, benchmark, and protocol. Availability differs by source; this is not every score on the website.Pro · Researchbenchlm_benchmark_history
Benchmark changesFirst recorded catalog definitions and results, additions, removals, and reintroductions where retained.Pro · Researchbenchlm_benchmark_events

Current ranking responses contain final BenchAlign outputs. Benchmark result queries contain licensed source scores, including GPQA Diamond, SWE-bench Verified, and FrontierMath evaluated by Epoch AI. BullshitBench v2 includes judge ratings on a 0–2 scale, refusal rates, and published ranks from one pinned snapshot. Source-reported runs retain their original units and dates; artifact publication timestamps do not establish run dates. Arena overall snapshot history contains the publisher’s ratings, ranks, and confidence intervals across retained dates, with scoring methods kept separate. Benchmark measurements and evaluation runs do not establish a leaderboard rank. Missing records stay missing.

Connect your tools.

  1. 01

    Create your Data account.

    Sign in by email. Free starts without a payment method.

  2. 02

    Create an API key.

    Keep it in your application’s or MCP client’s secret settings.

  3. 03

    Read current data or recorded history.

    Use REST, or connect an MCP client with Streamable HTTP.

Open the endpoint reference
Read the current coding ranking
curl 'https://data.benchlm.ai/v1/rankings/current?surface=coding&limit=10' \
  -H 'Authorization: Bearer YOUR_DATA_KEY'
MCP endpointhttps://data.benchlm.ai/mcp

Send Authorization: Bearer YOUR_DATA_KEY. Supported MCP protocol versions: 2026-07-28, 2025-11-25, and 2025-06-18.

View the MCP request
Call the current ranking MCP tool
curl 'https://data.benchlm.ai/mcp' \
  -H 'Authorization: Bearer YOUR_DATA_KEY' \
  -H 'Content-Type: application/json' \
  -H 'Accept: application/json, text/event-stream' \
  -H 'MCP-Protocol-Version: 2025-11-25' \
  --data '{"jsonrpc":"2.0","id":1,"method":"tools/call","params":{"name":"benchlm_list_current_rankings","arguments":{"surface":"coding","limit":10}}}'

Questions

Check the records you need before you choose a plan. Coverage and usage checks are free, including when your monthly reads are exhausted.

benchlm_get_usage
Does every benchmark have 12 months of history?

No. Retained Arena overall snapshots date back to August 2024, and source-reported benchmark runs reach January 2025. Other sources and saved BenchAlign rankings begin later. Research removes the history age cap; it does not fill missing records. Check the dates available for your exact source and model.

Why choose Research over Data Pro?

Research includes all retained eligible history and 500,000 successful reads per month for $249 USD. Data Pro includes three rolling calendar months and 100,000 reads for $49 USD. Both allow 60 requests per minute. Choose Research when older records or a larger monthly allowance matter to your work.

Can I check coverage before paying?

Yes. A free account gives you current rankings, licensed source results, model and benchmark catalogs, and coverage checks. The coverage tools identify available sources and date windows without using your monthly reads. Check your model and benchmark first, then choose the plan that includes the history you need.

What counts as a read?

One successful data page uses one read. Initialization, discovery, coverage and usage checks, and unsuccessful queries do not use the monthly allowance. All keys on an account share usage across REST and MCP.

When does my allowance reset?

Free renews on your account’s activation anniversary each month. Pro and Research follow your subscription period. Short months use their last day. Your account and usage responses show the exact monthly and request-rate reset times.

What happens when I reach a limit?

The API and MCP identify the monthly read allowance or request-rate limit, with a reset time and retry delay. Free accounts get a Data Pro upgrade link. Eligible Pro accounts can upgrade to Research with a prorated price difference and a fresh 500,000-read allowance for the remaining subscription period. You can receive usage emails at 75%, 90%, and 100%; manage them in your account. Usage and coverage checks stay available.

How much history can one query return?

Pro covers three rolling calendar months by UTC date; Research has no age cap on retained eligible records. Coverage varies by source. Historical pages contain up to 500 rows. Saved-state history uses 90-day windows and cursors. Source-run and publisher snapshot history take an exact board and source model ID, with date filters of up to 90 days. Snapshot history supports cursors and preserves publisher dates; run history preserves reported evaluation times. Check the relevant coverage tool first.

Commercial display or redistribution? Request a license.

Free, Pro, and Research are access plans for queries and internal analysis. Paid plans add reads and available history; they do not include a commercial display or redistribution license.

Publishing BenchLM-owned outputs or the compiled feed in a customer-facing commercial website, app, or product needs a separately agreed display license. Redistributing that material through an API, downloads, or resale needs separately agreed scope. Request a commercial license.

Individual quotations and attributed embeds keep their existing permissions. Independently licensed publisher measurements keep their own commercial-use and redistribution rights. Retain attribution, source notices, and licenses. Read the Data access terms.

Make the next comparison
with the history behind it.

Research gives you all retained eligible history for $249 USD/month. Check your sources and dates with a free account first.

Looking for the website’s existing downloads and their terms? Open the Data hub and downloads.