Skip to main content
BenchLM
Data

BenchLM Data

AI benchmarks, in your tools.

Query current rankings and benchmark results through REST or MCP. Add recorded history when you need it.

Compare plans

Create an account or sign in

New accounts start on Free. No card required.

We’ll email you a one-time code.

Current data is free. Add history when you need it.

Compare Data plans

Free

$0USD/month

1,00010 requests/minute

Current dataCurrent BenchAlign rankings, benchmark results, and catalogs.

Start free

Data Pro

$49USD/month

100,00060 requests/minute

Three rolling calendar monthsAvailable dated history, measured by UTC date.

Choose Data Pro

Research

$249USD/month

500,00060 requests/minute

All retained eligible historyOlder and undated records, where available. No history age cap.

Choose Research

Coverage varies by source, model, and date. Check it free before subscribing. Successful reads count toward your allowance; usage and coverage checks are free.

Also want model-change alerts?

Radar + Data includes Radar and Pro’s history window for $59 USD/month.

Review the bundle

Radar and Data have separate API keys and allowances. Existing Radar subscriptions keep their terms until period end before bundle checkout opens. Paid plans cover queries and internal analysis. See commercial-use terms.

Follow model performance over time.

A current ranking helps you choose today. Historical records let you inspect how the published results changed.

Compare published results across dates.

Follow Arena overall ratings, ranks, and confidence intervals across retained snapshots. Source-reported benchmark runs date back to January 2025. Research includes eligible records beyond 12 months.

Check the dates available

Keep the evidence with the score.

Inspect the exact source model, original units, and reported dates. Saved BenchAlign rankings retain their scoring version, so a methodology change stays visible in your analysis.

Read the scoring methodology

Put the records into your workflow.

Ask an MCP client how a model's rank changed, or query REST from a notebook or application. One API key works with both. Resolve model IDs and check coverage before requesting history.

See how to connect

Ask a question. Get the data behind it.

Use an MCP client for natural-language questions, or call the REST API from your own application.

“Which models lead the current coding ranking?”

All plans

Read final BenchAlign scores and ranks for a published surface. Search the catalogs and filter benchmarks by capability, including coding, reasoning, or math.

benchlm_list_current_rankings

“How has this model’s BenchAlign rank changed?”

Pro · Research

Read recorded scores and ranks over time, with the saved methodology version. Retrieve a ranking snapshot to compare the models recorded at a date.

benchlm_rank_history

“When was this benchmark first recorded?”

Pro · Research

Find catalog additions, removals, and reintroductions in the archive. Dates describe saved repository states, rather than a publisher’s release date.

benchlm_benchmark_events
See all available queries
Data coverage, plans, and MCP tools
QueryFields and coveragePlanMCP tool
Account usageCurrent plan, successful and remaining monthly reads, request-rate usage, and both reset times. Usage checks do not consume a monthly read.Freebenchlm_get_usage
Current BenchAlign rankingsFinal scores, published ranks, evidence labels, interval bounds, scoring date, and methodology version for each available surface.Freebenchlm_list_current_rankings
Current model rankingsOne model’s final score and rank across every available BenchAlign surface, with explicit gaps where it is unranked.Freebenchlm_get_current_model_rankings
Model catalogNames, creators, canonical IDs, source links, and sourced release dates with their date precision. Search by name or filter by creator.Freebenchlm_list_models / benchlm_get_model
Benchmark catalogNames, catalog IDs, source references, and curator categories and tags. Search by name or filter by category, tag, and relationship; discovery labels do not establish result coverage.Freebenchlm_list_benchmarks
Current benchmark resultsAll admitted current source results, with native units and exact profiles. Examples include Epoch AI, BullshitBench v2, and Arena overall. Search by category or model name.Freebenchlm_list_current_benchmark_results
Benchmark result coverageAvailable source boards, row counts, metrics, units, and source terms. Coverage checks do not use a read.Freebenchlm_benchmark_result_coverage
Published evaluation runsSource-reported benchmark runs for an exact board and model profile. Run dates are distinct from archive commit dates; undated records are identified.Pro · Researchbenchlm_benchmark_result_history
Published benchmark snapshotsFollow an exact source model’s published rating, rank, and confidence interval over time. Arena overall history preserves three scoring methods and publisher dates.Pro · Researchbenchlm_benchmark_result_snapshot_history
Historical coverageAvailable methods, retained date windows, and delivery coverage. Check this before requesting history.Freebenchlm_history_coverage
BenchAlign historySaved BenchLM scores, ranks, evidence labels, ranking surfaces, and methodology versions for an exact model.Pro · Researchbenchlm_rank_history
Ranking at a dateStored BenchAlign ranking rows as of a UTC date. The response identifies incomplete coverage and preserves the saved scoring version.Pro · Researchbenchlm_ranking_snapshot
Benchmark historyCleared measurements for an exact model, benchmark, and protocol. Availability differs by source; this is not every score on the website.Pro · Researchbenchlm_benchmark_history
Benchmark changesFirst recorded catalog definitions and results, additions, removals, and reintroductions where retained.Pro · Researchbenchlm_benchmark_events

Current ranking responses contain final BenchAlign outputs. Benchmark result queries contain source scores, including GPQA Diamond, SWE-bench Verified, and FrontierMath evaluated by Epoch AI. BullshitBench v2 includes judge ratings on a 0–2 scale, refusal rates, and published ranks from one pinned snapshot. Source-reported runs retain their original units and dates; artifact publication timestamps do not establish run dates. Arena overall snapshot history contains the publisher’s ratings, ranks, and confidence intervals across retained dates, with scoring methods kept separate. Benchmark measurements and evaluation runs do not establish a leaderboard rank. Missing records stay missing.

Connect your tools.

  1. 01

    Create your Data account.

    Sign in by email. Free starts without a payment method.

  2. 02

    Create an API key.

    Keep it in your application’s or MCP client’s secret settings.

  3. 03

    Read current data or recorded history.

    Use REST, or connect an MCP client with Streamable HTTP.

Open the endpoint reference
Read the current coding ranking
curl 'https://data.benchlm.ai/v1/rankings/current?surface=coding&limit=10' \
  -H 'Authorization: Bearer YOUR_DATA_KEY'
MCP endpointhttps://data.benchlm.ai/mcp

Send Authorization: Bearer YOUR_DATA_KEY. Supported MCP protocol versions: 2026-07-28, 2025-11-25, and 2025-06-18.

View the MCP request
Call the current ranking MCP tool
curl 'https://data.benchlm.ai/mcp' \
  -H 'Authorization: Bearer YOUR_DATA_KEY' \
  -H 'Content-Type: application/json' \
  -H 'Accept: application/json, text/event-stream' \
  -H 'MCP-Protocol-Version: 2025-11-25' \
  --data '{"jsonrpc":"2.0","id":1,"method":"tools/call","params":{"name":"benchlm_list_current_rankings","arguments":{"surface":"coding","limit":10}}}'

Questions

Check the records you need before you choose a plan. Coverage and usage checks are free, including when your monthly reads are exhausted.

benchlm_get_usage
Does every benchmark have 12 months of history?

No. Retained Arena overall snapshots date back to August 2024, and source-reported benchmark runs reach January 2025. Other sources and saved BenchAlign rankings begin later. Research removes the history age cap; it does not fill missing records. Check the dates available for your exact source and model.

Why choose Research over Data Pro?

Research includes all retained eligible history and 500,000 successful reads per month for $249 USD. Data Pro includes three rolling calendar months and 100,000 reads for $49 USD. Both allow 60 requests per minute. Choose Research when older records or a larger monthly allowance matter to your work.

Can I check coverage before paying?

Yes. A free account gives you current rankings, source results, model and benchmark catalogs, and coverage checks. The coverage tools identify available sources and date windows without using your monthly reads. Check your model and benchmark first, then choose the plan that includes the history you need.

What counts as a read?

One successful data page uses one read. Initialization, discovery, coverage and usage checks, and unsuccessful queries do not use the monthly allowance. All keys on an account share usage across REST and MCP.

When does my allowance reset?

Free renews on your account’s activation anniversary each month. Pro and Research follow your subscription period. Short months use their last day. Your account and usage responses show the exact monthly and request-rate reset times.

What happens when I reach a limit?

The API and MCP identify the monthly read allowance or request-rate limit, with a reset time and retry delay. Free accounts get a Data Pro upgrade link. Eligible Pro accounts can upgrade to Research with a prorated price difference and a fresh 500,000-read allowance for the remaining subscription period. You can receive usage emails at 75%, 90%, and 100%; manage them in your account. Usage and coverage checks stay available.

How much history can one query return?

Pro covers three rolling calendar months by UTC date; Research has no age cap on retained eligible records. Coverage varies by source. Historical pages contain up to 500 rows. Saved-state history uses 90-day windows and cursors. Source-run and publisher snapshot history take an exact board and source model ID, with date filters of up to 90 days. Snapshot history supports cursors and preserves publisher dates; run history preserves reported evaluation times. Check the relevant coverage tool first.

Commercial display or redistribution? Ask about commercial use.

Free, Pro, and Research are access plans for queries and internal analysis. Paid plans add reads and available history; they do not include commercial display or redistribution permission.

Publishing BenchLM-owned outputs or the compiled feed in a customer-facing commercial website, app, or product needs a separate display agreement. Redistributing that material through an API, downloads, or resale needs separately agreed scope. Request commercial-use permission.

Individual quotations and attributed embeds keep their existing permissions. Publisher measurements keep their own commercial-use and redistribution rights. Retain attribution, source notices, and applicable terms. Read the Data access terms.

Make the next comparison
with the history behind it.

Research gives you all retained eligible history for $249 USD/month. Check your sources and dates with a free account first.

613 benchmark pages, 882 model profiles, and 11 categories describe the website catalog. API coverage varies by source.
Looking for the website’s existing downloads and their terms? Open the Data hub and downloads.