Free
$0USD/month
1,00010 requests/minute
Current dataCurrent BenchAlign rankings, benchmark results, and catalogs.
Start freeAI benchmarks, in your tools.
Query current rankings and benchmark results through REST or MCP. Add recorded history when you need it.
Compare plansNew accounts start on Free. No card required.
We’ll email you a one-time code.
Current data is free. Add history when you need it.
$0USD/month
1,00010 requests/minute
Current dataCurrent BenchAlign rankings, benchmark results, and catalogs.
Start free$49USD/month
100,00060 requests/minute
Three rolling calendar monthsAvailable dated history, measured by UTC date.
Choose Data Pro$249USD/month
500,00060 requests/minute
All retained eligible historyOlder and undated records, where available. No history age cap.
Choose ResearchCoverage varies by source, model, and date. Check it free before subscribing. Successful reads count toward your allowance; usage and coverage checks are free.
Radar + Data includes Radar and Pro’s history window for $59 USD/month.
Radar and Data have separate API keys and allowances. Existing Radar subscriptions keep their terms until period end before bundle checkout opens. Paid plans cover queries and internal analysis. See commercial-use terms.
A current ranking helps you choose today. Historical records let you inspect how the published results changed.
Follow Arena overall ratings, ranks, and confidence intervals across retained snapshots. Source-reported benchmark runs date back to January 2025. Research includes eligible records beyond 12 months.
Check the dates availableInspect the exact source model, original units, and reported dates. Saved BenchAlign rankings retain their scoring version, so a methodology change stays visible in your analysis.
Read the scoring methodologyAsk an MCP client how a model's rank changed, or query REST from a notebook or application. One API key works with both. Resolve model IDs and check coverage before requesting history.
See how to connectUse an MCP client for natural-language questions, or call the REST API from your own application.
All plans
Read final BenchAlign scores and ranks for a published surface. Search the catalogs and filter benchmarks by capability, including coding, reasoning, or math.
benchlm_list_current_rankingsPro · Research
Read recorded scores and ranks over time, with the saved methodology version. Retrieve a ranking snapshot to compare the models recorded at a date.
benchlm_rank_historyPro · Research
Find catalog additions, removals, and reintroductions in the archive. Dates describe saved repository states, rather than a publisher’s release date.
benchlm_benchmark_events| Query | Fields and coverage | Plan | MCP tool |
|---|---|---|---|
| Account usage | Current plan, successful and remaining monthly reads, request-rate usage, and both reset times. Usage checks do not consume a monthly read. | Free | benchlm_get_usage |
| Current BenchAlign rankings | Final scores, published ranks, evidence labels, interval bounds, scoring date, and methodology version for each available surface. | Free | benchlm_list_current_rankings |
| Current model rankings | One model’s final score and rank across every available BenchAlign surface, with explicit gaps where it is unranked. | Free | benchlm_get_current_model_rankings |
| Model catalog | Names, creators, canonical IDs, source links, and sourced release dates with their date precision. Search by name or filter by creator. | Free | benchlm_list_models / benchlm_get_model |
| Benchmark catalog | Names, catalog IDs, source references, and curator categories and tags. Search by name or filter by category, tag, and relationship; discovery labels do not establish result coverage. | Free | benchlm_list_benchmarks |
| Current benchmark results | All admitted current source results, with native units and exact profiles. Examples include Epoch AI, BullshitBench v2, and Arena overall. Search by category or model name. | Free | benchlm_list_current_benchmark_results |
| Benchmark result coverage | Available source boards, row counts, metrics, units, and source terms. Coverage checks do not use a read. | Free | benchlm_benchmark_result_coverage |
| Published evaluation runs | Source-reported benchmark runs for an exact board and model profile. Run dates are distinct from archive commit dates; undated records are identified. | Pro · Research | benchlm_benchmark_result_history |
| Published benchmark snapshots | Follow an exact source model’s published rating, rank, and confidence interval over time. Arena overall history preserves three scoring methods and publisher dates. | Pro · Research | benchlm_benchmark_result_snapshot_history |
| Historical coverage | Available methods, retained date windows, and delivery coverage. Check this before requesting history. | Free | benchlm_history_coverage |
| BenchAlign history | Saved BenchLM scores, ranks, evidence labels, ranking surfaces, and methodology versions for an exact model. | Pro · Research | benchlm_rank_history |
| Ranking at a date | Stored BenchAlign ranking rows as of a UTC date. The response identifies incomplete coverage and preserves the saved scoring version. | Pro · Research | benchlm_ranking_snapshot |
| Benchmark history | Cleared measurements for an exact model, benchmark, and protocol. Availability differs by source; this is not every score on the website. | Pro · Research | benchlm_benchmark_history |
| Benchmark changes | First recorded catalog definitions and results, additions, removals, and reintroductions where retained. | Pro · Research | benchlm_benchmark_events |
Current ranking responses contain final BenchAlign outputs. Benchmark result queries contain source scores, including GPQA Diamond, SWE-bench Verified, and FrontierMath evaluated by Epoch AI. BullshitBench v2 includes judge ratings on a 0–2 scale, refusal rates, and published ranks from one pinned snapshot. Source-reported runs retain their original units and dates; artifact publication timestamps do not establish run dates. Arena overall snapshot history contains the publisher’s ratings, ranks, and confidence intervals across retained dates, with scoring methods kept separate. Benchmark measurements and evaluation runs do not establish a leaderboard rank. Missing records stay missing.
Sign in by email. Free starts without a payment method.
Keep it in your application’s or MCP client’s secret settings.
Use REST, or connect an MCP client with Streamable HTTP.
curl 'https://data.benchlm.ai/v1/rankings/current?surface=coding&limit=10' \
-H 'Authorization: Bearer YOUR_DATA_KEY'https://data.benchlm.ai/mcpSend Authorization: Bearer YOUR_DATA_KEY. Supported MCP protocol versions: 2026-07-28, 2025-11-25, and 2025-06-18.
curl 'https://data.benchlm.ai/mcp' \
-H 'Authorization: Bearer YOUR_DATA_KEY' \
-H 'Content-Type: application/json' \
-H 'Accept: application/json, text/event-stream' \
-H 'MCP-Protocol-Version: 2025-11-25' \
--data '{"jsonrpc":"2.0","id":1,"method":"tools/call","params":{"name":"benchlm_list_current_rankings","arguments":{"surface":"coding","limit":10}}}'Check the records you need before you choose a plan. Coverage and usage checks are free, including when your monthly reads are exhausted.
benchlm_get_usageNo. Retained Arena overall snapshots date back to August 2024, and source-reported benchmark runs reach January 2025. Other sources and saved BenchAlign rankings begin later. Research removes the history age cap; it does not fill missing records. Check the dates available for your exact source and model.
Research includes all retained eligible history and 500,000 successful reads per month for $249 USD. Data Pro includes three rolling calendar months and 100,000 reads for $49 USD. Both allow 60 requests per minute. Choose Research when older records or a larger monthly allowance matter to your work.
Yes. A free account gives you current rankings, source results, model and benchmark catalogs, and coverage checks. The coverage tools identify available sources and date windows without using your monthly reads. Check your model and benchmark first, then choose the plan that includes the history you need.
One successful data page uses one read. Initialization, discovery, coverage and usage checks, and unsuccessful queries do not use the monthly allowance. All keys on an account share usage across REST and MCP.
Free renews on your account’s activation anniversary each month. Pro and Research follow your subscription period. Short months use their last day. Your account and usage responses show the exact monthly and request-rate reset times.
The API and MCP identify the monthly read allowance or request-rate limit, with a reset time and retry delay. Free accounts get a Data Pro upgrade link. Eligible Pro accounts can upgrade to Research with a prorated price difference and a fresh 500,000-read allowance for the remaining subscription period. You can receive usage emails at 75%, 90%, and 100%; manage them in your account. Usage and coverage checks stay available.
Pro covers three rolling calendar months by UTC date; Research has no age cap on retained eligible records. Coverage varies by source. Historical pages contain up to 500 rows. Saved-state history uses 90-day windows and cursors. Source-run and publisher snapshot history take an exact board and source model ID, with date filters of up to 90 days. Snapshot history supports cursors and preserves publisher dates; run history preserves reported evaluation times. Check the relevant coverage tool first.
Free, Pro, and Research are access plans for queries and internal analysis. Paid plans add reads and available history; they do not include commercial display or redistribution permission.
Publishing BenchLM-owned outputs or the compiled feed in a customer-facing commercial website, app, or product needs a separate display agreement. Redistributing that material through an API, downloads, or resale needs separately agreed scope. Request commercial-use permission.
Individual quotations and attributed embeds keep their existing permissions. Publisher measurements keep their own commercial-use and redistribution rights. Retain attribution, source notices, and applicable terms. Read the Data access terms.
Research gives you all retained eligible history for $249 USD/month. Check your sources and dates with a free account first.
613 benchmark pages, 882 model profiles, and 11 categories describe the website catalog. API coverage varies by source.
Looking for the website’s existing downloads and their terms? Open the Data hub and downloads.