Skip to main content
BenchLM

Is that claim about an AI model true?

Workspace

Paste the claim. Get supported, contradicted, or not established, with the public rows the verdict rests on.

Evidence updated 379 models · 357 benchmarks
0 / 2000 · up to 5 sentences, one verdict each · not stored · a decision model reads each sentence, deterministic code checks it

Or set the claim by hand

Re-checking a claim you set by hand is free and does not use the reader.

How the claim checker decides

A claim is a small form: the claim shape, the model it is about, the model it is compared with, and the category or benchmark. A decision model fills that form from your sentence and nothing else. Deterministic code then checks the form against the same public data the model pages, benchmark pages, and pricing table show, on their stated basis. A row that rests on Estimated evidence, an unranked model, or a practical tie is reported as not established rather than rounded into a winner.

Looking for the right model rather than checking a claim? Use the LLM Selector or browse the best model for your job pages.

Questions

Can I paste several claims at once?

Yes. Up to five sentences are read in parallel, each as its own decision with its own verdict. Select a sentence to see its evidence or adjust its reading. Each sentence counts as one request against the claim checker’s own free allowance.

What can the claim checker verify?

Eight claim shapes: a model leads the overall leaderboard, leads a category leaderboard, scores higher than another model in a category, scores higher than another model on a named benchmark, is cheaper per token, has a larger documented context window, streams faster, or is an open-weight model. Anything else is reported as not checkable.

What do supported, contradicted, and not established mean?

Supported means the public data says what the claim says. Contradicted means it states the opposite. Not established means the data cannot settle it: a row is missing, at least one score rests on Estimated evidence or is unranked, the values are a practical tie, or the sentence made no checkable claim. A directional reading never names a winner.

Does an AI judge the claim?

No. A decision model only reads the sentence into a fixed form of multiple-choice answers: claim type, models, category, benchmark. The verdict is computed by deterministic code against the same public dataset the model, category, and benchmark pages use. You can correct the reading by hand, and that re-check costs nothing.

Can I call this from a script?

Yes. GET /api/tools/claim-check with type, subject, comparator, category, and benchmark returns the same verdict without the reader; add catalog=1 to list the model slugs, benchmark keys, and claim types.