Skip to main content
BenchLM

TabFact

We show this table for reference; we do not rank on it.

TabFact asks whether a statement is supported or refuted by a Wikipedia table. It tests decisions that combine language interpretation with structured evidence.

Original benchmark results

The original paper reports released-test accuracy, including simple, complex, and small-test subsets. GNN-TabFact supplies a later maintainer result on the released split. CodaLab’s challenge uses a separate hidden test and publishes six usernames with rounded scores, without identifying their models. Those submissions remain visible but cannot be assigned model profiles from the available source.

3 source tables · 14 evaluated configurations · reviewed October 1, 2026. Download full results (JSON)

Original released-test results

Human performance is a baseline and does not create a model profile. Released test results are distinct from the CodaLab challenge.

Published source

TabFact: A Large-scale Dataset for Table-based Fact Verification — Wenhu Chen and colleagues. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

Original released-test results. Published source metrics with original precision; unreported cells are not zero.
SystemValidation accuracy (%)Test accuracy (%)Simple test (%)Complex test (%)Small test (%)
BERT classifier w/o Table50.950.551.050.150.4
Table-BERT-Horizontal-F+T-Concatenate50.750.450.850.050.3
Table-BERT-Vertical-F+T-Template56.756.259.855.056.2
Table-BERT-Vertical-T+F-Template56.757.060.654.355.5
Table-BERT-Horizontal-F+T-Template66.065.179.058.167.9
Table-BERT-Horizontal-T+F-Template66.165.179.158.268.1
NSM w/ RL (Binary Reward)54.154.155.453.155.8
NSM w/ LPA-guided ML + RL63.263.577.456.166.9
LPA-Voting w/o Discriminator57.758.268.553.261.5
LPA-Weighted-Voting62.563.174.657.366.8
LPA-Ranking w/ Discriminator65.265.078.458.568.6
LPA-Ranking w/ Discriminator (Caption)65.165.378.758.568.9
Human Performance----92.1

GNN-TabFact maintainer comparison

Maintainer README comparison. The Table-BERT label does not identify which serialization variant; we preserve that source label separately.

Published source

TabFact: A Large-scale Dataset for Table-based Fact Verification — Wenhu Chen and colleagues. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

GNN-TabFact maintainer comparison. Published source metrics with original precision; unreported cells are not zero.
SystemDevelopment accuracy (%)Test accuracy (%)
GNN-TabFact72.172.2
Table-BERT66.165.1

CodaLab hidden-test submissions — model identity unavailable

All six public challenge submissions from the full results tab and CSV, retrieved October 1, 2026. Source scores are rounded 0–1 accuracies. Usernames and team labels do not identify evaluated models; none of these scores is assigned to a catalog model.

Published source

TabFact: A Large-scale Dataset for Table-based Fact Verification — Wenhu Chen and colleagues. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

CodaLab hidden-test submissions — model identity unavailable. Published source metrics with original precision; unreported cells are not zero.
RankSubmission usernameRounded score (0–1)
1chriszhao__0.78
2WENGSYX0.77
3eisenjulian0.74
4hongzhi0.63
5chino0.61
6wenhu0.57

Original benchmark results and configurations

The original paper reports released-test accuracy, including simple, complex, and small-test subsets. GNN-TabFact supplies a later maintainer result on the released split. CodaLab’s challenge uses a separate hidden test and publishes six usernames with rounded scores, without identifying their models. Those submissions remain visible but cannot be assigned model profiles from the available source.

Numeric results and configurations from the seven benchmark-owner papers and maintained result tables. This is not a census of all downstream papers or unpublished evaluations.

Snapshot

3 source tables14 configurationsDisplay only

About TabFact

Year

2019

Tasks

Fact verification against tables

Format

Decision classification

Difficulty

Depends on the evaluated split and label mapping

The original paper reports released-test accuracy, including simple, complex, and small-test subsets. GNN-TabFact supplies a later maintainer result on the released split. CodaLab’s challenge uses a separate hidden test and publishes six usernames with rounded scores, without identifying their models. Those submissions remain visible but cannot be assigned model profiles from the available source.

Freshness and provenance

Version

TabFact 2019

Refresh cadence

Static

Staleness state

Stale

Question availability

Public benchmark set

StaleDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

Which TabFact results are included?

This page includes 3 numeric result and configuration tables from TabFact’s original paper and available benchmark-owner updates, covering 14 evaluated systems or configurations. Source links, precision, metrics, and evaluation settings remain attached to each table. This coverage does not include every downstream paper or unpublished evaluation.

Can I compare these TabFact scores with the Perplexity panel?

Compare TabFact scores only when the sample, labels, input representation, and evaluation protocol match. Perplexity’s decision panel uses a fixed sample and its own harness. The original tables preserve different splits, metrics, and model configurations, so their numbers cannot establish a direct ranking against that panel.

Where can I download the TabFact results?

The results download on this page provides every imported TabFact table as JSON, with model links, published values, source URLs, and evaluation notes. Each linked configuration profile also exports its own result tables as JSON and numeric metrics as CSV. Missing source measurements remain unreported rather than zero.

Last updated: October 1, 2026 source review · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.