TabFact
We show this table for reference; we do not rank on it.
TabFact asks whether a statement is supported or refuted by a Wikipedia table. It tests decisions that combine language interpretation with structured evidence.
Original benchmark results
The original paper reports released-test accuracy, including simple, complex, and small-test subsets. GNN-TabFact supplies a later maintainer result on the released split. CodaLab’s challenge uses a separate hidden test and publishes six usernames with rounded scores, without identifying their models. Those submissions remain visible but cannot be assigned model profiles from the available source.
3 source tables · 14 evaluated configurations · reviewed October 1, 2026. Download full results (JSON)
Original released-test results
Human performance is a baseline and does not create a model profile. Released test results are distinct from the CodaLab challenge.
TabFact: A Large-scale Dataset for Table-based Fact Verification — Wenhu Chen and colleagues. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.
| System | Validation accuracy (%) | Test accuracy (%) | Simple test (%) | Complex test (%) | Small test (%) |
|---|---|---|---|---|---|
| BERT classifier w/o Table | 50.9 | 50.5 | 51.0 | 50.1 | 50.4 |
| Table-BERT-Horizontal-F+T-Concatenate | 50.7 | 50.4 | 50.8 | 50.0 | 50.3 |
| Table-BERT-Vertical-F+T-Template | 56.7 | 56.2 | 59.8 | 55.0 | 56.2 |
| Table-BERT-Vertical-T+F-Template | 56.7 | 57.0 | 60.6 | 54.3 | 55.5 |
| Table-BERT-Horizontal-F+T-Template | 66.0 | 65.1 | 79.0 | 58.1 | 67.9 |
| Table-BERT-Horizontal-T+F-Template | 66.1 | 65.1 | 79.1 | 58.2 | 68.1 |
| NSM w/ RL (Binary Reward) | 54.1 | 54.1 | 55.4 | 53.1 | 55.8 |
| NSM w/ LPA-guided ML + RL | 63.2 | 63.5 | 77.4 | 56.1 | 66.9 |
| LPA-Voting w/o Discriminator | 57.7 | 58.2 | 68.5 | 53.2 | 61.5 |
| LPA-Weighted-Voting | 62.5 | 63.1 | 74.6 | 57.3 | 66.8 |
| LPA-Ranking w/ Discriminator | 65.2 | 65.0 | 78.4 | 58.5 | 68.6 |
| LPA-Ranking w/ Discriminator (Caption) | 65.1 | 65.3 | 78.7 | 58.5 | 68.9 |
| Human Performance | - | - | - | - | 92.1 |
GNN-TabFact maintainer comparison
Maintainer README comparison. The Table-BERT label does not identify which serialization variant; we preserve that source label separately.
TabFact: A Large-scale Dataset for Table-based Fact Verification — Wenhu Chen and colleagues. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.
| System | Development accuracy (%) | Test accuracy (%) |
|---|---|---|
| GNN-TabFact | 72.1 | 72.2 |
| Table-BERT | 66.1 | 65.1 |
CodaLab hidden-test submissions — model identity unavailable
All six public challenge submissions from the full results tab and CSV, retrieved October 1, 2026. Source scores are rounded 0–1 accuracies. Usernames and team labels do not identify evaluated models; none of these scores is assigned to a catalog model.
TabFact: A Large-scale Dataset for Table-based Fact Verification — Wenhu Chen and colleagues. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.
| Rank | Submission username | Rounded score (0–1) |
|---|---|---|
| 1 | chriszhao__ | 0.78 |
| 2 | WENGSYX | 0.77 |
| 3 | eisenjulian | 0.74 |
| 4 | hongzhi | 0.63 |
| 5 | chino | 0.61 |
| 6 | wenhu | 0.57 |
Original benchmark results and configurations
The original paper reports released-test accuracy, including simple, complex, and small-test subsets. GNN-TabFact supplies a later maintainer result on the released split. CodaLab’s challenge uses a separate hidden test and publishes six usernames with rounded scores, without identifying their models. Those submissions remain visible but cannot be assigned model profiles from the available source.
Numeric results and configurations from the seven benchmark-owner papers and maintained result tables. This is not a census of all downstream papers or unpublished evaluations.
Snapshot
About TabFact
Year
2019
Tasks
Fact verification against tables
Format
Decision classification
Difficulty
Depends on the evaluated split and label mapping
The original paper reports released-test accuracy, including simple, complex, and small-test subsets. GNN-TabFact supplies a later maintainer result on the released split. CodaLab’s challenge uses a separate hidden test and publishes six usernames with rounded scores, without identifying their models. Those submissions remain visible but cannot be assigned model profiles from the available source.
Freshness and provenance
Version
TabFact 2019
Refresh cadence
Static
Staleness state
Stale
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
Which TabFact results are included?
This page includes 3 numeric result and configuration tables from TabFact’s original paper and available benchmark-owner updates, covering 14 evaluated systems or configurations. Source links, precision, metrics, and evaluation settings remain attached to each table. This coverage does not include every downstream paper or unpublished evaluation.
Can I compare these TabFact scores with the Perplexity panel?
Compare TabFact scores only when the sample, labels, input representation, and evaluation protocol match. Perplexity’s decision panel uses a fixed sample and its own harness. The original tables preserve different splits, metrics, and model configurations, so their numbers cannot establish a direct ranking against that panel.
Where can I download the TabFact results?
The results download on this page provides every imported TabFact table as JSON, with model links, published values, source URLs, and evaluation notes. Each linked configuration profile also exports its own result tables as JSON and numeric metrics as CSV. Missing source measurements remain unreported rather than zero.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.