# TabFact

> TabFact asks whether a statement is supported or refuted by a Wikipedia table. It tests decisions that combine language interpretation with structured evidence.

Canonical page: https://benchlm.ai/benchmarks/tabfact

- Category: [Decision Models](/decision-models)
- Last updated: October 1, 2026 source review

## About TabFact

- Year: 2019
- Tasks: Fact verification against tables
- Format: Decision classification
- Difficulty: Depends on the evaluated split and label mapping
- Paper: [TabFact: A Large-scale Dataset for Table-based Fact Verification](https://tabfact.github.io/)

The original paper reports released-test accuracy, including simple, complex, and small-test subsets. GNN-TabFact supplies a later maintainer result on the released split. CodaLab’s challenge uses a separate hidden test and publishes six usernames with rounded scores, without identifying their models. Those submissions remain visible but cannot be assigned model profiles from the available source.

TabFact is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Original benchmark results

The original paper reports released-test accuracy, including simple, complex, and small-test subsets. GNN-TabFact supplies a later maintainer result on the released split. CodaLab’s challenge uses a separate hidden test and publishes six usernames with rounded scores, without identifying their models. Those submissions remain visible but cannot be assigned model profiles from the available source.

[Download full results (JSON)](/api/data/decision-benchmarks?benchmark=tabFact)

### Original released-test results

Human performance is a baseline and does not create a model profile. Released test results are distinct from the CodaLab challenge.

TabFact: A Large-scale Dataset for Table-based Fact Verification — Wenhu Chen and colleagues. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://arxiv.org/abs/1909.02164)

| System | Validation accuracy (%) | Test accuracy (%) | Simple test (%) | Complex test (%) | Small test (%) |
| --- | --- | --- | --- | --- | --- |
| [BERT classifier w/o Table](/models/native-tabfact-bert-classifier-w-o-table) | 50.9 | 50.5 | 51.0 | 50.1 | 50.4 |
| [Table-BERT-Horizontal-F+T-Concatenate](/models/native-tabfact-table-bert-horizontal-f-t-concatenate) | 50.7 | 50.4 | 50.8 | 50.0 | 50.3 |
| [Table-BERT-Vertical-F+T-Template](/models/native-tabfact-table-bert-vertical-f-t-template) | 56.7 | 56.2 | 59.8 | 55.0 | 56.2 |
| [Table-BERT-Vertical-T+F-Template](/models/native-tabfact-table-bert-vertical-t-f-template) | 56.7 | 57.0 | 60.6 | 54.3 | 55.5 |
| [Table-BERT-Horizontal-F+T-Template](/models/native-tabfact-table-bert-horizontal-f-t-template) | 66.0 | 65.1 | 79.0 | 58.1 | 67.9 |
| [Table-BERT-Horizontal-T+F-Template](/models/native-tabfact-table-bert-horizontal-t-f-template) | 66.1 | 65.1 | 79.1 | 58.2 | 68.1 |
| [NSM w/ RL (Binary Reward)](/models/native-tabfact-nsm-w-rl-binary-reward) | 54.1 | 54.1 | 55.4 | 53.1 | 55.8 |
| [NSM w/ LPA-guided ML + RL](/models/native-tabfact-nsm-w-lpa-guided-ml-rl) | 63.2 | 63.5 | 77.4 | 56.1 | 66.9 |
| [LPA-Voting w/o Discriminator](/models/native-tabfact-lpa-voting-w-o-discriminator) | 57.7 | 58.2 | 68.5 | 53.2 | 61.5 |
| [LPA-Weighted-Voting](/models/native-tabfact-lpa-weighted-voting) | 62.5 | 63.1 | 74.6 | 57.3 | 66.8 |
| [LPA-Ranking w/ Discriminator](/models/native-tabfact-lpa-ranking-w-discriminator) | 65.2 | 65.0 | 78.4 | 58.5 | 68.6 |
| [LPA-Ranking w/ Discriminator (Caption)](/models/native-tabfact-lpa-ranking-w-discriminator-caption) | 65.1 | 65.3 | 78.7 | 58.5 | 68.9 |
| Human Performance | - | - | - | - | 92.1 |

### GNN-TabFact maintainer comparison

Maintainer README comparison. The Table-BERT label does not identify which serialization variant; we preserve that source label separately.

TabFact: A Large-scale Dataset for Table-based Fact Verification — Wenhu Chen and colleagues. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://github.com/wenhuchen/GNN-TabFact)

| System | Development accuracy (%) | Test accuracy (%) |
| --- | --- | --- |
| [GNN-TabFact](/models/native-tabfact-gnn-tabfact-maintainer) | 72.1 | 72.2 |
| [Table-BERT](/models/native-tabfact-table-bert-maintainer) | 66.1 | 65.1 |

### CodaLab hidden-test submissions — model identity unavailable

All six public challenge submissions from the full results tab and CSV, retrieved October 1, 2026. Source scores are rounded 0–1 accuracies. Usernames and team labels do not identify evaluated models; none of these scores is assigned to a catalog model.

TabFact: A Large-scale Dataset for Table-based Fact Verification — Wenhu Chen and colleagues. Numeric results transcribed and reformatted; evaluation notes and model links added by BenchLM. Published numeric precision is retained.

[Published source](https://competitions.codalab.org/competitions/21611/results/36309)

| Rank | Submission username | Rounded score (0–1) |
| --- | --- | --- |
| 1 | chriszhao__ | 0.78 |
| 2 | WENGSYX | 0.77 |
| 3 | eisenjulian | 0.74 |
| 4 | hongzhi | 0.63 |
| 5 | chino | 0.61 |
| 6 | wenhu | 0.57 |

## FAQ

### Which TabFact results are included?

This page includes 3 numeric result and configuration tables from TabFact’s original paper and available benchmark-owner updates, covering 14 evaluated systems or configurations. Source links, precision, metrics, and evaluation settings remain attached to each table. This coverage does not include every downstream paper or unpublished evaluation.

### Can I compare these TabFact scores with the Perplexity panel?

Compare TabFact scores only when the sample, labels, input representation, and evaluation protocol match. Perplexity’s decision panel uses a fixed sample and its own harness. The original tables preserve different splits, metrics, and model configurations, so their numbers cannot establish a direct ranking against that panel.

### Where can I download the TabFact results?

The results download on this page provides every imported TabFact table as JSON, with model links, published values, source URLs, and evaluation notes. Each linked configuration profile also exports its own result tables as JSON and numeric metrics as CSV. Missing source measurements remain unreported rather than zero.
