Gradient Boosting Still Rules Tabular Data. The Foundation Models Are Gaining Ground.
In 2022, a landmark benchmark settled years of argument: on typical tabular data, gradient-boosted trees beat deep learning, and the paper explained exactly why. Four years later, a new species of model trained only on synthetic tables is contesting the verdict. Here is the evidence, the scoreboard, and what it means for the model you train next.

Every few years, someone announces that deep learning has finally conquered the spreadsheet. Every few years, the spreadsheet wins. The 2022 round of this argument was unusually well refereed. Grinsztajn, Oyallon, and Varoquaux built a rigorous 45-dataset benchmark at NeurIPS 2022, gave every method an equal tuning budget, and delivered a verdict with reasons attached. Tree-based models were not just winning. The paper showed they were winning for structural causes, the kind that do not disappear with a better architecture or a longer training run.
That verdict has organized tabular machine learning ever since. And it is now, for the first time, genuinely under pressure. Not from bigger MLPs or cleverer transformers, but from a different species of model entirely: foundation models pre-trained on millions of synthetic tables that never train on your data at all. In September 2026 alone, three of them posted leaderboard scores above anything gradient boosting has managed on the same benchmark. This is a deep dive into what the 2022 verdict established, what has actually changed since, and how to choose the right model for your next tabular problem.
1. The 2022 verdict, in plain terms
The paper is "Why do tree-based models still outperform deep learning on typical tabular data?" (Grinsztajn, Oyallon, Varoquaux, NeurIPS 2022, Datasets and Benchmarks track, arXiv:2207.08815). Its headline result, quoted from the abstract: tree-based models "remain state of the art on medium-sized data (~10K samples) even without accounting for their superior speed." Random forests, gradient boosting, and XGBoost beat MLPs, ResNets, FT-Transformer, and SAINT across the benchmark. Crucially, the gap was not a tuning artifact. Giving the neural nets an exhaustive hyperparameter search did not close it.
The part that made the paper endure was the explanation. The authors proposed three structural reasons trees win, each backed by an experiment:
Trees learn jagged functions; neural nets are biased toward smooth ones. Neural networks learn low-frequency, smooth functions first, a phenomenon known as spectral bias (Rahaman et al., 2019). Real tabular targets are rarely smooth: churn flips at a contract boundary, default risk jumps at a credit-score threshold. The authors tested this directly by Gaussian-smoothing the training targets. Smoothing barely changed neural network performance but markedly degraded the trees, which tells you the trees were winning precisely by capturing the irregular, high-frequency patterns the smoothing erased.
Trees ignore junk features; MLPs drown in them. Real tables are full of uninformative columns. Tree ensembles are robust to them because every split is a feature-selection decision. MLP-like architectures degraded significantly as noise features increased. Notably, FT-Transformer held up better here, a hint that attention could select features too, but not well enough to win overall.
Rotation invariance is a bug, not a feature, for tables. MLPs and ResNets treat all linear combinations of features equally. But tabular columns have individual meaning: a "monthly charges" column is not interchangeable with a linear mix of "monthly charges" and "tenure." Decision trees split on one feature at a time, which is a better match for how tables actually work. The authors proved the point by randomly rotating the features, destroying the per-column structure. The performance order flipped: neural nets won on rotated data, trees won on real data.
A year earlier, Shwartz-Ziv and Armon ("Tabular Data: Deep Learning is Not All You Need," arXiv:2106.03253) had reached a compatible conclusion from the other direction: XGBoost outperformed the proposed deep tabular models across datasets, including the datasets the deep-learning papers themselves were evaluated on, and needed far less tuning. Their paper contained a second finding that matters more in 2026 than it did in 2021: an ensemble of XGBoost with a neural model beat XGBoost alone. That is the first crack in the wall, and everything since has widened it.
2. The plot twist: a foundation model that never trains on your data
TabPFN (Hollmann et al., published in Nature in January 2025) took a completely different route around the 2022 verdict. Instead of training a network on your table, the authors pre-trained a transformer on millions of synthetic tabular datasets generated from structural causal models. The result is a frozen network that treats your labeled training rows as context and labels your test rows in a single forward pass. No gradient descent on your data, no hyperparameter search, no cross-validation loop. You hand it the table and it reads the patterns out of the rows on the fly.
Why this matters is speed of a different kind. A tuned XGBoost pipeline costs you a search over depth, learning rate, regularization, and sampling. TabPFN costs you one forward pass. And on small tables it is genuinely competitive. An independent check by the consultancy INWT on a vehicle-pricing dataset found TabPFN beating XGBoost clearly at small training sizes, with the two converging to similar accuracy around 6,000 training rows. That "convergence point" framing is the honest way to read the result: the foundation model wins where data is scarce and tuning budgets are thin, which describes a large share of real analytics work.
3. The 2025 scoreboard: TabArena makes the contest legible
The field needed a referee for this new contest, and TabArena (2025) became one. It is a living benchmark: 51 manually curated real-world datasets (30 binary classification, 8 multiclass, 13 regression), evaluated with Elo ratings anchored on Random Forest at 1000, built from roughly 25 million individual model runs that took about 15 years of wall-clock time. Its headline conclusions, stated in the paper: with tuning and ensembling, the best deep learning methods (TabM, RealMLP) perform similar to or better than gradient-boosted trees; tabular foundation models dominate for small data even without tuning; and ensemble pipelines across model families are the state of the art.
Read that carefully, because the 2022 verdict is visible inside it. "With tuning and ensembling" is doing heavy work: the neural methods needed both to draw level with boosting. And the benchmark's scope is explicit about what it does not cover: it focuses on small-to-medium, independent-and-identically-distributed data, and explicitly leaves out temporal and grouped structure, very small datasets (under 500 rows), large datasets (over 250,000 rows), and distribution shift. The crown being contested is a specific crown: IID tables in the small-to-medium regime.
4. The September 2026 sprint: three new challengers
September 2026 compressed a year's worth of leaderboard movement into weeks. Three tabular foundation models posted scores that, if taken at face value, put them clearly ahead of tuned gradient boosting on TabArena:
TabPFN-3.5 (Prior Labs, mid-September 2026): a wider 220M-parameter in-context transformer, one multitask checkpoint covering classification and regression, handling up to 1M rows. Prior Labs reports its "Thinking" variant at 1910 Elo on TabArena, the base model at 1866, both ahead of TabFM+ at 1823, and claims the base model beats AutoGluon 1.6's extreme configuration by 130 Elo points in a fifth of the time. It also claims first place on BeyondArena, TALENT, and several other suites. The fine print from the same reports: on BeyondArena's grouped, temporal, and large-data slices, tuned and ensembled MLPs still lead.
LimiX-2 (Stable AI, September 2026): a 400M-parameter model at 1935 Elo on TabArena, a reported 117.4 points ahead of TabFM+, with pairwise win rates around 65-67% against TabFM and EXAONE Tabular. It also tops TALENT (1506) and BCCO (1432), the latter ahead of AutoGluon 1.6 at 1376.
Kumo Tabular (NVIDIA, late September 2026): an open foundation model claiming 1950 Elo on TabArena, first overall, while running 26 times faster than LimiX-2 on the same GPU. NVIDIA pre-trained three sizes on 35M, 71M, and 137M synthetic tables respectively. The honest caveats come from NVIDIA itself: the model handles numerical and categorical columns only (text, images, and timestamps need preprocessing), and accuracy may degrade when new data drifts from the context rows.
5. Where trees still win: the honest ledger
The foundation-model surge is real, but it is concentrated in the regime the 2022 paper studied and the regime TabArena measures. Outside that regime, the tree's advantages from 2022 still compound:
Large data. TabArena explicitly excludes tables above 250,000 rows. Gradient boosting trains on millions of rows on a CPU, and it has done so in production for a decade. Foundation models with in-context learning over thousands of rows are a GPU workload with a very different cost curve.
Latency and cost at serving time. A LightGBM model scores a row in milliseconds on commodity CPU. A transformer doing a forward pass conditioned on thousands of context rows needs a GPU and a budget conversation. For high-throughput scoring, this is not close.
Temporal and grouped structure. BeyondArena's own results show tuned MLPs still leading on grouped, temporal, and large-data slices. Time series, panel data, and anything with leakage-shaped structure remain tree and linear-model territory until proven otherwise.
Distribution shift. NVIDIA's own documentation warns that accuracy may degrade when query data drifts from context rows. Trees are not immune to shift, but a decade of production practice has built the monitoring and retraining playbooks for them.
Interpretability and compliance. In regulated settings, "here are the SHAP values from a tree ensemble" is an accepted language with auditors. "The frozen transformer read the context rows" is not, yet.
6. What this means for data and AI practitioners
The decision tree for choosing a model (the other kind of decision tree) is simpler than the discourse suggests:
- Start with LightGBM or XGBoost. It remains the fastest path to a strong baseline on any table, and the 2022 paper is why: the inductive biases of trees match the structure of tabular data. If your boosting baseline with honest validation already beats the business target, ship it.
- On small data with no tuning budget, try a foundation model. TabPFN-style models are now a legitimate first attempt below a few thousand rows, and the "no training" workflow is genuinely fast to evaluate. Compare it against your boosting baseline on the same split.
- Report the boosting baseline with an equal tuning budget. The enduring methodological lesson of 2022: most "we beat XGBoost" claims in the literature used weak baselines. Give the tree the same search budget you give the neural net, then compare.
- Ensemble across families when it counts. The 2021 finding that XGBoost plus a neural model beats either alone has survived every subsequent benchmark, including TabArena's. If a few extra points of accuracy justify the complexity, ensemble a booster with a foundation model rather than picking a winner.
- Do not re-platform on vendor Elo scores. Reproduce the claimed numbers on your own data first. The 2026 leaderboard is a press-release leaderboard until independent runs confirm it.
Against the foundation-model surge
The 2022 verdict was built on structural arguments, not on a temporary engineering lead. Smoothness bias, junk-feature robustness, and per-column meaning are properties of tables, not of 2022-era hardware. Foundation models route around these problems rather than solving them: in-context learning over synthetic priors is a different way of doing feature selection and function approximation, and it is strongest exactly where data is too scarce for the tree's advantages to compound. On the medium-to-large, IID-to-temporal, CPU-served problems that make up most production tabular ML, nothing in the 2026 numbers overturns Grinsztajn et al.
In favour of the foundation-model surge
The 2022 verdict's scope is also its vulnerability. "Medium-sized data (~10K samples)" is precisely the regime where TabPFN-style models now dominate, and "equal tuning budgets" cuts the other way when the challenger needs no tuning at all. The synthetic-pretraining bet has scaled further than skeptics expected: from TabPFN's small-data niche in early 2025 to 400M-parameter models and 137M synthetic tables by late 2026, with the accuracy-efficiency frontier moving every quarter. If your work lives in the small-data regime, the tree's structural advantages are real but no longer decisive.
What would make me wrong
Three falsifiable markers. First, if an independent, large-scale benchmark on non-IID, temporal tables above a million rows shows foundation models beating tuned gradient boosting, the "regime" defense collapses. Second, if GPU inference economics drop below CPU boosting at serving scale, the cost argument for trees evaporates. Third, if the vendor-reported Elo scores reproduce on TabArena's living leaderboard under independent runs, the 2026 ordering becomes fact rather than marketing. I will revisit this piece when any of the three happens.
Sources
- Grinsztajn, Oyallon, Varoquaux (2022). Why do tree-based models still outperform deep learning on typical tabular data? NeurIPS 2022, Datasets and Benchmarks track.
- Shwartz-Ziv, Armon (2021). Tabular Data: Deep Learning is Not All You Need.
- Hollmann et al. (2025). TabPFN, published in Nature, January 2025. Accessible explainer with benchmarks: Tabular Foundation Models: How TabPFN Predicts Without Training.
- TabArena (2025). TabArena: A Living Benchmark for Machine Learning on Tabular Data.
- Prior Labs (Sep 2026). TabPFN-3.5 technical report, summarized in Prior Labs Releases TabPFN-3.5.
- Stable AI (Sep 2026). LimiX-2 technical report, summarized in LimiX-2 Tops TabArena, Overtaking TabPFN-3.5.
- NVIDIA (Sep 2026). Kumo Tabular release, summarized in NVIDIA Releases Open Kumo Tabular Model.
Related on Everyday Data Science: I benchmarked pandas, Polars, and DuckDB on 3.4M taxi rows. DuckDB swept every query.
When did you last beat a tuned XGBoost on a real table, and what actually did it?
About the writer
Data Scientist & AI Researcher
Data scientist and AI researcher at Pace University. I coined Artificial Frictional Unemployment, and built the first machine learning model for crop yield prediction in Sierra Leone. Author of Understanding Agentic AI. I write about agentic systems and applied ML, with a bias toward what actually works, and who gets left out when it doesn't.
Found this useful? Passing it on to someone who builds is the best way to help the publication grow.
Built something worth sharing? Write it up for us →