Skip to main content

What Are Tabular Foundation Models? How TabPFN, TabICL, and TabFM Predict Without Training

Learn what tabular foundation models are, how in-context learning works, and when TabPFN or TabICL beat XGBoost on structured data, with a decision framework.
Oct 6, 2026  · 15 min read

Explore with AI

ChatGPTClaudePerplexity

Imagine two students sitting the same exam. The first student spends three weeks studying only one module, while the second has already practiced on millions of different exams and simply gets a sheet of worked examples on the day.

For most of machine learning’s past, tabular models have been the first student. Every new dataset meant cleaning columns, engineering features, training XGBoost or LightGBM, and tuning hyperparameters for hours.

Tabular foundation models, though, are the second student: you hand them your labeled rows as examples, and they predict the rest with no training at all.

In 2026, this idea went from research papers to real products, with SAP committing more than €1 billion to Prior Labs and Google releasing its own model. So in this article, we will look at:

  • What tabular foundation models are
  • Why tabular data was so hard for deep learning
  • How these models actually make predictions
  • The key models in 2026
  • How to try one in Python
  • When you should (and should not) use one

Don’t worry if you don’t have much deep learning experience. The code in this article uses the familiar scikit-learn interface, and if you want a quick refresher first, 8 Machine Learning Models Explained in 20 Minutes covers the basics.

Tabular Foundation Models TL;DR

  • Tabular foundation models such as TabPFN, TabICLv2, and Google’s TabFM are neural networks pretrained on millions of synthetic tables, so they can predict on your data without any training or tuning.
  • On small-to-medium datasets (roughly 300 to 100,000 rows) with random splits, they now beat tuned XGBoost, CatBoost, and LightGBM on most benchmarks.
  • Gradient-boosted trees still win on time-ordered, grouped, and very large datasets, and when you need fast CPU predictions or easy explanations.
  • TabICLv2 is the best place to start: it is open source, commercially usable, and installs with pip install tabicl.

What Are Tabular Foundation Models?

A tabular foundation model is a neural network pretrained on millions of synthetic tabular datasets that learns a general strategy for predicting a target column, then applies that strategy to a new table through in-context learning, without any per-dataset training.

The big change here is when the learning happens.

With XGBoost, the learning happens on your data every time you train a new model. With a tabular foundation model, the learning happened months ago during pretraining, and your data is simply the input. There are also a few terms you will keep seeing, so let me explain them in plain language:

  • In-context learning (ICL): the model looks at your labeled rows as examples and uses them to predict new rows, without changing its own weights.
  • Prior-fitted network (PFN): a model trained on data sampled from a prior, which is just a recipe for generating lots of different fake datasets.
  • Synthetic pretraining: training on made-up tables instead of real ones.
  • Zero-shot prediction: predicting on a brand-new dataset with no training or fine-tuning at all.

Now, let’s talk about how tabular foundation models compare to boosting methods and automated machine learning (AutoML):

  Tabular foundation models Gradient-boosted trees (XGBoost, LightGBM) AutoML (AutoGluon, H2O AutoML)
Training time on new data None Seconds to minutes, plus tuning Minutes to hours
Interpretability Limited (SHAP is possible but slow) Good (feature importance, fast SHAP) Varies, often hard-to-read ensembles
Best dataset size Hundreds to about 100K rows Thousands to many millions of rows Thousands to millions of rows
GPU needed? Recommended past a few thousand rows No Usually no
Small, randomly split datasets Best on current benchmarks Strong when tuned Strong but slow

Before we look at how they work, though, I want to quickly compare them with the most famous, widely used models: large language models.

Tabular foundation models vs. large language models

Large language models (LLMs) and tabular foundation models both use transformers, and both learn from examples in their input.

The difference is what they are built to read. An LLM reads a table as a long string of text, which means it has to work out from commas and spaces which numbers belong to the same column.

There is also a meaning problem.

The value 42 could be someone’s age or a revenue figure, but nothing in the number itself tells you which. A tabular foundation model is trained to treat each column as its own variable and each row as one example, so it focuses on how the columns relate to each other.

I like to think of it this way: an LLM is great at talking about your data, and a tabular foundation model is built to do the math on it.

The two can also work together, which is exactly how H2O.ai positions its tabH2O model: as the prediction tool an AI agent calls when it needs a number from a spreadsheet.

Why Was Tabular Data So Hard for Deep Learning?

Tabular data was hard for deep learning because tables have no shared structure that a neural network can learn once and reuse. A pixel is a pixel in every photo. A column called score in a financial dataset and a column called score in a football dataset have nothing in common.

A well-known 2022 paper by Léo Grinsztajn, Edouard Oyallon, and Gaël Varoquaux, Why do tree-based models still outperform deep learning on tabular data?, tested this on 45 datasets and found that tree-based models, especially gradient boosting, still beat deep learning on medium-sized tables of around 10,000 rows. The reasons they found are:

  • Sharp jumps: real patterns often have sudden steps (think tax brackets). Trees handle these easily, while neural networks are biased toward smooth functions.
  • Useless columns: tables often contain columns that do not matter. Trees simply ignore them, but they hurt neural networks noticeably.
  • Columns have individual meanings: each column is its own variable, and standard neural networks tend to blend columns together in ways that lose that meaning.

Two more practical problems make this worse. A single table can mix numbers, categories, rankings, and missing values, and most business tables only have hundreds or thousands of rows, which is far too little for a neural network trained from scratch.

I would argue that small datasets are the most important problem. A neural network trained from zero on 800 rows has nothing to fall back on, whereas a tree does not need any prior knowledge to work.

This is exactly what pretraining fixed. A tabular foundation model has already practiced on millions of small, messy, made-up tables before it ever sees yours, so it is not starting from zero.

How Do Tabular Foundation Models Work?

A tabular foundation model takes your labeled training rows and your unlabeled test rows together as one input and predicts the missing labels in a single forward pass. Its weights never change, just like an LLM answering a question after you give it a few examples in the prompt.

This also changes what the familiar scikit-learn methods mean.

When you call .fit() on TabICL, the TabICL documentation explains that it loads the pretrained checkpoint, preprocesses your data, and stores it. The real work happens during .predict().

The .fit() name stays so that the model slots into your existing scikit-learn code.

TabPFN overview: training on synthetic datasets with a loss (left) and predicting on a real dataset in a single forward pass (right), plus the 2D attention architecture

Figure 1: TabPFN is pretrained on millions of synthetic datasets, then predicts on a real table in one forward pass. The bottom panel shows how it attends across features and samples. Source: Hollmann et al., “Accurate predictions on small data with a tabular foundation model”, Nature (2025).

Synthetic pretraining

Synthetic pretraining means training the model on made-up tables instead of real ones, produced by a data generator that works like a generative model.

Think of a pilot training in a flight simulator. The pilot practices thousands of fake flights with different weather, airports, and faults, so that the first real flight does not feel new.

TabPFN works in a similar way.

Its creators wrote a generator that invents random “cause and effect” rules between columns (for example, column A affects column B, which affects the target), then produces rows from those rules.

During pretraining, the target column is hidden, the model tries to predict it, and this repeats millions of times with different rules, noise levels, and table sizes.

The important thing is that the model never learns facts about any real topic. It learns general skills such as spotting which columns matter, dealing with outliers, and knowing when to be unsure.

In fact, the TabICLv2 paper credits its new synthetic data generator as a key reason for its better performance.

Two ways to read a table

You do not need to understand every detail of the architectures, but it does help to know that there are two main designs:

  1. Look at every cell (TabPFN). TabPFN-2, published in Nature in 2025, keeps track of every single cell and compares cells across both rows and columns. This is very detailed but gets expensive as tables grow, which is why TabPFN-2 was limited to 10,000 rows and 500 features.
  2. Summarize each row first (TabICL). TabICL first squashes each row into one compact summary, then learns from those summaries instead of individual cells. Because it has far fewer things to compare, it can handle much bigger tables.

Note that both families are trained on synthetic data, so “prior-fitted network” describes how they are trained rather than a separate type of model.

TabFM architecture: a small table with training rows and one test row passes through alternating row and column attention, then row compression, then in-context learning to predict the missing label

Figure 2: TabFM mixes both designs: TabPFN-style attention over rows and columns, then TabICL-style row compression, then in-context learning for the missing label. Source: Google Research, “Introducing TabFM: A zero-shot foundation model for tabular data” (2026).

The catch: prediction gets slower instead

There is a trade-off here that catches a lot of people out.

XGBoost spends time training once, then predicts almost instantly.

A tabular foundation model skips training, but it has to read all of your training rows every time it predicts.

Think of a chef who does no prep work before service but has to reread the whole recipe book for every order.

For a small book, that is fine, but as your dataset grows, predictions get slower. Both TabICL and TabPFN can cache their “reading” of the training data to speed up repeated predictions, but the first pass still costs time.

What Are the Key Tabular Foundation Models in 2026?

The key tabular foundation models in 2026 are TabPFN, TabICLv2, Google’s TabFM, Fundamental’s NEXUS, and H2O.ai’s tabH2O. What I find interesting is how fast the business side moved: five months took this area from academic papers to major deals. Here is a brief timeline:

Now let’s walk through each one, starting with the model that started it all.

TabPFN (Prior Labs, now part of SAP)

TabPFN is the model that created this category. TabPFN-2 was published in Nature in January 2025 and beat tuned tree-based models on datasets of up to 10,000 rows.

Since then, Prior Labs has released new versions quickly. TabPFN-3 arrived in May 2026 and handles up to 1 million rows (and up to 200 features). It also added a Thinking mode (through its paid API) that spends extra time at the fitting stage to make better predictions.

The latest version, TabPFN-3.5, came out in September 2026, and its technical report claims first place on several major benchmarks, including on time-ordered and grouped data. I would treat those as the company’s own numbers until others confirm them.

One more thing is the license. TabPFN-2 can be used commercially as long as you credit Prior Labs, but versions 2.5 to 3.5 are non-commercial. That means you can experiment with them for free, but using them in a real product requires a paid license.

TabICLv2 (Inria)

TabICLv2 is the best fully open tabular foundation model right now. It was built at Inria, released in February 2026, and accepted at ICML 2026.

The headline result from the TabICLv2 paper is that, without any tuning, it beat RealTabPFN-2.5, the previous best model, even though that model had been tuned, ensembled, and fine-tuned on real data.

According to its GitHub repository, it also beats heavily tuned XGBoost, CatBoost, and LightGBM on around 80% of datasets in the TabArena benchmark.

It is also fast, handling 50,000 rows with 100 features in under 10 seconds on an H100 GPU, which is around 10 times faster than TabPFN-2.5.

It works best between 300 and 100,000 rows and can stretch to around 500,000 rows with some loss in accuracy.

In my opinion, this is the one most readers should start with.

You install it with pip install tabicl, it works like any scikit-learn model, and its permissive license lets you use it commercially without signing up for anything. The leaderboard changes fast, though, and TabPFN-3.5’s report now ranks itself ahead of TabICLv2.

Google TabFM

TabFM is Google Research’s tabular foundation model, released on June 30, 2026. What makes it different is where you can use it: it is built into BigQuery, Google’s cloud data warehouse, so you can get predictions with plain SQL.

-- Use past customers (with known churn) to predict churn for new customers
-- AI.PREDICT is in preview, so check the BigQuery docs for the latest syntax
SELECT *
FROM AI.PREDICT(
  TABLE my_dataset.customers_history,
  TABLE my_dataset.customers_new,
  label_col => 'churned'
);

This is great for people who know SQL but not Python.

However, the BigQuery documentation currently limits you to 20 feature columns and 10 classes, and Google plans to add token-based pricing on top of standard BigQuery charges from October 30, 2026. The model weights on Hugging Face are non-commercial, so commercial use goes through BigQuery.

Fundamental NEXUS

NEXUS is a closed, business-only model from Fundamental, a San Francisco startup founded by former DeepMind researchers. It launched in February 2026 with $255M in funding, and the company calls it a Large Tabular Model (LTM).

Businesses buy and run NEXUS through AWS, and Fundamental says Fortune 100 companies already use it for demand forecasting, pricing, and customer churn. However, there are no public weights or benchmark results you can check yourself, so I would treat it as an enterprise option.

H2O.ai tabH2O

tabH2O is H2O.ai’s model, and its whole idea is simple: send your data, get predictions back.

It cleans up missing values and categories for you and can run on a company’s own servers, which matters for banks and hospitals that cannot send data to the cloud.

What I like is that the tabH2O paper does not oversell it. It beat tuned CatBoost and LightGBM on the TALENT benchmark, but still ranked behind TabICLv2.

Summary of the key models

Model Creator Open weights? Max scale License Best for
TabPFN-2 Prior Labs Yes 10K rows Commercial with attribution Commercial use of a proven older model
TabPFN-3 / 3.5 Prior Labs (SAP) Yes, after accepting a license Up to 1M rows Non-commercial, paid API for business Top accuracy and research
TabICLv2 Inria Yes Best up to 100K rows Permissive open source Your default starting point
TabFM Google Research Yes 20 features in BigQuery Non-commercial weights SQL users on BigQuery
NEXUS Fundamental No Not disclosed Proprietary Large companies on AWS
tabH2O H2O.ai No Not disclosed Commercial API Teams with no ML setup, AI agents

How Do You Use a Tabular Foundation Model in Python?

You use a tabular foundation model in Python exactly like any other scikit-learn model.

The best way to see what it offers is to put it head to head with XGBoost, so let’s do that on scikit-learn’s built-in breast cancer dataset.

First, we install the packages:

pip install tabicl xgboost scikit-learn

Then we run both models on the same split:

from sklearn.datasets import load_breast_cancer
from sklearn.metrics import accuracy_score
from sklearn.model_selection import train_test_split
from tabicl import TabICLClassifier
from xgboost import XGBClassifier

# Load a small tabular dataset as a pandas DataFrame
X, y = load_breast_cancer(return_X_y=True, as_frame=True)
X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, stratify=y, random_state=42
)

# Tabular foundation model: fit() just stores the training rows
tfm = TabICLClassifier()
tfm.fit(X_train, y_train)
tfm_preds = tfm.predict(X_test)

# XGBoost baseline with reasonable, untuned settings
xgb = XGBClassifier(n_estimators=300, learning_rate=0.05, max_depth=4)
xgb.fit(X_train, y_train)
xgb_preds = xgb.predict(X_test)

print(f"TabICL accuracy:  {accuracy_score(y_test, tfm_preds):.3f}")
print(f"XGBoost accuracy: {accuracy_score(y_test, xgb_preds):.3f}")

This small dataset runs fine on a normal laptop CPU.

Both models will score very highly on something this clean, so please do not judge them on one split. The real test will be your own messy data.

If you want to try TabPFN instead, run pip install tabpfn, and then it is only a two-line change:

from tabpfn import TabPFNClassifier

tfm = TabPFNClassifier()  # The first run asks you to accept the license
tfm.fit(X_train, y_train)
tfm_preds = tfm.predict(X_test)

One important thing before you compare models on real data.

If your data has dates, split it by time instead of randomly. Otherwise, the model gets to peek at the future, and as you will see in the next section, that is exactly where tabular foundation models look better than they really are.

Where Do Tabular Foundation Models Work Well and Where Do They Struggle?

Tabular foundation models work well when your dataset is small to medium and the test rows look like the training rows. They struggle with time-ordered data, grouped data, very large tables, and strict explainability needs.

That “test rows look like training rows” idea has a name, IID (independent and identically distributed), and it is the single most important thing to check about your data.

Let’s start with where they work well:

  • Small-to-medium datasets: a few hundred to about 100,000 rows is the sweet spot.
  • Messy data: missing values and category columns are handled for you.
  • Quick experiments: you get a strong result in seconds, before writing any feature engineering code.
  • No ML setup: tools like BigQuery’s AI.PREDICT remove training and deployment entirely.

Now for the limitations, which I think matter more if you plan to use these in a real project.

They assume your data is IID

The BeyondArena benchmark, released in June 2026, tested 11 models on 142 datasets, including ones split by time (predicting the future from the past) and by group (predicting for a new hospital or country).

It found tabular foundation models do best on small and medium IID data, while tree-based and other deep learning models still lead on time-ordered, grouped, large, and very wide data. Since most business data (sales, fraud, churn) is ordered in time, this limitation shows up in most real projects.

TabPFN-3.5, released after BeyondArena, specifically claims to close this gap on time-ordered and grouped data. That is promising, but it has not yet been independently tested.

They are harder to explain

You can get SHAP values from both TabICL and TabPFN, but it takes much longer than SHAP on a tree model, and there are no tree splits to inspect. In areas like credit scoring, where regulators need to understand the model, this is a real problem.

They need more compute at prediction time

Anything past a few thousand rows really wants a GPU, and predictions get slower as your training data grows. A trained XGBoost model on a CPU will win on speed every time.

Licenses vary a lot

TabPFN-2 is commercial with credit, newer TabPFN versions are non-commercial, TabFM weights are non-commercial, and TabICLv2 is permissive.

BeyondArena chart of Elo by model family across task type, dataset size, dimensionality and feature type, showing tabular foundation models leading on small IID data but dropping on temporal and large datasets

Figure 3: Elo by model family (higher is better). Tabular foundation models lead on small IID data but drop on temporal and large datasets, where trees and MLPs hold up better. Blue is the best of TabICLv2, TabPFN-2.6, and TabDPT, not the newest TabPFN versions. Source: Purucker et al., “Beyond IID: How General Are Tabular Foundation Models, Really?” (2026), CC BY 4.0.

Hopefully, you will agree with me that tabular foundation models have raised the bar for how good a quick, no-effort baseline can be, but tree-based models are still the better choice in plenty of real situations.

When Should You Use a Tabular Foundation Model?

You should use a tabular foundation model when your dataset is small to medium and IID, and you do not need a fully explainable model or very fast predictions on a CPU. Here is the decision table I would use:

Scenario Recommended approach
Under 10K rows, IID, no explainability needs Start with TabICLv2 or TabPFN
10K to 100K rows, IID Test TabICLv2 and XGBoost, keep the winner
100K to 1M rows Start with XGBoost or LightGBM, try TabPFN-3 as a challenger
Data split by time or by group Start with tree-based models
Need SHAP, feature importance or auditing Tree-based model with SHAP
Fast predictions without a GPU Tree-based model
Quick proof of concept, no ML setup TabICLv2, or BigQuery AI.PREDICT if your data is already there
AI agent that needs predictions from tables An API like tabH2O or NEXUS

Before picking a model, I would recommend asking these three questions:

  1. Is my data IID? If rows are ordered by time or grouped by customer, store, or patient, be careful with any result from a random split.
  2. Do I need to explain the model? If a regulator or manager needs to audit it, lean toward trees.
  3. How fast do predictions need to be? If you are serving millions of predictions a day on CPUs, trees will be cheaper and faster.

And the last point I would like to make is this. Run a tabular foundation model first, because it only costs you a few minutes. If it cannot beat XGBoost on your data, you now know that spending time on feature engineering and tuning XGBoost is worth it.

Final Thoughts

Remember the two students from the start of this article? In 2026, the second student, the one who practiced on millions of exams, can finally beat the one who studied for weeks, at least on small and medium tables.

Tabular foundation models replace per-dataset training with pretraining, and replace fit() with learning from examples.

What surprises me most is how quickly this domain has improved. TabPFN-2 first showed in 2025 that a neural network could beat tuned trees on small tables, and TabICLv2 and TabPFN-3 have since pushed that to much larger ones.

The open problems now are time-ordered data, changing data, explainability, and licensing, which are exactly the things that decide whether a model makes it into a real product.

So my advice is to treat these models as your new starting point and pick the final tool based on your data, your limits, and where the model will run.

And if you want to master the tree-based models that still win in many real situations, check out our Machine Learning with Tree-Based Models in Python course and the Supervised Machine Learning in Python track. If you work in R, Machine Learning with Tree-Based Models in R covers the same ground.

Tabular Foundation Model FAQs

What is a tabular foundation model?

It is a neural network pretrained on millions of synthetic tables. It predicts on your dataset without training on it such as TabPFN, TabICLv2 and Google's TabFM.

How is TabPFN different from XGBoost?

XGBoost trains a new model on every dataset. TabPFN is pretrained once and reads your rows as examples at prediction time. So there is no training step, but predictions are slower.

Do I need a GPU to run popular tabular foundation models?

Do I need a GPU to run popular tabular foundation models?

Can I use tabular foundation models commercially?

It depends on the model and version. TabICLv2 is permissive. TabPFN-2 needs attribution, newer TabPFN versions and TabFM weights are non-commercial.

Do tabular foundation models handle missing values and categorical columns?

Yes. TabPFN and TabICL handle both without extra cleaning. You can skip imputation and one-hot encoding for a first run.


Vaibhav Mehra's photo
Author
Vaibhav Mehra
LinkedIn
Topics
Artificial Intelligence
Large Language Models

Top DataCamp Courses

Track

Developing Large Language Models

19 hr
Learn to develop large language models (LLMs) with PyTorch and Hugging Face, using the latest deep learning and NLP techniques.
See DetailsRight Arrow
Start Course
See MoreRight Arrow
Related

blog

Introduction to Foundation Models

Explore the concept of AI foundation models, focusing on their key characteristics, applications, and future in the AI era.
Andrea Valenzuela's photo

Andrea Valenzuela

10 min

blog

What are Foundation Models?

Discover the key technology that is powering the generative AI boom
Javier Canales Luna's photo

Javier Canales Luna

9 min

blog

Large Concept Models: A Guide With Examples

Learn what large concept models are, how they differ from LLMs, and how their architecture leads to improvements in language processing.
Amberle McKee's photo

Amberle McKee

8 min

Tutorial

What Are AI World Models? How They Work and 2026 Trends

Discover what AI world models are, how they differ from large language models, and how they help AI predict the future. Explore the latest 2026 industry trends.
Vaibhav Mehra's photo

Vaibhav Mehra

15 min

Tutorial

Google Project Genie Guide: Google's Foundation World Model Explained

A guide to Project Genie and Genie 3, Google’s interactive foundation world model, with hands-on examples.
Tim Lu's photo

Tim Lu

7 min

code-along

Introduction to Large Language Models with GPT & LangChain

Learn the fundamentals of working with large language models and build a bot that analyzes data.
Richie Cotton's photo

Richie Cotton

See MoreSee More