Skip to content

Latest commit

 

History

6 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Tabular Foundation Models Tutorial

Interactive website · Python notebook · Model landscape · Learning resources

Table as Prompt: An Interactive Guide to Tabular Foundation Models - accepted for presentation at the NeurIPS 2026 Education Track.

Tabular Foundation Models Tutorial: Interactive Guide, Papers and Code, and Benchmarks. Tabular in-context learning uses labeled examples and a new row's features as inputs to a pretrained Transformer with fixed weights to predict the new row's label.

What are tabular foundation models?

Tabular foundation models (TFMs) are pretrained predictors designed for reuse across tabular datasets. Conventional workflows fit and often tune a separate model for each dataset. PFN-style TFMs pretrain a shared predictor across sampled tasks. At inference, labeled rows provide context for predicting unlabeled query rows, without task-specific gradient updates. This adaptation through examples is tabular in-context learning.

Prior-data fitted networks (PFNs) learn from tasks sampled from a prior. During pretraining, the model predicts held-out query labels from context rows. The task distribution and expected query loss shape its learned inductive bias. TabPFN demonstrated this approach for small tabular classification tasks. Our guide uses TabICLv2 as a worked example; its architecture and inference procedure are model-specific, not shared by every TFM.

The broader goal is reusable prediction for structured data. The field now covers classification and regression across different table sizes, with open questions about robustness, scale, efficiency, and evaluation. Strong task-specific methods, including boosted trees, remain baselines for assessing where these models help.

Model landscape

TFMs differ in what they learn from and how they adapt. Synthetic-prior models learn across generated tasks; real-data models learn across existing tables. Other work brings text semantics into prediction or extends the size of the inference context.

Selected milestones through October 4, 2026. Each date links to a primary source and identifies the event: a release, an announcement, or a paper. Preprint dates are the first arXiv submission dates; they do not establish when code or weights became available. Team labels identify a lead institution or the project maintainer; papers list the full author affiliations.

Date Model / team Research direction Dated source Code / weights
2022-07-05 TabPFN · AutoML · Freiburg logo AutoML · Freiburg Synthetic-prior pretraining for in-context classification on small tables. Preprint GitHub
2024-10-23 TabDPT · Layer 6 logo Layer 6 Self-supervised pretraining on real tables, with retrieval for context selection. Preprint GitHub · Hugging Face
2025-01-08 TabPFNv2 · Prior Labs logo Prior Labs Extends tabular prediction to regression and larger datasets. Nature paper GitHub · Hugging Face
2025-02-08 TabICL · Inria logo Inria Builds row embeddings before in-context learning to handle larger tables. Preprint GitHub · Hugging Face
2025-05-23 TabSTAR · Technion Transfer learning with target-aware text representations. Preprint GitHub · Hugging Face
2025-06-12 ConTextTab · SAP logo SAP Combines semantic embeddings with tabular ICL and real-data pretraining. Preprint GitHub · Hugging Face
2025-07-22 Mitra · AWS · AutoGluon logo AWS · AutoGluon Uses a mixture of synthetic priors for classification and regression. Release announcement GitHub · Hugging Face
2025-09-03 LimiX · Stable AI logo Stable AI Models joint distributions over table variables and missingness. Preprint GitHub · Hugging Face
2025-11-11 TabPFN-2.5 · Prior Labs logo Prior Labs Expands table size and adds distillation into smaller predictors. Preprint GitHub · Hugging Face
2026-02-11 TabICLv2 · Inria logo Inria Adds regression and revises synthetic priors, attention, and pretraining. Preprint GitHub · Hugging Face
2026-05-13 TabPFN-3 · Prior Labs logo Prior Labs Scales in-context prediction to million-row datasets. Preprint GitHub · Hugging Face
2026-06-12 Nori · Synthefy logo Synthefy Synthetic-data pretraining for in-context regression with quantile predictions. Release announcement GitHub · Hugging Face
2026-06-30 TabFM · Google Research logo Google Research Google's synthetic-data model for in-context classification and regression. Announcement Project
2026-09-15 TabPFN-3.5 · Prior Labs logo Prior Labs Extends evaluation and capabilities to temporal, grouped, and mixed-modality tables. Release announcement GitHub · Hugging Face
2026-09-29 Kumo Tabular · NVIDIA logo NVIDIA Synthetic-data pretraining with column, row, and in-context attention for classification and regression. Release announcement GitHub · Hugging Face

Use the interactive guide

Open the guide in a modern browser. Follow the sequence from tabular prediction and adaptation through PFN theory, TabICLv2 training, and inference. Then try the browser playground and inspect the model's intermediate computations. Basic supervised learning is enough to get started.

TabICLv2 is the worked example. The guide distinguishes explanatory simulations from real model execution; browser inference runs locally. For a Python exercise, use the notebook or open it in Colab. Setup instructions cover the notebook environment.

Tutorials and learning resources

Resource What to learn
Tabular Foundation Models, Christoph Molnar An online book covering PFNs, in-context learning, pretraining, and practical prediction examples.
TabPFN documentation Official quickstarts and examples for applying TabPFN to classification and regression.
TabICL documentation Official usage instructions, configuration, and API reference.
nanoTabPFN A small educational implementation for studying the model and training loop.
nanoTabICL A compact implementation of the TabICLv2 architecture and a simplified synthetic-data prior.
Transformer Explainer An interactive introduction to attention and transformer computation, using a language model.

For a conceptual introduction, start with Molnar's book. To run a model on your own table, use the official documentation. To study how it is built, read the nano implementations alongside the papers below.

Papers and model implementations

These are selected entry points into PFN-based tabular prediction. Follow each project's documentation for available checkpoints, supported tasks, and terms of use.

Topic Paper Code and documentation
PFN foundations Transformers Can Do Bayesian Inference PFNs
TabPFN TabPFN-3 technical report Official code · Documentation and model access
TabICL TabICLv2 paper Official code and checkpoints · Documentation

Benchmarks and evaluation

Resource What it offers Sources
TabArena A maintained benchmark for comparing tabular models under documented evaluation settings. Paper · Code · Datasets · Leaderboard
TALENT A toolkit for comparing classical and deep tabular methods, with datasets and preprocessing options. Paper · Code and datasets
TabZilla An empirical study of when neural networks and boosted trees perform well on tabular data. Paper · Code and datasets

Read results with their dataset splits, tuning budgets, and ensembling settings. Scores from separate papers do not form a single ranking. Use these resources to choose an evaluation protocol and baselines for your own task.

Run this guide locally

Requires Node.js 20+, npm 10+, and Python 3. From the repository root:

cd materials/tabicl-explainer
npm ci
npm run build
cd ../website
./serve.sh 8000

Open localhost:8000. Serve over HTTP so that module workers and model assets load correctly.

Attribution and license

Browser model provenance documents the checkpoint and export. The explorer README credits the adaptation of Transformer Explainer.

Unless otherwise noted, original software and code are licensed under the Apache License, Version 2.0, and original educational content is licensed under Creative Commons Attribution 4.0 International. Copyright (c) 2026, Affirm, Inc. All rights reserved. See NOTICE for the project notice and THIRD_PARTY_NOTICES for the separate terms and attributions that apply to third-party software, model artifacts, adapted materials, and data.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages