Overview
NVIDIA Kumo Tabular is an open-source tabular foundation model released by NVIDIA on September 29, 2026. It belongs to the NVIDIA Kumo Structured model series, and its weights are available on Hugging Face. It is designed for classification and regression tasks on tabular data. Its core feature is that it requires no task-specific training, hyperparameter tuning, feature engineering, or missing value imputation. You only need to provide a labeled table and the rows to be predicted, and the model returns classification probabilities or numerical predictions in a single forward pass. The model uses a Transformer designed around tabular structure, combining column attention, row attention, and context attention, and offers three sizes from 28M to 215M parameters. Pretraining uses entirely synthetic tables generated by a structural causal model sampler.
Key Features
- Prediction in a single forward pass: Given a labeled table and the rows to be predicted, the model returns classification probabilities or numerical predictions in a single forward pass, with no task-specific training, hyperparameter tuning, feature engineering, or missing value imputation required. Both classification and regression are supported.
- Transformer architecture designed for tabular structure: It combines column attention, row attention, and context attention introduced by TabICL and TabPFN: cells are embedded via Fourier features, row embeddings alternate between column attention and row attention with rotary positions and are read out using four learnable CLS tokens, and in the final context Transformer, query rows attend only to context rows, so the context key-values can be computed once and reused.
- Length-aware attention temperature: It grows logarithmically with the number of keys, mitigating attention divergence as tables become larger.
- Regression outputs 999 quantiles: This provides point predictions and uncertainty estimates.
- Three parameter sizes and synthetic pretraining: It offers three sizes from 28M to 215M parameters. Pretraining uses entirely synthetic tables generated by a structural causal model sampler, covering missing values, coarsening, heavy tails, and other cases. The three stages scale progressively from 1,024 rows to 60,000 rows and up to 100 columns.
- Open-source license and commercial support: The code is at github.com/NVIDIA/structured-data-models, under the OpenMDW-1.1 license, which allows commercial use.
Use Cases
- In scenarios with only one labeled table and rows to be predicted, directly perform classification or regression prediction, eliminating training, hyperparameter tuning, and feature engineering.
- For regression tasks that require uncertainty estimates, the model's 999 quantile outputs can be used to obtain point predictions and uncertainty estimates.
- For scenarios with large tables where reducing repeated computation is desired, users can benefit from the design where context key-values are computed once and reused.
- For tabular data with missing values, coarsening, heavy tails, and similar conditions, the model can be used directly for prediction.
- For tabular modeling projects that require commercial deployment, this open-source model can be used under the OpenMDW-1.1 license.
Pros
- Classification and regression prediction can be completed in a single forward pass, with no training, hyperparameter tuning, or feature engineering required.
- No missing value imputation is needed, reducing the burden of tabular data preprocessing.
- Ranked first on TabArena with 1950 ELO, and ranked first on BeyondArena (1418 ELO), TALENT, and ScoringBench.
- 17x faster than LimiX-2 on a single RTX 6000 Pro.
- Regression outputs 999 quantiles, providing both point predictions and uncertainty estimates.
- Uses the OpenMDW-1.1 license, allows commercial use, and the weights are available on Hugging Face.
Pricing
This model is open source, uses the OpenMDW-1.1 license, allows commercial use, and its weights are available on Hugging Face. Specific pricing information is not provided; please refer to the official website.
Summary
NVIDIA Kumo Tabular is an open-source tabular foundation model released on September 29, 2026. It supports classification and regression in a single forward pass, with no training, hyperparameter tuning, feature engineering, or missing value imputation required. The model offers three sizes from 28M to 215M parameters, ranks first on TabArena, BeyondArena, TALENT, and ScoringBench, is 17x faster than LimiX-2 on a single RTX 6000 Pro, and uses the OpenMDW-1.1 license to allow commercial use.
Version History
- NVIDIA releases Kumo Tabular, an open tabular foundation model (2026-09-29): Given a table with labeled rows it returns class probabilities or regression predictions in a single forward pass, with no training, tuning or feature engineering. It comes in three sizes from 28M to 215M parameters, ranks first on TabArena at 1950 ELO and tops BeyondArena, TALENT and ScoringBench, licensed for commercial use under OpenMDW-1.1.