NVIDIA Kumo Tabular Sets a New Accuracy-Efficiency Frontier for Tabular Prediction
NVIDIA releases Kumo Tabular, a foundation model for tabular data that predicts labels without training or feature engineering. Weights and code are available on Hugging Face and GitHub, and the model tops four benchmarks. It runs via an open‑source library under a commercial‑friendly license.

NVIDIA released Kumo Tabular, an open foundation model designed to process tabular data without requiring traditional training or feature engineering. The model uses in-context learning to predict labels for new rows in a single forward pass, supporting both classification and regression tasks. Three versions exist, with parameter counts ranging from 28M to 215M. The weights and code are available for download, and the project is distributed under a commercial-friendly license. This release addresses the slow lifecycle of traditional machine learning pipelines, which typically demand extensive data preparation and hyperparameter tuning for every new task. By leveraging in-context learning, the system allows a pretrained model to solve tasks using only a few examples, similar to large language models for text. The developers claim the model ranks first on four specific benchmarks, including TabArena and BeyondArena. It also uses a causal simulation approach to generate its pretraining data, aiming to capture complex structural relationships within tables. The article does not specify the computational resources required for inference or the absolute time savings compared to gradient-boosted trees in real-world enterprise environments. While the text highlights top rankings on four benchmarks, it does not detail the specific margin of victory over existing methods. The source cuts off before describing the final steps of the data generation process, leaving the full scope of the causal simulation incomplete. No independent third-party evaluation of the accuracy claims is provided in this text.