
A practitioner's guide to tabular modeling, feature engineering, gradient boosting, deep tabular learning, causal inference, and decision intelligence.
Structured data is the substrate most deployed AI actually runs on: transactions, claims, customers, patients, loans, and the decisions made from them. This book is one connected journey through the theory, models, and engineering of tabular AI. It starts with the foundations of tables, targets, and leakage-safe validation, builds through linear models, gradient boosting, and deep tabular learning, then moves into imbalance and risk, explainability, causal inference, and decision optimization, before closing with the privacy, monitoring, and deployment concerns that govern real systems.
Each part stands on the one before it; together they span the full craft of turning structured records into reliable decisions.
Tables, schemas, and keys; data quality and missingness; target and label engineering; and the leakage-safe validation every honest result depends on.
5 chapters · 5 labsThe statistical spine: exploratory analysis, linear and generalized linear models, trees and ensembles, and the gradient boosting that remains tabular AI's baseline to beat.
5 chapters · 5 labsThe central craft: numeric and categorical encodings, time and event aggregation, point-in-time correctness, external and embedding features, and reusable feature stores.
5 chapters · 5 labsThe frontier beyond boosting: neural tabular networks, relational and graph-enhanced learning, AutoML and hyperparameter search, and the emerging idea of tabular foundation models.
5 chapters · 5 labsThe problems that dominate deployment: class imbalance, fraud and abuse, credit and risk scoring, anomaly detection, and survival and time-to-event modeling.
5 chapters · 5 labsMaking models understandable enough for high-stakes use: PDP, ICE, and SHAP; counterfactuals and recourse; error analysis; fairness; and governance.
5 chapters · 5 labsBeyond prediction into decisions: causal graphs, A/B testing, observational causal inference, uplift modeling, and optimizing actions under real constraints.
5 chapters · 5 labsImproving the data itself: data-centric methods, synthetic tabular generation, privacy-preserving learning, and LLM-assisted tabular workflows.
4 chapters · 4 labsTurning models into reliable systems: production scoring architecture, drift monitoring, continuous training and MLOps, human-in-the-loop review, cost and reliability.
5 chapters · 5 labsThe field synthesized across finance, healthcare, retail, and operations, the research frontier, and an end-to-end decision-intelligence capstone.
6 chapters · 6 labsFive habits, kept in every chapter from the first join to the final deployed decision.
Time-aware and entity-aware splits, point-in-time features, and explicit leakage detectors run through every chapter, because an optimistic split is the most common way tabular results lie.
Regularized linear models and gradient boosting come before deep and graph models, so you always know what a new method actually has to beat.
Pitfalls, math asides, real industry examples, and cross-references are typeset as distinct boxes, so you can read deep or skim fast and never miss a trap.
Each chapter closes with a hands-on lab on real datasets (credit, fraud, churn, clinical), from quick checks to portfolio-ready projects.
Every model connects to a cost matrix, threshold, or policy; the book measures utility and decision quality, not accuracy in the abstract.
Building Tabular AI is one of the connected books in the series, each a deep, build-it-yourself guide to a major field of AI.
Hands-On AI Science is a series of in-depth guides to the major fields of artificial intelligence. Every book goes deep into the theory, models, and internals, covering the classical foundations and the most recent ideas, then shows you how to build each one in Python with the modern libraries and tools that get the job done. The writing stays plain and light (illustrations, analogies, mental models, worked examples, and a little fun) without trading away rigor or coverage. Each volume is self-contained and complete enough to anchor a full course on its subject.
From Structured Data to Decision Intelligence.
You are here