Classical machine learning on tabular data with scikit-learn 1.9.x (Python 3.14): the estimator
API, preprocessing, pipelines, model choice, cross-validation, tuning, metrics and persistence.
Arrays are covered in NumPy, DataFrames in pandas,
and neural networks in PyTorch.
Setup
uv add scikit-learn pandas skops # or pip install in a venvuv run python -c "import sklearn; sklearn.show_versions()"
seeds anything random (splits, trees, init); pass an int for reproducible runs
n_jobs=-1
use every core (joblib); None means 1 unless set by joblib.parallel_config
sklearn.set_config(...)
global switches: transform_output, enable_metadata_routing, array_api_dispatch
Estimator API
Every object follows the same contract: hyperparameters go in the constructor, fit learns from
data and returns self, and whatever was learned is stored in attributes ending in _.
Method / attribute
On
Does
Est(**hyperparams)
all
stores the arguments verbatim; no validation, no work
fit(X, y)
all
learns state; returns self so calls chain
predict(X)
predictors
labels (classifiers) or values (regressors)
predict_proba(X)
most classifiers
class probabilities, shape (n, n_classes), columns in classes_ order
decision_function(X)
linear models, SVMs
raw scores; use for ROC AUC when there is no predict_proba
transform(X)
transformers
returns the transformed features
fit_transform(X, y=None)
transformers
fit then transform in one pass (often faster); train data only
fit_predict(X)
clusterers
fit and return cluster labels
score(X, y)
predictors
accuracy (classifiers), R² (regressors)
get_params() / set_params(**p)
all
read/change hyperparameters; set_params needs a refit
from sklearn.model_selection import train_test_splitX_train, X_test, y_train, y_test = train_test_split( X, y, test_size=0.2, # float = fraction, int = rows stratify=y, # keep class ratios in both parts random_state=42, # reproducible shuffle)
Situation
Do
Classification
stratify=y, especially with rare classes
Rows grouped (per user, per patient)
GroupShuffleSplit / GroupKFold so a group never spans train and test
Time series
shuffle=False or TimeSeriesSplit: train on the past, test on the future
Tuning hyperparameters
keep a final test set untouched; tune with CV on the training part
Very small data
skip the hold-out; report cross-validated scores
Preprocessing
All live in sklearn.preprocessing or sklearn.impute and follow fit/transform.
Transformer
Does
Use when
StandardScaler
zero mean, unit variance
linear models, SVM, kNN, PCA, neural nets
MinMaxScaler
rescale to [0, 1]
bounded inputs; sensitive to outliers
RobustScaler
median / IQR
heavy outliers
MaxAbsScaler
divide by max abs
sparse data (keeps zeros)
PowerTransformer / QuantileTransformer
make skewed features Gaussian-like / uniform
long tails (income, counts)
OneHotEncoder(handle_unknown="ignore")
one column per category
nominal categories, low cardinality
OrdinalEncoder
category to integer
ordered categories; tree models
TargetEncoder
category to (cross-fitted) mean target
high-cardinality categories
LabelEncoder
encode y labels
targets only, never features
SimpleImputer(strategy=...)
fill with mean/median/most_frequent/constant
missing values; add_indicator=True keeps a "was missing" flag
KNNImputer
fill from nearest rows
small data, correlated features
IterativeImputer
model each feature from the others
needs from sklearn.experimental import enable_iterative_imputer
PolynomialFeatures / SplineTransformer
interactions / smooth non-linear bases
linear models on curved data
KBinsDiscretizer
bucket numeric features
coarse non-linearity
FunctionTransformer(np.log1p)
wrap any function
stateless feature tweaks
Tree ensembles (RandomForest*, HistGradientBoosting*) do not need scaling, and
HistGradientBoosting* handles NaN and categorical columns natively.
Pipelines & ColumnTransformer
A Pipeline chains transformers and a final estimator into one estimator, so every preprocessing
step is refit inside each CV fold and nothing from validation data leaks into training.
auto-names steps from class names (standardscaler)
make_column_selector(dtype_include="number")
pick columns by dtype or regex
Pipeline(..., memory="cache/")
cache fitted transformers during searches
Leak
Fix
Scaling / imputing on all data, then splitting
put the transformer in the pipeline
Feature selection before CV
put SelectKBest etc. in the pipeline
Target encoding outside CV
TargetEncoder in the pipeline (it cross-fits internally)
Shuffled CV on time series
TimeSeriesSplit
Duplicates or same entity in train and test
group-aware splitters
Choosing a model
Task
Estimator
Good for
Scale features?
Baseline
DummyClassifier / DummyRegressor
the score to beat
no
Regression
LinearRegression, Ridge, Lasso, ElasticNet
fast, interpretable; Lasso zeroes features
yes (for penalties)
Classification
LogisticRegression
strong linear baseline, calibrated-ish probabilities
yes
Both
DecisionTreeClassifier / Regressor
explainable rules; overfits alone
no
Both
RandomForestClassifier / Regressor
robust default, little tuning
no
Both
HistGradientBoostingClassifier / Regressor
best default on tabular data; fast on 10k+ rows; NaN and categories built in
no
Both
GradientBoostingClassifier / Regressor
small data; slower than the Hist variant
no
Both
SVC / SVR (RBF kernel)
small-to-medium data, non-linear boundaries; scales badly past ~10k rows
yes
Both
LinearSVC, SGDClassifier / SGDRegressor
large or sparse (text) data
yes
Both
KNeighborsClassifier / Regressor
low-dimensional, local structure
yes
Classification
GaussianNB, MultinomialNB
tiny data; word counts (MultinomialNB)
no
Clustering
KMeans / MiniBatchKMeans
round clusters, known k
yes
Clustering
HDBSCAN, DBSCAN
arbitrary shapes, noise points, unknown k
yes
Clustering
AgglomerativeClustering, GaussianMixture
hierarchies; soft assignments
yes
Dimensionality
PCA, TruncatedSVD (sparse)
compression, decorrelation, plotting
yes
Anomalies
IsolationForest, LocalOutlierFactor
outlier scores
depends
The deprecations to know about: LogisticRegression(penalty=...) is deprecated since 1.8. Use
l1_ratio (0 is L2, 1 is L1) and C=np.inf for no penalty. SVC(probability=True) is
deprecated in 1.9; see calibration.
model-agnostic: score drop when one column is shuffled; run on held-out data
feature_importances_ (trees)
impurity-based (MDI); biased toward high-cardinality and continuous features, computed on training data
coef_ (linear)
comparable only after scaling; correlated features share weight
PartialDependenceDisplay.from_estimator
how predictions change with one or two features
SHAP (third party)
per-prediction attributions
from sklearn.inspection import permutation_importancer = permutation_importance( model, df, target, # use a held-out set in practice n_repeats=10, scoring="roc_auc", random_state=0,)ranked = sorted( zip(df.columns, r.importances_mean, r.importances_std), key=lambda t: -t[1],)for name, mean, std in ranked: print(f"{name:8} {mean:.3f} ± {std:.3f}")
Permutation importance works on the raw input columns of a pipeline, so city is scored as one
feature rather than per one-hot column. Correlated features hide each other's importance.
Probability calibration
A classifier is calibrated when "0.8" means right about 80% of the time. Trees, boosting, SVMs
and naive Bayes often are not.
Tool
Does
CalibratedClassifierCV(est, method="sigmoid")
Platt scaling; fine with little data
method="isotonic"
non-parametric; needs 1000+ calibration samples
method="temperature"
1.8+; one parameter, preserves ranking, good for multiclass
cv=5
fits the base model and calibrator per fold
CalibratedClassifierCV(FrozenEstimator(fitted))
calibrate an already fitted model on new data (replaces the removed cv="prefit")
pickle underneath: loading runs arbitrary code; efficient with big arrays
pickle
no
same risk, same version constraints
skops.io.dump / load
yes, with an allow-list
refuses unknown types until you trust them explicitly
ONNX (skl2onnx)
yes
serve from other runtimes (ONNX Runtime, including the browser)
Only unpickle files you created or fully trust.
Pin the scikit-learn version: loading in another version raises InconsistentVersionWarning
and may give wrong results. Save the training code and data version too.
import skops.io as siosio.dump(model, "model.skops")unknown = sio.get_untrusted_types(file="model.skops")print(unknown) # review before trustingloaded = sio.load("model.skops", trusted=unknown)assert (loaded.predict(df) == model.predict(df)).all()
Output config & routing
Feature
API
Notes
DataFrame output
est.set_output(transform="pandas")
also "polars"; global: set_config(transform_output="pandas")
Metadata routing
set_config(enable_metadata_routing=True)
pass sample_weight, groups through pipelines and searches
Request metadata
est.set_fit_request(sample_weight=True)
each consumer opts in; also set_score_request
Array API
set_config(array_api_dispatch=True)
experimental; some estimators accept PyTorch/CuPy arrays and stay on GPU; needs array-api-compat and SCIPY_ARRAY_API=1
Callbacks
est.set_callbacks(ProgressBar())
experimental in 1.9 (sklearn.callback); few estimators support it
Sparse type
set_config(sparse_interface="sparray")
1.9+; SciPy sparse arrays instead of matrices
import sklearnfrom sklearn.linear_model import LogisticRegressionfrom sklearn.model_selection import cross_validatefrom sklearn.pipeline import make_pipelinefrom sklearn.preprocessing import StandardScalerpre.set_output(transform="pandas")pre.fit_transform(df).head() # named DataFrame columnswith sklearn.config_context(enable_metadata_routing=True): w = np.where(y == 1, 2.0, 1.0) scaler = StandardScaler().set_fit_request( sample_weight=False # opt out explicitly ) lr = ( LogisticRegression() .set_fit_request(sample_weight=True) .set_score_request(sample_weight=True) ) routed = make_pipeline(scaler, lr) routed.fit(X, y, sample_weight=w) # reaches only LR cross_validate( routed, X, y, params={"sample_weight": w} )
Unrequested metadata raises UnsetMetadataPassedError, so a weight is never dropped silently.
Recipes
Tabular pipeline on mixed columns
When a DataFrame mixes numbers, categories and gaps: gradient boosting with native categorical support needs almost no preprocessing.
When a trained pipeline ships to another process; skops refuses types you haven't approved.
from pathlib import Pathimport skops.io as siofrom sklearn.datasets import make_classificationfrom sklearn.linear_model import LogisticRegressionfrom sklearn.pipeline import make_pipelinefrom sklearn.preprocessing import StandardScalerX, y = make_classification(random_state=0)pipe = make_pipeline(StandardScaler(), LogisticRegression())pipe.fit(X, y)path = Path("pipe.skops")sio.dump(pipe, path)# later, in the serving process:untrusted = sio.get_untrusted_types(file=path)assert untrusted == [], f"review: {untrusted}"restored = sio.load(path)assert (restored.predict(X) == pipe.predict(X)).all()
Custom transformer
When a feature step needs learned state and should work inside Pipeline, clone and grid search.
import numpy as npfrom numpy.typing import ArrayLike, NDArrayfrom sklearn.base import BaseEstimator, TransformerMixinfrom sklearn.utils.validation import ( check_is_fitted, validate_data,)class QuantileClipper(TransformerMixin, BaseEstimator): """Clip each column to its fitted quantile range.""" def __init__(self, q: float = 0.01) -> None: self.q = q # store params as given, no logic def fit(self, X: ArrayLike, y: object = None): Xv = validate_data(self, X) # sets n_features_in_ self.low_ = np.quantile(Xv, self.q, axis=0) self.high_ = np.quantile(Xv, 1 - self.q, axis=0) return self def transform(self, X: ArrayLike) -> NDArray[np.float64]: check_is_fitted(self) Xv = validate_data(self, X, reset=False) return np.clip(Xv, self.low_, self.high_)
TransformerMixin supplies fit_transform. Pass-through column names come from
OneToOneFeatureMixin, and sklearn.utils.estimator_checks.check_estimator tests the contract.