../

scikit-learn

Classical machine learning on tabular data with scikit-learn 1.9.x (Python 3.14): the estimator API, preprocessing, pipelines, model choice, cross-validation, tuning, metrics and persistence. Arrays are covered in NumPy, DataFrames in pandas, and neural networks in PyTorch.

Setup

uv add scikit-learn pandas skops  # or pip install in a venv
uv run python -c "import sklearn; sklearn.show_versions()"
ConventionMeaning
Xfeatures, shape (n_samples, n_features): NumPy array, SciPy sparse, pandas or polars DataFrame
ytarget, shape (n_samples,) (or (n_samples, n_outputs))
random_state=0seeds anything random (splits, trees, init); pass an int for reproducible runs
n_jobs=-1use every core (joblib); None means 1 unless set by joblib.parallel_config
sklearn.set_config(...)global switches: transform_output, enable_metadata_routing, array_api_dispatch

Estimator API

Every object follows the same contract: hyperparameters go in the constructor, fit learns from data and returns self, and whatever was learned is stored in attributes ending in _.

Method / attributeOnDoes
Est(**hyperparams)allstores the arguments verbatim; no validation, no work
fit(X, y)alllearns state; returns self so calls chain
predict(X)predictorslabels (classifiers) or values (regressors)
predict_proba(X)most classifiersclass probabilities, shape (n, n_classes), columns in classes_ order
decision_function(X)linear models, SVMsraw scores; use for ROC AUC when there is no predict_proba
transform(X)transformersreturns the transformed features
fit_transform(X, y=None)transformersfit then transform in one pass (often faster); train data only
fit_predict(X)clusterersfit and return cluster labels
score(X, y)predictorsaccuracy (classifiers), R² (regressors)
get_params() / set_params(**p)allread/change hyperparameters; set_params needs a refit
clone(est)allfresh unfitted copy with the same hyperparameters
name_ (trailing underscore)fittedlearned state: coef_, classes_, mean_, n_features_in_, feature_names_in_
estimator_api.py
import numpy as np
from sklearn.base import clone
from sklearn.linear_model import LogisticRegression
from sklearn.preprocessing import StandardScaler
 
rng = np.random.default_rng(0)
X = rng.normal(size=(200, 3))       # (n_samples, n_features)
y = (X[:, 0] + X[:, 1] > 0).astype(int)
 
scaler = StandardScaler()           # hyperparameters only
Xs = scaler.fit_transform(X)        # learn mean/std, apply
print(scaler.mean_, scaler.scale_)  # learned: trailing _
 
clf = LogisticRegression(C=1.0).fit(Xs, y)  # returns self
clf.predict(Xs[:3])                 # array([1, 0, 1])
clf.predict_proba(Xs[:3])           # (3, 2), rows sum to 1
clf.score(Xs, y)                    # mean accuracy
clf.coef_, clf.intercept_, clf.classes_
 
fresh = clone(clf).set_params(C=0.1)  # unfitted copy

Splitting data

from sklearn.model_selection import train_test_split
 
X_train, X_test, y_train, y_test = train_test_split(
    X, y,
    test_size=0.2,       # float = fraction, int = rows
    stratify=y,          # keep class ratios in both parts
    random_state=42,     # reproducible shuffle
)
SituationDo
Classificationstratify=y, especially with rare classes
Rows grouped (per user, per patient)GroupShuffleSplit / GroupKFold so a group never spans train and test
Time seriesshuffle=False or TimeSeriesSplit: train on the past, test on the future
Tuning hyperparameterskeep a final test set untouched; tune with CV on the training part
Very small dataskip the hold-out; report cross-validated scores

Preprocessing

All live in sklearn.preprocessing or sklearn.impute and follow fit/transform.

TransformerDoesUse when
StandardScalerzero mean, unit variancelinear models, SVM, kNN, PCA, neural nets
MinMaxScalerrescale to [0, 1]bounded inputs; sensitive to outliers
RobustScalermedian / IQRheavy outliers
MaxAbsScalerdivide by max abssparse data (keeps zeros)
PowerTransformer / QuantileTransformermake skewed features Gaussian-like / uniformlong tails (income, counts)
OneHotEncoder(handle_unknown="ignore")one column per categorynominal categories, low cardinality
OrdinalEncodercategory to integerordered categories; tree models
TargetEncodercategory to (cross-fitted) mean targethigh-cardinality categories
LabelEncoderencode y labelstargets only, never features
SimpleImputer(strategy=...)fill with mean/median/most_frequent/constantmissing values; add_indicator=True keeps a "was missing" flag
KNNImputerfill from nearest rowssmall data, correlated features
IterativeImputermodel each feature from the othersneeds from sklearn.experimental import enable_iterative_imputer
PolynomialFeatures / SplineTransformerinteractions / smooth non-linear baseslinear models on curved data
KBinsDiscretizerbucket numeric featurescoarse non-linearity
FunctionTransformer(np.log1p)wrap any functionstateless feature tweaks

Tree ensembles (RandomForest*, HistGradientBoosting*) do not need scaling, and HistGradientBoosting* handles NaN and categorical columns natively.

Pipelines & ColumnTransformer

A Pipeline chains transformers and a final estimator into one estimator, so every preprocessing step is refit inside each CV fold and nothing from validation data leaks into training.

pipeline.py
import numpy as np
import pandas as pd
from sklearn.compose import ColumnTransformer
from sklearn.impute import SimpleImputer
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import Pipeline, make_pipeline
from sklearn.preprocessing import (
    OneHotEncoder,
    StandardScaler,
)
 
rng = np.random.default_rng(0)
n = 400
df = pd.DataFrame({
    "age": rng.integers(18, 80, n).astype(float),
    "income": rng.lognormal(10, 0.5, n),
    "city": rng.choice(["ams", "ber", "par"], n),
})
df.loc[rng.random(n) < 0.1, "age"] = np.nan
noise = rng.lognormal(0, 0.3, n)
target = (df["income"] * noise > 25_000).astype(int)
 
num = make_pipeline(
    SimpleImputer(strategy="median"), StandardScaler()
)
cat = OneHotEncoder(
    handle_unknown="ignore", sparse_output=False
)
pre = ColumnTransformer(
    [
        ("num", num, ["age", "income"]),
        ("cat", cat, ["city"]),
    ],
    remainder="drop",   # or "passthrough"
)
model = Pipeline(
    [("pre", pre), ("clf", LogisticRegression())]
)
model.fit(df, target)
AccessGives
model["clf"], model.named_steps["clf"]a step
model[:-1]sub-pipeline of every step but the last
model.set_params(clf__C=0.5)nested param: step__param, deeper with more __
model.set_params(clf="passthrough")disable or swap a step
model[:-1].get_feature_names_out()output column names, e.g. num__age, cat__city_ams
make_pipeline(a, b)auto-names steps from class names (standardscaler)
make_column_selector(dtype_include="number")pick columns by dtype or regex
Pipeline(..., memory="cache/")cache fitted transformers during searches
LeakFix
Scaling / imputing on all data, then splittingput the transformer in the pipeline
Feature selection before CVput SelectKBest etc. in the pipeline
Target encoding outside CVTargetEncoder in the pipeline (it cross-fits internally)
Shuffled CV on time seriesTimeSeriesSplit
Duplicates or same entity in train and testgroup-aware splitters

Choosing a model

TaskEstimatorGood forScale features?
BaselineDummyClassifier / DummyRegressorthe score to beatno
RegressionLinearRegression, Ridge, Lasso, ElasticNetfast, interpretable; Lasso zeroes featuresyes (for penalties)
ClassificationLogisticRegressionstrong linear baseline, calibrated-ish probabilitiesyes
BothDecisionTreeClassifier / Regressorexplainable rules; overfits aloneno
BothRandomForestClassifier / Regressorrobust default, little tuningno
BothHistGradientBoostingClassifier / Regressorbest default on tabular data; fast on 10k+ rows; NaN and categories built inno
BothGradientBoostingClassifier / Regressorsmall data; slower than the Hist variantno
BothSVC / SVR (RBF kernel)small-to-medium data, non-linear boundaries; scales badly past ~10k rowsyes
BothLinearSVC, SGDClassifier / SGDRegressorlarge or sparse (text) datayes
BothKNeighborsClassifier / Regressorlow-dimensional, local structureyes
ClassificationGaussianNB, MultinomialNBtiny data; word counts (MultinomialNB)no
ClusteringKMeans / MiniBatchKMeansround clusters, known kyes
ClusteringHDBSCAN, DBSCANarbitrary shapes, noise points, unknown kyes
ClusteringAgglomerativeClustering, GaussianMixturehierarchies; soft assignmentsyes
DimensionalityPCA, TruncatedSVD (sparse)compression, decorrelation, plottingyes
AnomaliesIsolationForest, LocalOutlierFactoroutlier scoresdepends

The deprecations to know about: LogisticRegression(penalty=...) is deprecated since 1.8. Use l1_ratio (0 is L2, 1 is L1) and C=np.inf for no penalty. SVC(probability=True) is deprecated in 1.9; see calibration.

Cross-validation

from sklearn.model_selection import (
    StratifiedKFold,
    cross_val_predict,
    cross_val_score,
)
 
cv = StratifiedKFold(
    n_splits=5, shuffle=True, random_state=0
)
scores = cross_val_score(
    model, df, target, cv=cv, scoring="roc_auc"
)
print(f"{scores.mean():.3f} ± {scores.std():.3f}")
oof = cross_val_predict(model, df, target, cv=cv)  # per row
SplitterUse for
cv=5 (int)StratifiedKFold for classifiers, KFold otherwise; no shuffle
KFold(shuffle=True)regression, i.i.d. rows
StratifiedKFoldclassification, keeps class ratios per fold
GroupKFold / StratifiedGroupKFoldrows from one group stay in one fold (groups= array)
LeaveOneGroupOutone fold per group (per site, per subject)
TimeSeriesSplit(n_splits, gap=)expanding window; training always precedes test
ShuffleSplit / StratifiedShuffleSplitmany random train/test splits
RepeatedStratifiedKFoldrepeats with new shuffles for tighter estimates
PredefinedSplita fixed validation set expressed as CV
FunctionReturns
cross_val_scoreone score per fold (one metric)
cross_validatedict of test_<metric>, fit_time, optionally train scores and fitted estimators
cross_val_predictout-of-fold predictions for every row (for plots and reports, not a score)
learning_curve / validation_curvescores vs training size / vs one hyperparameter
SearchTriesUse when
GridSearchCVevery combinationfew parameters, few values
RandomizedSearchCVn_iter samples from lists or distributionsmany parameters; continuous ranges
HalvingGridSearchCVgrid, successive halvingbig grids; bad candidates get few resources
HalvingRandomSearchCVrandom candidates, halvingthe fast default for large spaces

Halving searches are still experimental: from sklearn.experimental import enable_halving_search_cv.

from scipy.stats import loguniform
from sklearn.model_selection import RandomizedSearchCV
 
search = RandomizedSearchCV(
    model,
    {"clf__C": loguniform(1e-3, 1e2)},  # step__param
    n_iter=20,
    cv=5,
    scoring="roc_auc",
    n_jobs=-1,
    random_state=0,
)
search.fit(df, target)
search.best_params_, search.best_score_
best = search.best_estimator_   # refit on all data
Attribute / paramMeaning
best_params_, best_score_winning combination and its mean CV score
best_estimator_refit on the full training data (refit=True)
cv_results_dict of arrays; pd.DataFrame(search.cv_results_)
scoring={"auc": ..., "f1": ...}several metrics; then refit="auc" picks the winner
error_score="raise"fail loudly instead of scoring broken fits as NaN
Param grid as a list of dictsdifferent grids per model or per step value

Metrics

Scorer strings follow "higher is better", so error metrics are negated: "neg_mean_absolute_error".

ClassificationUse when
accuracy_scorebalanced classes, equal error costs
balanced_accuracy_scoreimbalanced classes; mean of per-class recall
precision_scorefalse positives are costly (spam filter)
recall_scorefalse negatives are costly (disease screening)
f1_score(average=...)one number trading precision against recall; macro treats classes equally
roc_auc_scoreranking quality across thresholds; needs scores/probabilities
average_precision_scorePR AUC; better than ROC AUC for rare positives
log_loss, brier_score_lossquality of the probabilities themselves
matthews_corrcoefsingle balanced summary for binary imbalanced data
confusion_matrix, classification_reportper-class breakdown
RegressionUse when
mean_absolute_errorerror in target units, robust to outliers
root_mean_squared_errortarget units; punishes large errors
mean_squared_errorthe loss most models minimize
r2_scorefraction of variance explained (1 is perfect, 0 is "predict the mean")
mean_absolute_percentage_errorrelative error; breaks near zero targets
median_absolute_errortypical error, ignores outliers
ClusteringNeeds labels?Measures
silhouette_scorenocohesion vs separation, -1 to 1
davies_bouldin_scorenolower is better
calinski_harabasz_scorenohigher is better
adjusted_rand_scoreyesagreement with ground truth, chance-corrected
normalized_mutual_info_scoreyesshared information with ground truth

Class imbalance

LeverHow
Stratifystratify=y in splits, StratifiedKFold in CV
Honest metricbalanced accuracy, PR AUC, recall/precision per class; never plain accuracy
Reweight classesclass_weight="balanced" (linear, SVM, trees, HistGradientBoosting*)
Reweight rowsfit(X, y, sample_weight=w)
Move the thresholdTunedThresholdClassifierCV(est, scoring="f1") picks the cut-off by CV
Fixed business thresholdFixedThresholdClassifier(est, threshold=0.3)
Resampleimbalanced-learn's SMOTE etc. inside its own imblearn.pipeline.Pipeline, never before the split
from sklearn.model_selection import (
    TunedThresholdClassifierCV,
)
 
tuned = TunedThresholdClassifierCV(
    model.set_params(clf__class_weight="balanced"),
    scoring="balanced_accuracy",
    cv=5,
).fit(df, target)
tuned.best_threshold_, tuned.best_score_

Feature importance

MethodNotes
permutation_importance(est, X_val, y_val)model-agnostic: score drop when one column is shuffled; run on held-out data
feature_importances_ (trees)impurity-based (MDI); biased toward high-cardinality and continuous features, computed on training data
coef_ (linear)comparable only after scaling; correlated features share weight
PartialDependenceDisplay.from_estimatorhow predictions change with one or two features
SHAP (third party)per-prediction attributions
from sklearn.inspection import permutation_importance
 
r = permutation_importance(
    model, df, target,     # use a held-out set in practice
    n_repeats=10, scoring="roc_auc", random_state=0,
)
ranked = sorted(
    zip(df.columns, r.importances_mean, r.importances_std),
    key=lambda t: -t[1],
)
for name, mean, std in ranked:
    print(f"{name:8} {mean:.3f} ± {std:.3f}")

Permutation importance works on the raw input columns of a pipeline, so city is scored as one feature rather than per one-hot column. Correlated features hide each other's importance.

Probability calibration

A classifier is calibrated when "0.8" means right about 80% of the time. Trees, boosting, SVMs and naive Bayes often are not.

ToolDoes
CalibratedClassifierCV(est, method="sigmoid")Platt scaling; fine with little data
method="isotonic"non-parametric; needs 1000+ calibration samples
method="temperature"1.8+; one parameter, preserves ranking, good for multiclass
cv=5fits the base model and calibrator per fold
CalibratedClassifierCV(FrozenEstimator(fitted))calibrate an already fitted model on new data (replaces the removed cv="prefit")
CalibrationDisplay.from_estimatorreliability diagram
brier_score_loss, log_losscompare before/after
from sklearn.calibration import CalibratedClassifierCV
from sklearn.frozen import FrozenEstimator
from sklearn.svm import SVC
 
X_fit, X_cal, y_fit, y_cal = train_test_split(
    Xs, y, test_size=0.3, stratify=y, random_state=0
)
svc = SVC().fit(X_fit, y_fit)  # probability=True deprecated
calibrated = CalibratedClassifierCV(
    FrozenEstimator(svc), method="sigmoid"
).fit(X_cal, y_cal)
calibrated.predict_proba(Xs[:3])

Saving models

FormatLoad is safe?Notes
joblib.dump / joblib.loadnopickle underneath: loading runs arbitrary code; efficient with big arrays
picklenosame risk, same version constraints
skops.io.dump / loadyes, with an allow-listrefuses unknown types until you trust them explicitly
ONNX (skl2onnx)yesserve from other runtimes (ONNX Runtime, including the browser)
  • Only unpickle files you created or fully trust.
  • Pin the scikit-learn version: loading in another version raises InconsistentVersionWarning and may give wrong results. Save the training code and data version too.
import skops.io as sio
 
sio.dump(model, "model.skops")
unknown = sio.get_untrusted_types(file="model.skops")
print(unknown)               # review before trusting
loaded = sio.load("model.skops", trusted=unknown)
assert (loaded.predict(df) == model.predict(df)).all()

Output config & routing

FeatureAPINotes
DataFrame outputest.set_output(transform="pandas")also "polars"; global: set_config(transform_output="pandas")
Metadata routingset_config(enable_metadata_routing=True)pass sample_weight, groups through pipelines and searches
Request metadataest.set_fit_request(sample_weight=True)each consumer opts in; also set_score_request
Array APIset_config(array_api_dispatch=True)experimental; some estimators accept PyTorch/CuPy arrays and stay on GPU; needs array-api-compat and SCIPY_ARRAY_API=1
Callbacksest.set_callbacks(ProgressBar())experimental in 1.9 (sklearn.callback); few estimators support it
Sparse typeset_config(sparse_interface="sparray")1.9+; SciPy sparse arrays instead of matrices
import sklearn
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import cross_validate
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
 
pre.set_output(transform="pandas")
pre.fit_transform(df).head()     # named DataFrame columns
 
with sklearn.config_context(enable_metadata_routing=True):
    w = np.where(y == 1, 2.0, 1.0)
    scaler = StandardScaler().set_fit_request(
        sample_weight=False           # opt out explicitly
    )
    lr = (
        LogisticRegression()
        .set_fit_request(sample_weight=True)
        .set_score_request(sample_weight=True)
    )
    routed = make_pipeline(scaler, lr)
    routed.fit(X, y, sample_weight=w)  # reaches only LR
    cross_validate(
        routed, X, y, params={"sample_weight": w}
    )

Unrequested metadata raises UnsetMetadataPassedError, so a weight is never dropped silently.

Recipes

Tabular pipeline on mixed columns

When a DataFrame mixes numbers, categories and gaps: gradient boosting with native categorical support needs almost no preprocessing.

import numpy as np
import pandas as pd
from sklearn.ensemble import HistGradientBoostingClassifier
from sklearn.model_selection import train_test_split
 
rng = np.random.default_rng(1)
n = 1_000
X = pd.DataFrame({
    "age": rng.integers(18, 80, n).astype(float),
    "plan": pd.Categorical(rng.choice(["free", "pro"], n)),
    "visits": rng.poisson(5, n).astype(float),
})
X.loc[rng.random(n) < 0.1, "age"] = np.nan   # NaN is fine
rule = (X["plan"] == "pro") & (X["visits"] > 4)
y = (rule ^ (rng.random(n) < 0.05)).astype(int)  # noise
 
X_tr, X_te, y_tr, y_te = train_test_split(
    X, y, stratify=y, random_state=0
)
hgb = HistGradientBoostingClassifier(
    categorical_features="from_dtype",  # category dtype
    random_state=0,
).fit(X_tr, y_tr)
print(f"test accuracy {hgb.score(X_te, y_te):.3f}")

Cross-validate with several metrics

When one number hides trade-offs and you want train vs test scores to spot overfitting.

import pandas as pd
from sklearn.datasets import make_classification
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import cross_validate
 
X, y = make_classification(
    n_samples=600, weights=[0.8], random_state=0
)
res = cross_validate(
    RandomForestClassifier(random_state=0),
    X, y,
    cv=5,
    scoring=["roc_auc", "balanced_accuracy", "f1"],
    return_train_score=True,
    n_jobs=-1,
)
summary = pd.DataFrame(res).mean().round(3)
print(summary)   # fit_time, test_roc_auc, train_roc_auc...

Tune a gradient boosting model

When defaults are close but you want the last few points; successive halving discards bad candidates early.

from scipy.stats import loguniform, randint
from sklearn.datasets import make_classification
from sklearn.ensemble import HistGradientBoostingClassifier
from sklearn.experimental import (
    enable_halving_search_cv,  # noqa: F401 (side effect)
)
from sklearn.model_selection import HalvingRandomSearchCV
 
X, y = make_classification(n_samples=2_000, random_state=0)
space = {
    "learning_rate": loguniform(0.01, 0.3),
    "max_leaf_nodes": randint(8, 64),
    "min_samples_leaf": randint(5, 50),
    "l2_regularization": loguniform(1e-4, 1.0),
}
search = HalvingRandomSearchCV(
    HistGradientBoostingClassifier(random_state=0),
    space,
    factor=3,              # keep the best third each round
    scoring="roc_auc",
    random_state=0,
    n_jobs=-1,
).fit(X, y)
print(search.best_params_, round(search.best_score_, 3))

Confusion matrix and classification report

When you need per-class precision/recall on a held-out set, not a single score.

from sklearn.datasets import make_classification
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import (
    classification_report,
    confusion_matrix,
)
from sklearn.model_selection import train_test_split
 
X, y = make_classification(
    n_samples=800, n_classes=3, n_informative=4,
    random_state=0,
)
X_tr, X_te, y_tr, y_te = train_test_split(
    X, y, stratify=y, random_state=0
)
pred = LogisticRegression().fit(X_tr, y_tr).predict(X_te)
print(confusion_matrix(y_te, pred))  # rows true, cols pred
print(classification_report(y_te, pred, digits=3))
# plot: ConfusionMatrixDisplay.from_predictions(y_te, pred)

Persist and reload a pipeline

When a trained pipeline ships to another process; skops refuses types you haven't approved.

from pathlib import Path
 
import skops.io as sio
from sklearn.datasets import make_classification
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
 
X, y = make_classification(random_state=0)
pipe = make_pipeline(StandardScaler(), LogisticRegression())
pipe.fit(X, y)
 
path = Path("pipe.skops")
sio.dump(pipe, path)
# later, in the serving process:
untrusted = sio.get_untrusted_types(file=path)
assert untrusted == [], f"review: {untrusted}"
restored = sio.load(path)
assert (restored.predict(X) == pipe.predict(X)).all()

Custom transformer

When a feature step needs learned state and should work inside Pipeline, clone and grid search.

import numpy as np
from numpy.typing import ArrayLike, NDArray
from sklearn.base import BaseEstimator, TransformerMixin
from sklearn.utils.validation import (
    check_is_fitted,
    validate_data,
)
 
class QuantileClipper(TransformerMixin, BaseEstimator):
    """Clip each column to its fitted quantile range."""
 
    def __init__(self, q: float = 0.01) -> None:
        self.q = q          # store params as given, no logic
 
    def fit(self, X: ArrayLike, y: object = None):
        Xv = validate_data(self, X)   # sets n_features_in_
        self.low_ = np.quantile(Xv, self.q, axis=0)
        self.high_ = np.quantile(Xv, 1 - self.q, axis=0)
        return self
 
    def transform(self, X: ArrayLike) -> NDArray[np.float64]:
        check_is_fitted(self)
        Xv = validate_data(self, X, reset=False)
        return np.clip(Xv, self.low_, self.high_)

TransformerMixin supplies fit_transform. Pass-through column names come from OneToOneFeatureMixin, and sklearn.utils.estimator_checks.check_estimator tests the contract.

References