A scikit-learn Pipeline chains data-preparation steps and a model into a single object. It is the simplest way to avoid data leakage and to ship a model reliably.
Why Pipelines
- No leakage: each step is fitted only on training folds during cross-validation.
- Reproducibility: the exact same transformations run at training and prediction time.
- Deployment: save one object that takes raw inputs and returns predictions.
- Tuning: search over preprocessing and model hyperparameters together.
Mixed Column Types
A ColumnTransformer applies different steps to different columns:
from sklearn.compose import ColumnTransformer
from sklearn.pipeline import Pipeline, make_pipeline
from sklearn.impute import SimpleImputer
from sklearn.preprocessing import OneHotEncoder, StandardScaler
from sklearn.ensemble import HistGradientBoostingClassifier
prepare = ColumnTransformer([
("num", make_pipeline(SimpleImputer(strategy="median"), StandardScaler()), numeric_cols),
("cat", make_pipeline(SimpleImputer(strategy="most_frequent"),
OneHotEncoder(handle_unknown="ignore")), categorical_cols),
])
model = Pipeline([("prepare", prepare), ("clf", HistGradientBoostingClassifier())])
model.fit(X_train, y_train)
Tuning Through a Pipeline
Parameter names use double underscores: clf__learning_rate, prepare__num__simpleimputer__strategy.
Saving and Loading
Save the fitted pipeline with joblib.dump and load it where predictions are needed. Record library versions, because pickled objects can break across versions — and only load files you trust.
Custom Steps
Wrap your own feature logic in a FunctionTransformer or a small class with fit and transform, so it stays inside the pipeline too.