Hyperparameters can make a real difference to model quality. Tuning them systematically beats guessing.
Methods
- Grid search: try every combination of a few values per hyperparameter. Exhaustive but grows quickly.
- Random search: sample combinations at random. For the same budget it usually finds better settings than a grid, because only a few hyperparameters tend to matter.
- Bayesian optimisation: builds a model of how settings affect performance and picks promising ones next. Libraries such as Optuna make this easy.
- Successive halving / Hyperband: try many settings cheaply, then give more resources to the best.
What to Tune
Focus on the few that matter most:
- Gradient boosting: learning rate, number of trees (via early stopping), depth or leaves, subsampling.
- Random forest: number of trees, max features, min samples per leaf.
- Neural networks: learning rate, batch size, architecture size, regularisation.
- Linear models: regularisation strength.
Doing It Honestly
- Tune with cross-validation on the training data.
- Keep the test set untouched until the final evaluation.
- Search on a log scale for rates and regularisation strengths (0.001, 0.01, 0.1…).
- Beware overfitting the validation set when running very many trials; confirm on the test set once.
Know When to Stop
Better features or more data usually beat extra tuning. If improvements are within the noise between cross-validation folds, stop.
from sklearn.model_selection import RandomizedSearchCV
search = RandomizedSearchCV(model, param_distributions, n_iter=50, cv=5, random_state=0).fit(X, y)