§ 3.2Module 3

K-Fold Cross-Validation

On this page

3.2 — K-Fold Cross-Validation

Recall first. Why is selecting a model using the training score optimistic? What does a validation fold contribute that training data alone cannot?

The procedure

In k-fold cross-validation, partition the training data into k folds. Repeat k times: hold one fold out for validation, fit on the other k−1, and score on the held-out fold. Average the scores:

CV_score = (1/k) Σₖ score(fitted on folds except k, fold k)

Every training example is used for validation once and fitting k−1 times. Cross-validation estimates performance for model comparison and hyperparameter selection; it does not replace an untouched final test set.1

Typical choices are k=5 or 10. Larger k uses more training data per fit but costs more computation and can make estimates more correlated; smaller k is cheaper but each fit has less training data. There is no universally best k.

Split design is part of correctness

Preprocessing must be fitted inside each training fold. A pipeline applies imputation, scaling, feature selection, and estimator fitting without allowing the validation fold to influence transformations.1

Worked trace

With 12 observations and k=3, folds are F₁,F₂,F₃, each containing four observations:

Mean CV score is (0.80+0.70+0.90)/3=0.80. The spread warns that performance varies by fold. After choosing the model, refit it on all available training data and evaluate once on the untouched test set.

Nested selection and leakage

If many hyperparameters are tuned against the same CV results, the CV estimate can become optimistic because it has guided selection. Nested CV uses inner folds for tuning and outer folds for estimating the full selection procedure, especially when a separate test set is unavailable. In ordinary coursework, the minimum safe answer is: keep a final test set untouched, and put all learned preprocessing inside the CV pipeline.1

Exercise

A medical dataset has five records per patient. The team randomly distributes rows into five folds. What is the problem and which split would you choose? Why might stratification alone not solve it?

Revealed answer

Rows from one patient can occur in both training and validation, so the model sees patient-specific information and the estimate is optimistic. Use group k-fold with patient ID as the group (and consider class balance separately if supported). Stratification preserves class proportions but does not prevent the same patient crossing folds.

Exam lens

Draw k rounds, state fit-on-k−1/validate-on-1, average scores, and reserve the test set. Add one sentence on stratification, grouped data, or time ordering to show evaluation awareness.

Rapid revision checklist

Key takeaways

Sources

Footnotes

  1. scikit-learn, Cross-validation User Guide and common pitfalls. ↩ ↩2 ↩3