K-Fold Cross-Validation
On this page
3.2 — K-Fold Cross-Validation
Recall first. Why is selecting a model using the training score optimistic? What does a validation fold contribute that training data alone cannot?
The procedure
In k-fold cross-validation, partition the training data into k folds. Repeat k times: hold one fold out for validation, fit on the other k−1, and score on the held-out fold. Average the scores:
CV_score = (1/k) Σₖ score(fitted on folds except k, fold k)
Every training example is used for validation once and fitting k−1 times. Cross-validation estimates performance for model comparison and hyperparameter selection; it does not replace an untouched final test set.1
Typical choices are k=5 or 10. Larger k uses more training data per fit but costs more computation and can make estimates more correlated; smaller k is cheaper but each fit has less training data. There is no universally best k.
Split design is part of correctness
- Stratified k-fold: approximately preserves class proportions in each fold; useful for classification.
- Group k-fold: keeps all observations from one person, machine, or patient in one fold to avoid entity leakage.
- Time-series split: respects temporal order; never train on future to validate the past.
- Repeated k-fold: repeats random partitions to assess stability.
Preprocessing must be fitted inside each training fold. A pipeline applies imputation, scaling, feature selection, and estimator fitting without allowing the validation fold to influence transformations.1
Worked trace
With 12 observations and k=3, folds are F₁,F₂,F₃, each containing four observations:
- Round 1: fit on
F₂+F₃, validate onF₁; score 0.80. - Round 2: fit on
F₁+F₃, validate onF₂; score 0.70. - Round 3: fit on
F₁+F₂, validate onF₃; score 0.90.
Mean CV score is (0.80+0.70+0.90)/3=0.80. The spread warns that performance varies by fold. After choosing the model, refit it on all available training data and evaluate once on the untouched test set.
Nested selection and leakage
If many hyperparameters are tuned against the same CV results, the CV estimate can become optimistic because it has guided selection. Nested CV uses inner folds for tuning and outer folds for estimating the full selection procedure, especially when a separate test set is unavailable. In ordinary coursework, the minimum safe answer is: keep a final test set untouched, and put all learned preprocessing inside the CV pipeline.1
Exercise
A medical dataset has five records per patient. The team randomly distributes rows into five folds. What is the problem and which split would you choose? Why might stratification alone not solve it?
Revealed answer
Rows from one patient can occur in both training and validation, so the model sees patient-specific information and the estimate is optimistic. Use group k-fold with patient ID as the group (and consider class balance separately if supported). Stratification preserves class proportions but does not prevent the same patient crossing folds.
Exam lens
Draw k rounds, state fit-on-k−1/validate-on-1, average scores, and reserve the test set. Add one sentence on stratification, grouped data, or time ordering to show evaluation awareness.
Rapid revision checklist
- Can I describe every fold’s training and validation sets?
- Can I calculate an average CV score?
- Can I choose stratified, group, or time-aware splitting?
- Can I explain pipeline leakage?
- Can I distinguish CV validation from final test evaluation?
Key takeaways
- k-fold CV reuses training data efficiently for validation and model selection.
- The split must match dependence, class balance, and time structure.
- Learned preprocessing belongs inside each training fold.
- A final untouched test set is still needed for an honest final estimate.
Sources
Footnotes
-
scikit-learn, Cross-validation User Guide and common pitfalls. ↩ ↩2 ↩3