Different Ways to Combine Classifiers
On this page
3.9 — Different Ways to Combine Classifiers
Recall first. You have three classifiers with different strengths. Would you combine their hard labels, their probabilities, or train another model on their outputs? What could go wrong with each choice?
Combination is a second model-design problem
Let classifiers output labels h₁(x),...,h_M(x) or scores/probabilities p_m(y|x). Common combination rules are:
- Hard majority voting: choose the class with most votes. Simple and robust, but ties need a rule and all votes receive equal weight.
- Weighted voting: choose the class with largest
Σw_m 1[h_m(x)=c], where weights are selected on validation data. Do not tune weights on the final test set. - Soft/probability voting: average or weighted-average class probabilities and choose the largest. It uses confidence information but assumes scores are comparable; calibration may be needed.
- Max rule: use the largest class score across classifiers; sensitive to badly scaled or overconfident scores.
- Stacking: train a meta-classifier on base predictions (and possibly original features). It can learn when each classifier is reliable.
- Cascades: apply a cheap model first and send uncertain cases to a more expensive model; useful when latency or cost matters.12
Stacking without leakage
The meta-model must not learn from base predictions generated on the same rows used to fit those base models. Otherwise it can exploit overfit predictions and report an unrealistic score. Use out-of-fold predictions: split training data, fit each base classifier on inner training folds, predict the held-out fold, and concatenate those predictions as meta-training features. After the design is fixed, refit base models on all training data and evaluate the complete stack on an untouched test set.1
Worked voting trace
Three classifiers produce:
hard labels: A, B, A → hard vote A
P(A): 0.55, 0.95, 0.60 → soft mean = 0.70 → A
If validation weights are [0.2,0.5,0.3], weighted P(A)=0.2(0.55)+0.5(0.95)+0.3(0.60)=0.755, still A. Suppose classifier B is systematically overconfident: the soft vote may be worse than hard voting unless probabilities are calibrated. Stacking could learn to discount it using appropriate out-of-fold evidence.
When combination helps
Combination helps when classifiers have complementary inductive biases or errors. If all use the same features and make the same mistakes, averaging cannot recover missing information. More models also mean more computation, more monitoring, and more opportunities for leakage. Always compare the ensemble against the best single classifier and a simple baseline on the same split.
Exercise
Four classifiers output [A,A,B,C]. Is there a unique hard-vote winner? If a validation procedure gives weights [0.1,0.2,0.5,0.2] to those classifiers, what are weighted votes for A, B, and C?
Revealed answer
Hard voting is a three-way tie: A, B, and C each have one or more? A has two votes, B one, C one, so hard winner is A. Weighted votes: A 0.1+0.2=0.3, B 0.5, C 0.2; weighted winner is B. (The weights should be learned only from validation/CV data and checked on held-out data.)
Exam lens
List hard vote, weighted vote, soft vote, max rule, stacking, and cascade. For stacking, emphasize out-of-fold meta-features to prevent leakage. State the accuracy–diversity and calibration trade-offs.
Rapid revision checklist
- Can I distinguish hard and soft voting?
- Can I calculate a weighted vote?
- Can I explain stacking in one diagram?
- Can I explain why in-sample base predictions leak?
- Can I state when combining classifiers may not help?
Key takeaways
- Classifiers can be combined at the label, score, probability, or learned meta-model level.
- Soft voting uses confidence but depends on calibration and comparable scales.
- Stacking is powerful only when meta-training predictions are out-of-fold.
- Diversity and honest validation determine whether combination improves generalization.