§ 3.9Module 3

Different Ways to Combine Classifiers

On this page

3.9 — Different Ways to Combine Classifiers

Recall first. You have three classifiers with different strengths. Would you combine their hard labels, their probabilities, or train another model on their outputs? What could go wrong with each choice?

Combination is a second model-design problem

Let classifiers output labels h₁(x),...,h_M(x) or scores/probabilities p_m(y|x). Common combination rules are:

  1. Hard majority voting: choose the class with most votes. Simple and robust, but ties need a rule and all votes receive equal weight.
  2. Weighted voting: choose the class with largest Σw_m 1[h_m(x)=c], where weights are selected on validation data. Do not tune weights on the final test set.
  3. Soft/probability voting: average or weighted-average class probabilities and choose the largest. It uses confidence information but assumes scores are comparable; calibration may be needed.
  4. Max rule: use the largest class score across classifiers; sensitive to badly scaled or overconfident scores.
  5. Stacking: train a meta-classifier on base predictions (and possibly original features). It can learn when each classifier is reliable.
  6. Cascades: apply a cheap model first and send uncertain cases to a more expensive model; useful when latency or cost matters.12

Stacking without leakage

The meta-model must not learn from base predictions generated on the same rows used to fit those base models. Otherwise it can exploit overfit predictions and report an unrealistic score. Use out-of-fold predictions: split training data, fit each base classifier on inner training folds, predict the held-out fold, and concatenate those predictions as meta-training features. After the design is fixed, refit base models on all training data and evaluate the complete stack on an untouched test set.1

Worked voting trace

Three classifiers produce:

hard labels: A, B, A  → hard vote A
P(A):        0.55, 0.95, 0.60 → soft mean = 0.70 → A

If validation weights are [0.2,0.5,0.3], weighted P(A)=0.2(0.55)+0.5(0.95)+0.3(0.60)=0.755, still A. Suppose classifier B is systematically overconfident: the soft vote may be worse than hard voting unless probabilities are calibrated. Stacking could learn to discount it using appropriate out-of-fold evidence.

When combination helps

Combination helps when classifiers have complementary inductive biases or errors. If all use the same features and make the same mistakes, averaging cannot recover missing information. More models also mean more computation, more monitoring, and more opportunities for leakage. Always compare the ensemble against the best single classifier and a simple baseline on the same split.

Exercise

Four classifiers output [A,A,B,C]. Is there a unique hard-vote winner? If a validation procedure gives weights [0.1,0.2,0.5,0.2] to those classifiers, what are weighted votes for A, B, and C?

Revealed answer

Hard voting is a three-way tie: A, B, and C each have one or more? A has two votes, B one, C one, so hard winner is A. Weighted votes: A 0.1+0.2=0.3, B 0.5, C 0.2; weighted winner is B. (The weights should be learned only from validation/CV data and checked on held-out data.)

Exam lens

List hard vote, weighted vote, soft vote, max rule, stacking, and cascade. For stacking, emphasize out-of-fold meta-features to prevent leakage. State the accuracy–diversity and calibration trade-offs.

Rapid revision checklist

Key takeaways

Sources

Footnotes

  1. scikit-learn, Voting Classifier and Stacking. ↩ ↩2

  2. Wolpert, “Stacked Generalization,” Neural Networks 5 (1992); Hastie, Tibshirani & Friedman, ESL, chapters 8–10. ↩