IAT 1 Question Bank
On this page
Machine Learning Question Bank: Exam-Ready Answers
These structured answers cover all 31 questions in the IAT-1 question bank. In an examination, an answer can be expanded or shortened according to the available time; retain the definition, key formula or algorithm, example, and conclusion.
Module I: Foundations of Machine Learning
1. Define Machine Learning. Explain the different types of Machine Learning with suitable examples.
Machine Learning (ML) is a branch of Artificial Intelligence in which a computer system learns patterns from data and uses those patterns to make predictions or decisions, instead of being programmed with a separate rule for every possible situation. A useful formal definition is: a program learns from experience E with respect to a task T and performance measure P if its performance on T, measured by P, improves with E. For example, an email classifier learns from previously labelled emails and improves its spam-detection accuracy as it receives more training examples.
| Type | Training information | Main objective | Example |
|---|---|---|---|
| Supervised learning | Input features and correct target labels | Learn a mapping from inputs to known outputs | Predict house price or classify an email as spam/not spam |
| Unsupervised learning | Input features without target labels | Discover hidden structure or groups | Customer segmentation using clustering |
| Reinforcement learning | Rewards or penalties obtained from actions | Learn a policy that maximises long-term reward | Robot navigation or game playing |
| Semi-supervised learning | A small labelled set and a large unlabelled set | Use both types of data to improve prediction | Classifying medical images when only some are labelled |
In supervised learning, the dataset contains pairs (xᵢ, yᵢ). Regression predicts a continuous value, such as temperature or salary, while classification predicts a class, such as disease/no disease. Linear regression, decision trees, and neural networks can be used in this category.
In unsupervised learning, only xᵢ is available. Algorithms such as K-Means clustering, hierarchical clustering, and Principal Component Analysis identify groups, lower-dimensional representations, or unusual observations. For example, an online store can group customers according to purchasing behaviour without being given the groups in advance.
In reinforcement learning, an agent interacts with an environment. At each step, it observes a state, takes an action, receives a reward, and moves to a new state. The agent gradually learns a policy π(a|s) that chooses useful actions. A robot receiving positive reward for reaching a destination is a simple example.
Thus, the choice of learning type depends on the available feedback: labelled outcomes suggest supervised learning, unlabelled data suggests unsupervised learning, and sequential decisions with rewards suggest reinforcement learning.
2. Explain the steps involved in developing a Machine Learning application using a suitable example.
Developing an ML application is an iterative process rather than simply choosing an algorithm. The important steps are given below, using email spam classification as an example.
- Define the problem and success criterion. State exactly what must be predicted and how success will be measured. In spam filtering, the task is binary classification: spam or legitimate email. Because wrongly blocking a genuine email is costly, precision, recall, and false-positive rate may be more important than accuracy alone.
- Collect relevant data. Gather a sufficiently large and representative set of emails. Each record may contain the subject, message body, sender information, links, attachments, and a trusted label.
- Understand and clean the data. Inspect class balance, missing values, duplicates, corrupted records, and possible data leakage. Duplicate emails should not be placed in both training and test sets because this can make the test score unrealistically high.
- Pre-process and engineer features. Convert raw text into numerical features using TF-IDF, word n-grams, or embeddings. Remove or normalize unsuitable content, and encode categorical variables such as sender domain when appropriate.
- Split the data. Create training, validation, and test sets. The training set fits the model, the validation set helps choose hyperparameters, and the untouched test set provides the final estimate of generalisation performance. Stratified splitting is useful when spam and non-spam examples are imbalanced.
- Select and train a baseline model. Start with a majority-class classifier or logistic regression. Then train more advanced models if they are justified. The training process minimises a loss function, such as logistic loss.
- Tune and validate. Use cross-validation, threshold tuning, and hyperparameter search. For example, the regularisation strength of logistic regression and the decision threshold can be selected using validation data.
- Evaluate and interpret. Use a confusion matrix, precision, recall, F1-score, and possibly ROC-AUC or precision-recall AUC. Examine errors to learn whether newsletters, phishing links, or multilingual messages cause problems.
- Deploy and monitor. Integrate the model into the mail server or application. Record prediction quality, latency, data drift, and changes in spam behaviour.
- Maintain responsibly. Protect private email content, control access to data, document limitations, and provide a mechanism to correct false decisions. The process repeats as new labelled data becomes available.
3. Differentiate between supervised, unsupervised, and reinforcement learning with real-world applications.
The three learning paradigms differ mainly in the kind of feedback available to the algorithm and the type of problem being solved.
| Aspect | Supervised learning | Unsupervised learning | Reinforcement learning |
|---|---|---|---|
| Data | Labelled examples | Unlabelled observations | State, action, reward sequences |
| Feedback | Correct answer is supplied during training | No explicit correct answer | Delayed reward or penalty |
| Main goal | Predict a target for new inputs | Discover structure or representation | Learn a policy for sequential decisions |
| Typical tasks | Classification and regression | Clustering, dimensionality reduction, anomaly detection | Control, planning, games, robotics |
| Example | Predict whether a transaction is fraudulent | Group customers by behaviour | Teach a robot to pick an object |
In supervised learning, a model minimises a loss between its prediction and the known target. A bank may train a classifier using past loan applications labelled as default or non-default. A regression model may predict electricity demand from weather and calendar features.
In unsupervised learning, the algorithm finds useful regularities without labels. A retailer may use K-Means to group customers into purchasing segments. Anomaly-detection methods may identify unusual network activity by learning the normal pattern. Since no target is provided, quality must be judged using clustering measures, reconstruction error, visual inspection, or business usefulness.
In reinforcement learning, an agent repeatedly interacts with an environment and learns from consequences rather than from a list of correct actions. The agent seeks to maximise the expected discounted return:
Gₜ = Σₖ₌₀^∞ γᵏ rₜ₊ₖ₊₁
where γ controls the importance of future rewards. A warehouse robot may receive a reward for placing an item correctly and a penalty for collisions. Supervised learning normally treats examples as independent cases, whereas reinforcement learning involves exploration, delayed consequences, and the exploration–exploitation trade-off.
4. Explain overfitting and underfitting. How can these problems be reduced in Machine Learning models?
Overfitting occurs when a model learns the training data too closely, including noise, random fluctuations, and accidental relationships. It achieves very low training error but performs poorly on unseen data. A very deep decision tree that memorises individual training records is a common example.
Underfitting occurs when a model is too simple to represent the important relationship in the data. It has high training error and also high validation or test error. A straight-line model used for a strongly curved relationship may underfit.
| Condition | Training error | Validation/test error | Main cause |
|---|---|---|---|
| Underfitting | High | High | Model too simple, weak features, excessive regularisation |
| Good fit | Low | Low and similar to training error | Suitable complexity and useful features |
| Overfitting | Very low | High | Model too complex, noisy data, leakage, or insufficient data |
Overfitting can be reduced by collecting more representative training data, removing duplicates and leakage, reducing model complexity, and applying regularisation. Examples include limiting a tree’s max_depth, increasing min_samples_leaf, using Ridge or Lasso penalties, reducing neural-network parameters, and applying dropout. Cross-validation helps select an appropriate complexity level. Early stopping stops training when validation performance stops improving.
Underfitting can be reduced by using a more expressive model, adding informative features, reducing excessive regularisation, training for longer when optimisation has not converged, and allowing a deeper tree or a higher-degree feature representation. Increasing complexity should be guided by validation performance rather than training accuracy alone.
A reliable workflow is to plot learning curves, compare training and validation errors, and check whether the model has seen enough data. Overfitting is a problem of high variance, while underfitting is commonly associated with high bias.
5. Compare training error and generalisation error. Why is generalisation more important than training accuracy?
Training error is the average loss obtained on the observations used to fit the model. If D_train = {(xᵢ,yᵢ)}ᵐᵢ₌₁, the empirical training error is
R̂_train(f) = (1/m) Σᵢ L(f(xᵢ), yᵢ)
where L is a loss function such as squared error or zero-one classification loss. Generalisation error is the expected loss on new examples drawn from the real data-generating distribution:
R(f) = E_(X,Y)~P [L(f(X), Y)]
It is estimated using a validation set, test set, or cross-validation. The difference between training performance and unseen-data performance is the generalisation gap.
| Feature | Training error | Generalisation error |
|---|---|---|
| Data used | Examples used for fitting | New, unseen examples |
| What it measures | Fit to known data | Ability to transfer learned patterns |
| Effect of complexity | Usually decreases as complexity increases | May decrease and then increase |
| Main risk | Memorisation can look successful | Reveals overfitting and real-world failure |
Training accuracy is not sufficient because a model can memorise the training records. For example, a decision tree grown until every leaf contains one training example may obtain 100% training accuracy, but its rules may not work for new patients or customers. A high training score can also be caused by data leakage, where target or test-set information accidentally enters the features.
Generalisation is more important because the purpose of ML is to make reliable predictions on future or unseen cases. A model with 94% training accuracy and 90% test accuracy is often more useful than a model with 99% training accuracy and 70% test accuracy. To estimate generalisation honestly, use a representative hold-out test set, perform preprocessing inside a cross-validation pipeline, avoid repeated tuning on the test set, and report uncertainty where possible.
6. Explain the Bias–Variance Trade-off with a neat diagram.
Bias is the error caused by overly simple assumptions in the learning algorithm. A high-bias model systematically misses important relationships and tends to underfit. Variance is the error caused by sensitivity to the particular training sample. A high-variance model changes substantially when the training data changes and tends to overfit.
For a regression problem, the expected prediction error at a point can be expressed as
Expected squared error = Bias² + Variance + irreducible noise (σ²)
The irreducible noise is caused by randomness or missing information and cannot be eliminated completely by choosing a more complex model.
Expected
test error
^ total error
| /\
| / \
| bias² / \ variance
| \ / \
| \______________/ \
| noise floor ----------------
+--------------------------------------> Model complexity
too simple best too complex
underfitting overfitting
When model complexity increases from a very low level, bias usually decreases because the model can represent more patterns. At the same time, variance usually increases because the model can react to noise. The best model lies near the minimum validation error, where the combined bias and variance are acceptably small.
The trade-off can be managed by selecting model complexity with cross-validation, using regularisation, pruning trees, collecting more data, reducing noisy features, and using ensemble methods. Bagging mainly reduces variance by averaging unstable models, while boosting often reduces bias by adding learners that correct previous errors. The objective is not the lowest training error; it is the best expected performance on unseen data.
7. A model achieves 99% training accuracy but only 72% testing accuracy. Analyze the problem and suggest suitable remedies.
The large gap between 99% training accuracy and 72% testing accuracy is a strong sign of overfitting or a problem in the evaluation pipeline. The generalisation gap is 27 percentage points. The model has learned the training examples very well, but much of what it learned does not transfer to unseen examples.
First check whether the test set is representative and whether the split was performed correctly. If the test set contains a different population, time period, device type, or class distribution, the issue may be distribution shift rather than ordinary overfitting. Check for duplicate or near-duplicate records across the sets, and check whether a small or noisy test set gives an unstable estimate.
Suitable remedies are:
- Check data leakage. Remove future information, target-derived values, and identifiers that allow the model to memorise records. Fit imputers, encoders, and scalers only on training folds.
- Use a reliable split. Apply stratified splitting for classification, group splitting when records from the same person or device are related, and time-based splitting for forecasting.
- Reduce model complexity. Use a shallower tree, fewer parameters, fewer polynomial terms, or stronger regularisation. For trees, control
max_depth,min_samples_split, andmin_samples_leaf. - Tune with cross-validation. Do not tune on the final test set. Select the simplest model whose validation performance is competitive.
- Improve the data. Collect more labelled examples, remove incorrect labels and duplicates, and add representative cases from under-covered classes.
- Use feature selection or dimensionality reduction. Remove irrelevant and highly noisy variables while preserving meaningful features.
- Consider ensembling. Bagging or a regularised ensemble may reduce variance, but ensembling cannot replace correction of leakage or poor data quality.
- Evaluate suitable metrics. If classes are imbalanced, accuracy can be misleading; inspect precision, recall, F1-score, ROC-AUC, and the confusion matrix.
Finally, compare training, cross-validation, and test learning curves. If both training and validation accuracy are low, the issue may be underfitting or weak features. If training remains high while validation remains low, stronger variance control is needed.
8. Discuss any five real-world applications of Machine Learning and justify why ML is suitable for each application.
1. Medical diagnosis and risk prediction
ML can analyse symptoms, laboratory values, medical images, and patient history to estimate the probability of a disease or adverse event. It is suitable because medical data is high-dimensional and may contain complex interactions that are difficult to express using fixed rules. A supervised classifier can learn from labelled cases, while anomaly detection can highlight unusual scans. The model should support, not replace, clinicians; sensitivity, calibration, interpretability, and fairness are important.
2. Fraud detection
Banks and payment systems process millions of transactions with features such as amount, location, time, merchant, and device. ML is suitable because fraud patterns change and are difficult to encode manually. A classifier or anomaly-detection model can score transactions in real time. Since fraud is usually rare, precision-recall analysis, cost-sensitive learning, and careful treatment of false positives are needed.
3. Recommendation systems
Streaming services and online stores recommend products, films, or songs using user history, item attributes, and behaviour of similar users. ML is suitable because preferences are individual, dynamic, and too numerous for hand-written rules. Collaborative filtering, matrix factorisation, and neural ranking models can learn latent preferences. Evaluation should include ranking quality, diversity, novelty, and business outcomes.
4. Speech recognition and natural-language processing
Voice assistants convert speech to text, and language models classify sentiment, translate text, or answer questions. ML is suitable because language contains variations in accent, vocabulary, context, and syntax that are difficult to program exhaustively. Models learn statistical relationships from large speech and text corpora. Accuracy should be checked across languages, accents, and user groups.
5. Predictive maintenance
Sensors from machines record vibration, temperature, pressure, and operating conditions. ML can predict probable failure or estimate remaining useful life. It is suitable because failure signatures may be nonlinear and may develop gradually. A model can help schedule maintenance before breakdown, reducing downtime and cost. Time-aware validation and anomaly detection are essential because random splitting can leak information from the future.
Across these applications, ML is most useful when sufficient historical data exists, the target or feedback can be defined, the environment is reasonably stable, and predictions can improve a real decision. Data quality, privacy, bias, monitoring, and human oversight remain necessary even when model accuracy is high.
9. Describe the process of data pre-processing. Why is it important before model training?
Data pre-processing converts raw data into a form that a learning algorithm can use reliably. It is important because algorithms learn from the supplied representation; missing values, inconsistent units, outliers, duplicated examples, and poorly encoded categories can produce misleading patterns and unstable models.
- Data inspection and profiling: understand rows and columns, data types, target distribution, ranges, unique values, missingness, duplicates, and possible leakage. Histograms, box plots, and correlation plots help identify problems.
- Data cleaning: correct invalid values, remove exact duplicates, standardise units, resolve inconsistent category names, and decide how to handle outliers. Outliers should not be deleted automatically because some are genuine and important.
- Handling missing values: remove rows only when missingness is rare and random; otherwise use training-only median imputation for numerical values or the most frequent/“unknown” category for categorical values. Add missingness indicators when informative.
- Encoding categorical variables: use one-hot encoding for nominal categories, ordinal encoding only when order is meaningful, and suitable leakage-controlled target or learned embeddings for high-cardinality features.
- Scaling numerical features: standardisation gives approximately zero mean and unit variance; min-max scaling maps values to a fixed interval. Scaling is important for distance-based models, gradient optimisation, and regularised linear models. Tree-based models generally do not require it.
- Feature engineering and selection: create meaningful variables, such as age from date of birth or rolling averages for time series. Remove irrelevant, redundant, or leakage-prone features.
- Splitting and transformation discipline: split before learning preprocessing statistics. A pipeline should fit the imputer, scaler, and encoder only on the training portion and then apply them to validation and test data.
- Class imbalance treatment: use stratified splits, class weights, suitable resampling within training folds, or threshold adjustment. Never oversample the test set.
Good pre-processing improves data quality, convergence, fairness, reproducibility, and generalisation. It also makes the final model easier to deploy because the exact transformation steps are documented.
10. Explain the major challenges faced while developing Machine Learning applications and propose suitable solutions.
ML projects face challenges throughout the data, modelling, deployment, and governance lifecycle.
| Challenge | Why it matters | Suitable solution |
|---|---|---|
| Poor or insufficient data | The model learns unreliable patterns | Define data requirements, clean labels, collect representative data, and use quality checks |
| Missing values and inconsistent formats | Algorithms may fail or learn spurious relationships | Use validated pipelines, documented imputation, and consistent schemas |
| Class imbalance | Accuracy may hide poor minority-class performance | Use stratified validation, class weights, resampling, threshold tuning, and precision-recall metrics |
| Overfitting and underfitting | The model may fail on new data or fail to learn the task | Use cross-validation, regularisation, feature engineering, and complexity control |
| Data leakage | Test information makes performance look unrealistically high | Split early and fit learned transformations inside the training pipeline |
| Distribution drift | Future data may differ from historical data | Monitor feature and label drift, use time-aware validation, and retrain when needed |
| Interpretability and trust | Users may not accept unexplained decisions | Use interpretable models where possible and provide explanations and calibration |
| Bias and fairness | Historical data may reproduce unequal treatment | Audit subgroup metrics, improve representation, use fairness-aware methods, and retain human review |
| Privacy and security | Personal or sensitive data can be exposed | Minimise data, control access, encrypt, anonymise where appropriate, and follow policy |
| Deployment and scalability | A good notebook model may be slow or fragile in production | Package preprocessing and model together, test latency, version artifacts, and monitor |
| Reproducibility | Results may change across runs or environments | Track code, data versions, random seeds, parameters, and experiments |
A practical solution is an end-to-end ML lifecycle: define the decision, build a representative dataset, create a reproducible pipeline, validate against the real deployment setting, test safety and fairness, deploy gradually, and monitor performance. Technical accuracy alone does not guarantee a successful application; the model must also be useful, reliable, maintainable, and acceptable to its users.
Module II: Regression, Classification, and Decision Trees
11. Explain Linear Regression with its mathematical model and assumptions.
Linear Regression models the relationship between a continuous dependent variable and one or more input variables. In simple linear regression, one feature x is used:
yᵢ = β₀ + β₁xᵢ + εᵢ
ŷᵢ = β₀ + β₁xᵢ
Here, yᵢ is the observed target, β₀ is the intercept, β₁ is the slope, and εᵢ is the error term. For p features, multiple linear regression is
ŷᵢ = β₀ + β₁xᵢ₁ + ··· + βₚxᵢₚ = xᵢᵀβ.
The ordinary least-squares method chooses coefficients that minimise the sum of squared errors:
J(β) = Σᵢ (yᵢ − ŷᵢ)²
If X includes a column of ones and the inverse exists, the least-squares solution is
β̂ = (XᵀX)⁻¹Xᵀy.
In practice, numerical solvers such as QR decomposition or gradient descent are preferred to directly computing an inverse. The coefficient βⱼ represents the expected change in the target for a one-unit increase in feature xⱼ, holding the other features constant.
Main assumptions
- Linearity: the conditional mean of the target is approximately a linear combination of the predictors.
- Independence: observations or error terms are independent, especially in time-series or repeated-measurement settings.
- Zero conditional mean:
E[ε|X] = 0, meaning important systematic predictors have not been omitted. - Homoscedasticity: error variance is approximately constant across fitted values. A funnel-shaped residual plot suggests a violation.
- Low multicollinearity: predictors should not be nearly exact linear combinations of each other; otherwise coefficients become unstable.
- Normality of errors for classical inference: normal errors are mainly needed for reliable small-sample confidence intervals and hypothesis tests, not for obtaining least-squares predictions.
Model quality can be evaluated using MAE, MSE, RMSE, and R², together with residual plots. Feature scaling is not required for ordinary least squares itself, but it can help optimisation and coefficient comparison. Linear regression is simple, fast, interpretable, and a useful baseline, but it may be inadequate for strong nonlinear relationships or complex interactions.
12. Differentiate between Linear Regression, Multiple Linear Regression, and Logistic Regression.
| Aspect | Simple linear regression | Multiple linear regression | Logistic regression |
|---|---|---|---|
| Number of inputs | One feature | Two or more features | One or more features |
| Target | Continuous | Continuous | Usually a binary class |
| Output | Any real value | Any real value | Probability between 0 and 1, then a class label |
| Link/function | Identity | Identity | Sigmoid/logistic function |
| Typical loss | Squared-error loss | Squared-error loss | Log loss/cross-entropy |
| Example | Salary from experience | House price from area, rooms, and age | Default or no default |
In simple linear regression, a line is fitted to describe the average relationship between one predictor and a continuous target. Multiple linear regression extends the same idea to several predictors. Its coefficients measure the partial effect of each feature while holding the other included features fixed.
Logistic regression first computes a linear score
z = β₀ + βᵀx
and transforms it using the sigmoid function:
σ(z) = 1/(1 + e⁻ᶻ).
The output is the estimated probability of class 1. A threshold, often 0.5 but selected according to application costs, converts the probability into a class prediction. Logistic regression is trained by minimising log loss, not ordinary squared error. Its decision boundary is linear in the input features unless nonlinear features or interactions are added. The word linear refers to the linear decision score or log-odds:
log(p/(1−p)) = β₀ + Σⱼ βⱼxⱼ.
Therefore, linear regression is used for continuous quantities, multiple linear regression is the multi-feature version of linear regression, and logistic regression is used for probabilistic classification. Ridge or Lasso regularisation can be applied to all of them to improve generalisation.
13. Explain the construction of a Decision Tree using the Gini Index.
A Decision Tree is a supervised model that recursively divides the feature space into increasingly pure groups. Internal nodes contain tests such as “Weather = Sunny,” branches represent outcomes of the test, and leaves contain the predicted class.
For a node containing observations from K classes, let pₖ be the fraction belonging to class k. The Gini impurity is
Gini(S) = 1 − Σₖ₌₁ᴷ pₖ².
A pure node has Gini impurity 0. For a candidate split of S into child subsets S₁,…,Sᵣ, the weighted impurity is
Gini_split = Σⱼ (|Sⱼ|/|S|) Gini(Sⱼ)
Gain = Gini(S) − Gini_split.
Construction procedure
- Place all training records at the root and calculate the root Gini impurity.
- Generate candidate tests for each feature. For a categorical feature, candidates may be category-based; for a numerical feature, candidates are thresholds between sorted values.
- Calculate weighted Gini impurity for every candidate.
- Select the candidate with the lowest value and divide the data.
- Repeat the process recursively for each non-pure child node.
- Stop when a node is pure, no useful split remains, or a rule such as maximum depth, minimum samples per split, or minimum impurity decrease is reached.
- Optionally apply post-pruning to remove branches that do not improve validation performance.
For example, if a node contains six records with four “Yes” and two “No,” then
Gini = 1 − (4/6)² − (2/6)² = 0.4444.
CART generally constructs binary trees and uses a greedy split search. The greedy method is efficient, although it does not guarantee the globally optimal tree. Unrestricted growth can overfit, so pruning and hyperparameter controls are important.
14. Explain CART algorithm with suitable examples.
CART means Classification and Regression Trees. It is a decision-tree algorithm that solves both classification and regression problems. CART builds a binary tree, so every internal node is divided into two child nodes. It chooses the split that gives the greatest reduction in impurity or loss.
For classification, CART commonly uses the Gini Index. For regression, it commonly minimises weighted mean squared error. If the target values in a node are y₁,…,yₙ, the node prediction is their mean ȳ, and the node impurity is
MSE(S) = (1/n) Σᵢ (yᵢ − ȳ)².
CART procedure
- Start with all observations at the root.
- For every feature and possible threshold, form a binary split such as
xⱼ ≤ tandxⱼ > t. - Compute the weighted child impurity.
- Select the split with the largest impurity reduction.
- Recursively apply the procedure to each child.
- Stop using conditions such as maximum depth, minimum samples per leaf, or insufficient impurity decrease.
- Prune the tree if a smaller tree gives better validation performance.
Classification example: A bank wants to classify a loan as safe or risky. CART may first split on income ≤ 50,000. The left child may then split on credit_score ≤ 650, while the right child may split on existing debt. A leaf predicts the majority class.
Regression example: If the target is house price, CART may split on area ≤ 1,000 square feet. Each leaf predicts the average price of houses reaching that leaf. A further split may use the number of bedrooms or location score.
CART handles nonlinear relationships and feature interactions without requiring a linear equation. Its advantages are interpretability and little need for scaling. Its limitations include instability, piecewise-constant predictions, and overfitting when the tree is too deep. Random forests and boosting reduce some of this instability by combining multiple trees.
15. Calculate Accuracy, Precision, Recall, Sensitivity, Specificity, and F1-score.
The given confusion matrix is:
| Actual / Predicted | Positive | Negative | Total |
|---|---|---|---|
| Positive | TP = 50 | FN = 10 | 60 |
| Negative | FP = 5 | TN = 35 | 40 |
| Total | 55 | 45 | 100 |
The formulas are:
Accuracy = (TP + TN)/(TP + TN + FP + FN)
Precision = TP/(TP + FP)
Recall = TP/(TP + FN)
Sensitivity = Recall
Specificity = TN/(TN + FP)
F1 = 2(Precision)(Recall)/(Precision + Recall)
| Metric | Calculation | Result |
|---|---|---|
| Accuracy | (50+35)/100 | 0.85 = 85% |
| Precision | 50/(50+5) | 0.9091 = 90.91% |
| Recall | 50/(50+10) | 0.8333 = 83.33% |
| Sensitivity | Same as recall | 0.8333 = 83.33% |
| Specificity | 35/(35+5) | 0.875 = 87.50% |
| F1-score | 2(0.9091)(0.8333)/(0.9091+0.8333) | 0.8696 = 86.96% |
The classifier correctly predicts 85% of all observations. When it predicts “positive,” it is correct about 90.91% of the time. It detects 83.33% of actual positive cases and correctly rejects 87.50% of actual negative cases. The F1-score gives a balanced summary of precision and recall. Accuracy may be insufficient when classes are highly imbalanced; the importance of sensitivity, specificity, and precision depends on the application’s error costs.
16. Explain ROC Curve and discuss its significance in evaluating classification models.
A Receiver Operating Characteristic (ROC) curve evaluates a binary classifier over many decision thresholds. The horizontal axis is the False Positive Rate (FPR) and the vertical axis is the True Positive Rate (TPR), also called recall or sensitivity:
TPR = TP/(TP+FN)
FPR = FP/(FP+TN) = 1 − Specificity
Many classifiers produce a probability or decision score rather than only a class label. At a high threshold, few cases are called positive, so both TPR and FPR may be low. As the threshold decreases, more cases are called positive, increasing both TPR and FPR. Plotting the pairs (FPR, TPR) creates the ROC curve.
True Positive Rate
1.0 | near ideal
| ********
| **
| **
0.0 +------------------------------
0.0 1.0 False Positive Rate
diagonal line = random ranking
The Area Under the ROC Curve (ROC-AUC) summarises ranking ability. An AUC of 1.0 represents perfect separation, while an AUC of 0.5 is approximately random ranking. A higher ROC curve is generally better because it provides a higher true-positive rate for the same false-positive rate.
The ROC curve is significant because it evaluates all thresholds, helps compare probability-producing classifiers, shows the trade-off between detecting positives and creating false alarms, and supports threshold selection based on application costs. However, ROC-AUC does not choose the operating threshold and does not show whether probabilities are calibrated. For highly imbalanced datasets, the precision-recall curve may be more informative because it focuses on the positive class and false discoveries. Final evaluation should include a confusion matrix at the selected threshold and metrics related to the real decision cost.
17. Create a Decision Tree using Gini Index for the Play Golf data.
The data is:
| Day | Weather | Temperature | Humidity | Wind | Play Golf |
|---|---|---|---|---|---|
| D1 | Sunny | Hot | High | Weak | NO |
| D2 | Sunny | Hot | High | Strong | NO |
| D3 | Cloudy | Hot | High | Weak | YES |
| D4 | Rain | Mild | High | Weak | YES |
| D5 | Rain | Cool | Normal | Weak | YES |
| D6 | Rain | Cool | Normal | Strong | NO |
| D7 | Cloudy | Cool | Normal | Strong | YES |
| D8 | Sunny | Cool | Normal | Weak | YES |
The target has five YES and three NO records. Therefore, the root Gini impurity is
Gini(root) = 1 − (5/8)² − (3/8)² = 0.46875.
For the candidate binary split Weather = Sunny versus Weather ≠ Sunny, the Sunny branch contains D1, D2, and D8, with two NO and one YES:
Gini(Sunny) = 1 − (2/3)² − (1/3)² = 0.4444.
The non-Sunny branch contains D3–D7, with four YES and one NO:
Gini(not Sunny) = 1 − (4/5)² − (1/5)² = 0.32.
Weighted Gini = (3/8)(0.4444) + (5/8)(0.32) ≈ 0.3667.
Gain ≈ 0.46875 − 0.3667 = 0.10208.
For the Sunny branch, split on Temperature = Hot: D1,D2 are NO and D8 is YES. For the non-Sunny branch, split on Wind = Weak versus Strong. The Weak child D3,D4,D5 is pure YES. The Strong child D6,D7 is then split on Weather: Rain gives NO and Cloudy gives YES.
Weather = Sunny?
├── yes: Temperature = Hot?
│ ├── yes → NO
│ └── no → YES
└── no: Wind = Weak?
├── yes → YES
└── no: Weather = Rain?
├── yes → NO
└── no (Cloudy) → YES
All leaves are pure for these eight observations. The exact tree can differ if ties are resolved differently or if a different split representation is used. Purity on a tiny dataset is not evidence of generalisation, so a real implementation should use stopping rules or pruning.
18. Compare Decision Trees and Logistic Regression for a medical diagnosis problem. Justify your choice.
| Criterion | Decision tree | Logistic regression |
|---|---|---|
| Relationship learned | Nonlinear, rule-based partitions | Linear relationship in log-odds |
| Interpretability | If–then paths and thresholds | Coefficients and odds ratios |
| Probability output | Leaf class proportions; may require calibration | Direct probability model; still requires validation/calibration |
| Feature preparation | Little scaling required; categorical handling depends on implementation | Encoding and scaling may be important |
| Interactions | Learned automatically through branches | Must be added explicitly or represented through features |
| Stability | Can be unstable and overfit when deep | Usually stable with suitable regularisation |
A decision tree is attractive when clinicians need a simple rule path, when important relationships are nonlinear, or when feature interactions such as “high age and high blood pressure” matter. A shallow, pruned tree can be presented as a screening aid. However, a single tree may be unstable, may produce poorly calibrated probabilities, and can create misleading rules from a small dataset.
Logistic regression is attractive when the number of patients is limited relative to the number of features, a mostly additive relationship is plausible, and a stable risk score is required. The coefficient βⱼ changes the log-odds by βⱼ per unit of xⱼ; e^βⱼ is the corresponding odds ratio when other variables are fixed. Ridge or Lasso regularisation can reduce instability and handle correlated predictors.
Justified choice: begin with regularised logistic regression as the primary baseline if predictors are structured clinical variables and calibrated risk probabilities are important. Compare it with a shallow, pruned decision tree using cross-validated sensitivity, specificity, calibration, subgroup fairness, and clinical interpretability. If strong nonlinearities are present, the tree may perform better; the final model should be externally validated and treated as decision support rather than an autonomous diagnosis.
19. Create a Decision Tree using Gini Index for the given decision data.
The data is:
| Example | Weather | Parents | Cash | Exam | Decision |
|---|---|---|---|---|---|
| 1 | Sunny | Visit | Rich | Yes | Cinema |
| 2 | Sunny | No-visit | Rich | No | Tennis |
| 3 | Windy | Visit | Rich | No | Cinema |
| 4 | Rainy | Visit | Poor | Yes | Cinema |
| 5 | Rainy | No-visit | Rich | No | Stay-in |
| 6 | Rainy | Visit | Poor | No | Cinema |
| 7 | Windy | No-visit | Poor | Yes | Cinema |
| 8 | Windy | No-visit | Rich | Yes | Shopping |
| 9 | Windy | Visit | Rich | No | Cinema |
| 10 | Sunny | No-visit | Rich | No | Tennis |
| 11 | Sunny | No-visit | Poor | Yes | Tennis |
The target classes are cinema (6), tennis (3), stay-in (1), and shopping (1). Hence,
Gini(root) = 1 − (6/11)² − (3/11)² − (1/11)² − (1/11)²
≈ 0.6116.
The best root split is Parents = Visit versus Parents = No-visit. The visit branch contains examples 1, 3, 4, 6, and 9; all have decision Cinema, so its Gini is 0. The no-visit branch contains examples 2, 5, 7, 8, 10, and 11; its class counts are tennis = 3 and one each of stay-in, cinema, and shopping:
Gini(no-visit) = 1 − (3/6)² − 3(1/6)² = 0.6667.
Weighted Gini = (5/11)(0) + (6/11)(0.6667) = 0.3636.
Gain ≈ 0.6116 − 0.3636 = 0.2479.
Within the no-visit branch, Weather = Sunny gives a pure Tennis child (examples 2, 10, 11). The remaining Rainy/Windy examples are split by Weather: Rainy gives Stay-in (example 5), while Windy gives examples 7 and 8. Finally, Cash splits those examples: Poor gives Cinema and Rich gives Shopping.
Parents = Visit?
├── yes → CINEMA
└── no: Weather = Sunny?
├── yes → TENNIS
└── no: Weather = Rainy?
├── yes → STAY-IN
└── no (Windy): Cash = Rich?
├── yes → SHOPPING
└── no (Poor) → CINEMA
The exact tree may vary if a multiway split is allowed or if tie-breaking rules differ. The tree perfectly fits this small dataset, so validation and pruning remain important.
20. Derive the Gradient Descent update rule for Linear Regression.
Let the training set contain m examples and n features. Include an intercept by defining xᵢ₀ = 1. The linear model is
ŷᵢ = θ₀xᵢ₀ + θ₁xᵢ₁ + ··· + θₙxᵢₙ = θᵀxᵢ.
Use the half mean-squared-error cost function:
J(θ) = (1/(2m)) Σᵢ (ŷᵢ − yᵢ)².
For a particular parameter θⱼ, differentiate the cost:
∂J/∂θⱼ = (1/m) Σᵢ (ŷᵢ − yᵢ)xᵢⱼ, j = 0,1,…,n.
The factor 1/2 cancels the 2 produced during differentiation. For learning rate α, the update rule is
θⱼ ← θⱼ − (α/m) Σᵢ (ŷᵢ − yᵢ)xᵢⱼ.
In matrix form,
θ ← θ − (α/m) Xᵀ(Xθ − y).
Gradient descent moves parameters opposite to the gradient. The algorithm is:
- Initialise the coefficients, often to zero or small values.
- Compute predictions for all training examples.
- Calculate residuals
ŷᵢ − yᵢ. - Calculate the gradient.
- Update all parameters simultaneously.
- Repeat until the cost decrease is small or a maximum number of iterations is reached.
If α is too small, convergence is slow; if it is too large, the cost may oscillate or diverge. Feature scaling often improves convergence when features have very different magnitudes. Batch gradient descent uses all examples per update, stochastic gradient descent uses one example, and mini-batch gradient descent uses a small batch.
21. Demonstrate how Regularisation can improve model performance.
Regularisation adds a penalty for model complexity to the training objective. Instead of minimising only training loss, the model minimises
Objective = Training loss + λ × Complexity penalty,
where λ ≥ 0 controls the strength of the penalty. The intercept is normally not penalised.
For linear regression, Ridge uses an L2 penalty:
J_Ridge = (1/(2m))Σᵢ(yᵢ−ŷᵢ)² + λΣⱼ₌₁ⁿ θⱼ².
Lasso uses an L1 penalty:
J_Lasso = (1/(2m))Σᵢ(yᵢ−ŷᵢ)² + λΣⱼ₌₁ⁿ |θⱼ|.
Ridge shrinks correlated coefficients toward zero but usually does not make them exactly zero. Lasso can make some coefficients exactly zero, so it performs a form of feature selection. Elastic Net combines L1 and L2 penalties and is useful when there are many correlated features.
The effect can be demonstrated with a high-degree polynomial model. Without regularisation, the curve can pass through nearly every training point, producing very low training MSE but large test MSE. As λ increases, large coefficients are penalised, the curve becomes smoother, and test error may decrease even though training error increases slightly. If λ becomes too large, the model becomes too simple and underfits.
Error
^ Training error ________
| /
| /
| Validation /\
| error / \________
+------------------------------> Regularisation strength
too little suitable too much
overfit underfit
Regularisation improves numerical stability when features are correlated and helps high-dimensional models. It must be combined with feature scaling for coefficient-based models so that units do not affect the penalty unfairly. Select the regularisation strength using cross-validation and use the final test set only once for unbiased evaluation.
Module III: Ensemble Learning and Model Validation
22. Define Ensemble Learning. Explain why ensemble methods generally outperform single models.
Ensemble Learning combines predictions from multiple base learners to produce one final prediction. The base learners may be decision trees, linear models, neural networks, or different algorithm families. The central idea is that several models can compensate for one another’s errors when their errors are not perfectly correlated.
For classification, common combination rules include majority voting, weighted voting, probability averaging, and stacking. For regression, predictions are commonly averaged or combined by a meta-model. If independent estimators have similar variance σ², averaging M of them can reduce variance approximately to σ²/M. In practice, the reduction is smaller when models are correlated, so diversity is important.
Ensembles generally outperform a single model for four reasons:
- Variance reduction: averaging high-variance learners such as deep decision trees makes the final prediction less sensitive to the particular training sample. Bagging and random forests use this principle.
- Bias reduction: sequential methods such as boosting add learners that focus on previous errors, allowing the combined model to represent difficult patterns.
- Error correction: a model that is wrong in one region may be corrected by another model with a different decision boundary or feature representation.
- Robustness: a weak individual learner or noisy feature may have less effect when many learners contribute to the final output.
For an ensemble of classifiers h₁,…,hₘ, hard voting is
ŷ = mode{h₁(x), …, hₘ(x)}.
Soft voting averages class probabilities:
p̂ₖ(x) = (1/M) Σₘ pₘₖ(x), ŷ = argmaxₖ p̂ₖ(x).
Ensembles do not automatically outperform every single model. They can increase computation, reduce interpretability, amplify common bias, and overfit if poorly tuned or if all learners make the same errors. Cross-validation, calibration, leakage control, and a simple baseline are still required. The main ensemble families include bagging, random forests, boosting, voting, and stacking.
23. Explain K-Fold Cross Validation with a neat diagram.
K-Fold Cross Validation is a resampling method used to estimate how a model is likely to perform on unseen data and to select models or hyperparameters. The dataset is divided into K approximately equal, non-overlapping folds. In each round, one fold is used for validation and the other K−1 folds are used for training. The process is repeated until every fold has served as the validation fold once.
Dataset divided into five folds
Round 1: [VALID] [TRAIN] [TRAIN] [TRAIN] [TRAIN]
Round 2: [TRAIN] [VALID] [TRAIN] [TRAIN] [TRAIN]
Round 3: [TRAIN] [TRAIN] [VALID] [TRAIN] [TRAIN]
Round 4: [TRAIN] [TRAIN] [TRAIN] [VALID] [TRAIN]
Round 5: [TRAIN] [TRAIN] [TRAIN] [TRAIN] [VALID]
If the score in round r is sᵣ, the cross-validation estimate is
s̄ = (1/K) Σᵣ₌₁ᴷ sᵣ.
The spread of fold scores indicates sensitivity to the split. Typical choices are K=5 or K=10, although the correct choice depends on dataset size and computational cost. For classification, Stratified K-Fold preserves approximately the same class proportions in each fold. For grouped observations, such as several records from the same patient, Group K-Fold prevents the same group from appearing in both training and validation. For time-dependent data, ordinary random K-Fold can leak future information; a time-series split is safer.
All learned preprocessing steps must be fitted separately inside each training fold. Otherwise, information from the validation fold can leak into the model. A final untouched test set is still recommended: cross-validation supports model selection, while the test set provides the final independent estimate. For nested model selection, an inner cross-validation loop tunes hyperparameters and an outer loop estimates performance.
24. Differentiate between Bagging and Boosting.
Both Bagging and Boosting combine multiple base learners, but they differ in how learners are trained and how their predictions are combined.
| Aspect | Bagging | Boosting |
|---|---|---|
| Full name | Bootstrap Aggregating | Sequential error-correcting ensemble |
| Training | Learners trained independently, often in parallel | Learners trained sequentially |
| Data relationship | Bootstrap samples drawn from training data | Later learners focus on residuals or misclassified observations |
| Main effect | Primarily reduces variance | Often reduces bias and can reduce both bias and variance |
| Combination | Average or majority vote | Weighted sum or weighted vote |
| Typical base learner | Deep decision trees | Shallow trees or decision stumps |
| Sensitivity | Relatively robust to noise and parallelisable | Can be sensitive to noisy labels and outliers |
| Examples | Bagged trees, Random Forest | AdaBoost, Gradient Boosting, XGBoost |
In bagging, bootstrap datasets are created by sampling training records with replacement. A separate model is trained on each sample. For regression,
f̂_bag(x) = (1/M) Σₘ f̂ₘ(x).
For classification, majority voting or probability averaging is used. Because the models are trained independently, bagging can be parallelised efficiently.
In boosting, the model is built as an additive sequence:
Fₘ(x) = F₀(x) + Σₘ αₘhₘ(x).
Each new learner is selected to improve the current ensemble. AdaBoost increases the weights of previously misclassified observations. Gradient boosting fits a new learner to the negative gradient or residual of the loss function. XGBoost adds regularisation and efficient implementation details.
Bagging is often preferred when the base model has high variance, such as a deep tree, and when robustness and parallel training are important. Boosting may give higher predictive accuracy on structured data but requires careful learning-rate, depth, number-of-estimators, and early-stopping control.
25. Explain Random Forest algorithm with suitable examples.
A Random Forest is an ensemble of decision trees trained using two kinds of randomness: bootstrap samples of observations and random subsets of features. This reduces the correlation among trees and usually improves generalisation compared with a single unpruned tree.
Algorithm
- Choose the number of trees
B. - For each tree, draw a bootstrap sample from the training data.
- At every split in that tree, randomly select only a subset of the available features.
- Find the best split using the selected features, commonly by Gini impurity for classification or squared error for regression.
- Grow the tree, subject to limits such as minimum leaf size.
- Combine predictions from all trees: use majority vote or averaged probabilities for classification, and the mean prediction for regression.
For a regression forest,
f̂_RF(x) = (1/B) Σᵦ Tᵦ(x).
For classification,
ŷ = mode{T₁(x), T₂(x), …, T_B(x)}.
An important feature is the out-of-bag (OOB) sample. A bootstrap sample contains roughly 63% unique training observations on average, so the remaining observations can estimate the tree’s prediction error without a separate validation set. OOB evaluation must still be used carefully when there is time dependence, grouping, or leakage.
Example: In loan-risk prediction, one tree may split first on income, another on credit score, and another on debt-to-income ratio. Because each tree sees a different sample and feature subset, their errors are less synchronised. Their combined prediction is generally more stable. In a regression application such as house-price prediction, the forest averages the predictions of many trees.
Advantages include strong performance on tabular data, automatic interaction learning, little need for scaling, resistance to overfitting compared with a single tree, and useful feature-importance tools. Limitations include higher memory and prediction cost, reduced interpretability, possible bias toward dominant classes or high-cardinality features, and less smooth probability calibration. Hyperparameters such as number of trees, maximum depth, max_features, and minimum leaf size should be selected using validation.
26. Explain Boosting, Decision Stumps, and XGBoost.
Boosting is an ensemble strategy that builds a strong learner from a sequence of weak learners. A weak learner performs only slightly better than random guessing, but each new learner is trained to correct errors made by the current ensemble. The final prediction is a weighted combination of all learners.
A Decision Stump is a decision tree with only one split. For example, a stump may ask whether income is greater than 50,000 and assign one prediction to each side. Stumps have high bias and low complexity, but many stumps can be combined by boosting. In AdaBoost, incorrectly classified observations receive larger weights, so the next stump pays more attention to them:
F(x) = Σₘ αₘhₘ(x).
Gradient Boosting generalises the error-correction idea to differentiable loss functions. Starting with an initial prediction F₀(x), each tree is fitted to the negative gradient of the loss at current predictions. The update is
Fₘ(x) = Fₘ₋₁(x) + ηhₘ(x),
where η is the learning rate. Small learning rates usually require more trees but can improve generalisation.
XGBoost, or Extreme Gradient Boosting, is an optimised and regularised gradient-boosted-tree system. It builds an additive ensemble of CART-style trees and uses a second-order approximation of the loss involving first derivatives gᵢ and second derivatives hᵢ. Its objective combines loss with a tree-complexity penalty:
Objective = Σᵢ l(yᵢ, ŷᵢ) + Σₖ Ω(fₖ).
XGBoost uses shrinkage, column subsampling, row sampling, regularisation, efficient split finding, and early stopping. These features can improve accuracy and training speed on structured/tabular data. Boosting is powerful but needs careful control of tree depth, learning rate, number of trees, subsampling, regularisation, and stopping criteria. It can overfit noisy labels or extreme outliers.
27. Compare Random Forest and XGBoost with respect to bias, variance, accuracy, and computational complexity.
Random Forest and XGBoost are both tree ensembles, but Random Forest is primarily a parallel bagging method, while XGBoost is a sequential boosting method.
| Criterion | Random Forest | XGBoost |
|---|---|---|
| Training principle | Bootstrap samples and random feature subsets | Add trees sequentially to reduce current loss |
| Bias | Usually lower than a single shallow model; may retain bias if trees are constrained | Often very low after enough boosting rounds |
| Variance | Strong variance reduction through averaging and decorrelation | Can be higher if trees are too deep or rounds are too many; regularisation controls it |
| Accuracy | Strong, robust baseline for tabular data | Often higher peak accuracy after careful tuning |
| Noise sensitivity | Generally more robust to noisy labels and outliers | More sensitive because later trees may chase noise |
| Parallelism | Trees can be trained independently | Tree rounds are sequential, although split and feature operations are parallelised |
| Main hyperparameters | Number of trees, depth, features per split, leaf size | Number of rounds, learning rate, depth, subsampling, regularisation |
| Computation | Simpler to tune and predictable; memory grows with trees/depth | More tuning and potentially longer sequential training; implementation is highly efficient |
Random Forest reduces variance by averaging many diverse high-variance trees. It is a strong first model when the dataset contains noise, parallel training is valuable, or limited tuning time is available. Its predictions may be less accurate than a well-tuned boosting model, and probability calibration may require extra work.
XGBoost reduces residual error stage by stage and includes explicit regularisation. It often produces excellent accuracy on structured data, learns nonlinearities and interactions, and supports missing-value handling and early stopping. However, its sequential nature, larger hyperparameter space, and sensitivity to data leakage or noisy labels make validation especially important.
The common practical pattern is that Random Forest is a robust, lower-variance baseline, while XGBoost can reach lower bias and higher accuracy but may overfit without regularisation. Complexity depends on the number of samples, features, trees, depth, and implementation; no single universal Big-O comparison captures every configuration.
28. Explain Boosting. Compare it with Bagging.
Boosting constructs an additive model in stages. Let Fₘ₋₁(x) be the current ensemble after m−1 learners. A new learner hₘ(x) is chosen to reduce the current loss, and the model becomes
Fₘ(x) = Fₘ₋₁(x) + ηαₘhₘ(x),
where η is the learning rate and αₘ may be a learner weight. In AdaBoost, examples are reweighted according to classification mistakes. In gradient boosting, a new learner approximates the negative gradient of the loss, often a residual for squared-error regression.
Boosting process
- Initialise a simple model or constant prediction.
- Measure current errors or loss gradients.
- Train a weak learner to focus on those errors.
- Multiply its contribution by a learning rate or weight.
- Add it to the ensemble.
- Repeat for a chosen number of rounds, using validation-based early stopping when appropriate.
| Feature | Boosting | Bagging |
|---|---|---|
| Learner dependence | Sequential; later learners depend on earlier ones | Independent; learners can be trained in parallel |
| Main target | Correct bias and residual errors | Reduce variance through averaging |
| Data use | Reweight examples or fit residuals | Bootstrap samples |
| Typical learners | Stumps or shallow trees | Often deep trees |
| Noise response | Can chase noisy observations | Usually more robust |
| Examples | AdaBoost, Gradient Boosting, XGBoost | Bagged trees, Random Forest |
Bagging is particularly effective for unstable models. If each tree has a different error due to sampling, averaging reduces variance. Boosting can model difficult structure efficiently because each learner is directed toward current weaknesses, but it is more sensitive to learning rate, depth, number of rounds, and noisy labels. Bagging is easier to parallelise and often requires less tuning. Boosting is often more accurate on well-prepared tabular data, but requires careful validation, regularisation, and threshold selection.
29. A company has developed two different classifiers. Analyse different ensemble techniques to combine their predictions and recommend the most suitable approach.
Assume the company has classifiers C₁ and C₂, each producing either a class label or a probability. First evaluate them separately on the same validation folds and check whether their errors are complementary. A classifier with slightly lower accuracy may still be valuable if it correctly identifies cases missed by the other model.
Hard voting: each classifier returns a class label and the majority class is selected. With only two classifiers, a tie-breaking rule is required, such as using the more reliable classifier or a third baseline. Hard voting is simple but discards probability information.
Soft voting or probability averaging: if both models output calibrated probabilities, compute
p_combined(y=1|x) = w p₁(y=1|x) + (1−w)p₂(y=1|x),
where 0 ≤ w ≤ 1 is selected using validation data. Predict the positive class when the combined probability exceeds an application-specific threshold. Soft voting is usually preferable to hard voting because it uses confidence information.
Weighted voting: assign a larger weight to the classifier with better validated performance, calibration, or lower cost-weighted error. The weight must be learned without overfitting the validation data.
Stacking: use predictions of the two classifiers as input features to a meta-classifier. The meta-classifier learns when to trust each base model. Its training features must be out-of-fold predictions; using predictions from models trained on the same records creates leakage.
Blending: similar to stacking, but a separate hold-out set is used to train the meta-model. It is simpler but uses less data for base-model training.
Recommendation: start with calibrated soft voting using a validation-selected weight and threshold. It is transparent, works well when the two classifiers have complementary errors, and avoids the complexity of a meta-model. Use stacking only if enough labelled data is available and cross-validated experiments show a reliable improvement. Select the final method using nested or repeated stratified cross-validation, and compare accuracy, precision, recall, F1-score, ROC-AUC, calibration, latency, and subgroup performance. Keep the simpler approach if its performance is statistically and operationally equivalent.
30. Evaluate the advantages and limitations of Boosting over Bagging for fraud detection applications.
Fraud detection is commonly an imbalanced, adversarial, and evolving classification problem. The positive class may be rare, false positives can inconvenience legitimate customers, and false negatives can cause financial loss.
Advantages of Boosting over Bagging
Boosting often achieves higher predictive accuracy because each new tree focuses on transactions that the current ensemble handles poorly. This can improve detection of subtle fraud patterns involving merchant, device, location, timing, and spending behaviour. Boosting learns nonlinear interactions and can optimise differentiable loss functions, class weights, or ranking objectives. Its probability scores can prioritise transactions for manual review, and its threshold can be adjusted to match the cost of false positives and false negatives. Boosting may use shallow trees efficiently and often performs very well on structured tabular data. XGBoost adds regularisation, row/column subsampling, and early stopping, which can help control overfitting.
Limitations of Boosting over Bagging
Boosting is sequential, so training is less naturally parallel than bagging. It has a larger hyperparameter space and may be sensitive to noisy labels, outliers, and concept drift. If later learners chase unusual legitimate transactions, false positives can rise. A model tuned too aggressively may overfit historical fraud patterns and fail when criminals change behaviour. Bagging methods such as Random Forest are often more robust and easier to tune. They provide a strong baseline, although bagging may have higher bias than a well-tuned boosting model and may miss some complex ranking structure.
Recommended fraud-detection design
Use a time-aware training/validation split because future transactions must not influence past predictions. Handle class imbalance using class weights, carefully designed sampling within training periods, or cost-sensitive objectives. Evaluate precision, recall, F1-score, PR-AUC, recall at a fixed false-positive rate, expected financial cost, and calibration. Monitor drift, investigate false positives, and retrain periodically. A practical choice is to compare a calibrated Random Forest baseline with regularised XGBoost, then deploy the model that gives the best validated cost-sensitive performance while meeting latency, explainability, and monitoring requirements.
31. Design an ensemble learning solution for spam email classification. Justify the choice of base learners, ensemble technique, validation strategy, and evaluation metrics.
Problem definition and data
The task is binary classification: label an email as spam or legitimate. The dataset should contain message text, subject, sender and domain features, URLs, attachment indicators, timestamps, and reliable labels. Duplicate or near-duplicate messages must be grouped before splitting. Personal content should be minimised, access-controlled, and securely stored.
Base learners
Use complementary models rather than many nearly identical models:
- Regularised logistic regression with TF-IDF word and character n-grams. It is fast, interpretable, and strong for sparse text. Character n-grams help detect obfuscated words and unusual URLs.
- Multinomial Naive Bayes. It is computationally efficient and often effective for word-count or TF-IDF features. Its conditional-independence assumption is unrealistic but can provide diverse errors.
- A tree-based model such as Random Forest or XGBoost on metadata. It can learn nonlinear relationships involving sender reputation, number of links, message length, time, attachment type, and authentication results.
- Optionally, a compact text-embedding classifier can capture semantic similarity, but it should be included only if latency, privacy, and data volume justify it.
Ensemble technique
Use probability-level stacking. Train the text models and metadata model separately. Generate out-of-fold probabilities for every training email. Train a regularised logistic-regression meta-learner on these probabilities and selected reliability features. The final model outputs a spam probability. Probability calibration using a validation set is important because the threshold determines how much legitimate mail is blocked.
If the dataset is small, use calibrated soft voting instead of stacking. A weighted average is easier to explain and less likely to overfit:
p_spam = w₁p_TFIDF-LR + w₂p_NB + w₃p_metadata,
where w₁ + w₂ + w₃ = 1.
Select weights on validation folds.
Validation strategy
Use a time-aware split because spam tactics change. Keep the most recent period as a final test set. Within the earlier period, use stratified or grouped cross-validation, ensuring that campaigns, sender domains, and near-duplicate messages do not occur across training and validation folds. All vectorisers, imputers, encoders, and resampling operations must be fitted inside each training fold. Tune the decision threshold according to the relative cost of spam reaching the inbox versus a legitimate email being quarantined.
Evaluation metrics
Report confusion-matrix metrics, but do not rely on accuracy because spam datasets are often imbalanced. Use precision, recall, F1-score, PR-AUC, false-positive rate for legitimate emails, and recall at an operationally acceptable false-positive rate. Also report calibration metrics such as log loss or Brier score, latency, memory use, and subgroup results for languages, domains, and message types. Review false positives manually and monitor performance after deployment.
The recommended solution is therefore a calibrated, time-validated stacking ensemble of TF-IDF logistic regression, Naive Bayes, and a metadata tree model. It combines complementary text and behavioural signals while remaining faster and easier to audit than a large opaque model.
References
- Scikit-learn documentation on supervised learning, decision trees, ensembles, model evaluation, preprocessing, and cross-validation.
- Standard machine-learning terminology for supervised learning, decision trees, ensemble methods, and model evaluation.