Multivariate Linear Regression
On this page
2.2 — Multivariate Linear Regression
Recall first. Write the simple-regression model from 2.1. Now list two features that could jointly predict exam score, and predict what happens to the meaning of a slope when the other feature is held fixed.
From one predictor to many
Multivariate (more precisely, multiple) linear regression predicts a continuous response from p predictors:
yᵢ = β₀ + β₁xᵢ₁ + β₂xᵢ₂ + ... + βₚxᵢₚ + εᵢ
ŷᵢ = β₀ + Σⱼ βⱼxᵢⱼ
“Multivariate” is often used in courses for multiple predictors; strictly, multivariable means one response with many predictors, while multivariate regression can mean several responses. Follow the syllabus convention, but state the model clearly. The coefficients are estimated by ordinary least squares (OLS):
minimize_β RSS(β) = Σᵢ(yᵢ − ŷᵢ)²
With a column of ones added to X, the matrix form is y = Xβ + ε. When XᵀX is invertible, the OLS solution is:
β̂ = (XᵀX)⁻¹Xᵀy
In software, a stable least-squares solver is preferred to explicitly forming the inverse. OLS and the linear-model notation are described in the [scikit-learn linear-model guide]1 and in Hastie, Tibshirani & Friedman’s Elements of Statistical Learning (ESL).2
Interpretation: the phrase that earns marks
βⱼ is the expected change in y for a one-unit increase in xⱼ, holding the other included predictors constant. For example,
score = 20 + 4(study_hours) + 0.3(attendance_percent)
The coefficient 4 means four score points per additional study hour at the same attendance percentage. It does not prove that studying causes four extra points. If two predictors are correlated, “holding one fixed” may describe a rare or unrealistic comparison.
β₀is predictedywhen every predictor is zero; it may be outside the data’s meaningful range.- A positive coefficient indicates a positive conditional association; a negative coefficient indicates a negative one.
- Units matter: changing hours to minutes changes the coefficient’s numerical value but not predictions.
- A coefficient is not automatically an importance ranking. Scale, units, noise, and collinearity affect its size.
Collinearity
If predictors contain nearly redundant information, XᵀX is ill-conditioned or singular. Coefficients can then become large, unstable, and difficult to interpret even when predictions are reasonable. Check correlations, condition numbers, or variance inflation factors; remove, combine, or regularize redundant predictors when justified. Do not interpret a large coefficient as strong evidence without checking uncertainty and the design matrix.2
Assumptions and diagnostics
For prediction, the most important practical checks are that the training examples represent deployment and that the conditional relationship is sufficiently modeled. Classical coefficient tests and confidence intervals commonly rely on:
- Linearity/additivity: the conditional mean is linear in the chosen predictors; interactions or curves may be needed.
- Independent errors: residuals are not systematically correlated, especially in time or grouped data.
- Constant variance: residual spread is roughly stable (homoscedasticity).
- Normal errors: mainly needed for exact small-sample inference, not as a requirement for OLS predictions themselves.
- No perfect multicollinearity: predictors do not duplicate an exact linear combination.
Residual-vs-fitted plots reveal curvature and unequal spread; residuals in collection order reveal dependence; Q–Q plots help inspect normality. Add an interaction such as hours × attendance only if domain reasoning and validation support it—not because a larger formula is automatically better.23
Worked trace
Use two predictors and three observations:
x₁ hours | x₂ practice tests | y score |
|---|---|---|
| 1 | 0 | 3 |
| 0 | 1 | 2 |
| 2 | 1 | 5 |
The model ŷ = β₀ + β₁x₁ + β₂x₂ fits all three points. From rows 1 and 2, β₀ + β₁ = 3 and β₀ + β₂ = 2. Row 3 gives β₀ + 2β₁ + β₂ = 5. Substitution gives β₀ = 1, β₁ = 2, β₂ = 1.
So ŷ = 1 + 2x₁ + x₂. For a student with x₁=3, x₂=2, prediction is 1 + 6 + 2 = 9. The interpretation of β₁=2 is conditional: two more predicted score units for one more hour at the same number of practice tests. Real datasets usually have more rows than coefficients; OLS then finds the least-squares compromise rather than an exact fit.
Exercise
A fitted model is price = 50 + 0.2(area) − 3(age) + 10(near_station), where area is in m², age in years, and near_station is 0/1. Interpret the age and station coefficients. What is the prediction for area 100, age 5, near station 1?
Revealed answer
Holding area and station status fixed, each additional year is associated with a 3-unit decrease in predicted price. At the same area and age, being near the station changes prediction by +10 units. Prediction: 50 + 0.2(100) − 3(5) + 10 = 65.
Exam lens
Write: model → OLS objective → matrix solution → conditional interpretation → assumptions → one limitation (collinearity or leakage). Always say “holding other predictors constant.” Distinguish a coefficient’s association from causation and from marginal correlation.
Rapid revision checklist
- Can I write
ŷ = Xβand the OLS objective? - Can I interpret a slope with all other predictors held fixed?
- Can I explain why collinearity destabilizes coefficients?
- Can I name which assumptions matter for prediction versus inference?
- Can I calculate a prediction from several coefficients?
Key takeaways
- Multiple linear regression is a linear combination of many predictors.
- OLS chooses coefficients minimizing residual sum of squares.
- Coefficients are conditional associations, not automatic causal effects or feature-importance scores.
- Collinearity can harm interpretation even when predictive accuracy looks acceptable.
Sources
Footnotes
-
scikit-learn, Linear Models User Guide and LinearRegression API. ↩
-
Hastie, Tibshirani & Friedman, The Elements of Statistical Learning, 2nd ed., free official PDF. ↩ ↩2 ↩3
-
Penn State STAT 501, Multiple Linear Regression. ↩