§ 2.3Module 2

Logistic Regression

On this page

2.3 — Logistic Regression

Recall first. Linear regression can output any real number. Why is that unsuitable as a probability? What function could squeeze a score into (0,1)?

The classification idea

Logistic regression is a supervised classification model that first forms a linear score and then maps it to a probability. For binary target y ∈ {0,1}:

z = β₀ + β₁x₁ + ... + βₚxₚ
p(x) = P(y=1 | x) = σ(z) = 1/(1+e^(−z))

The sigmoid is increasing, approaches 0 for very negative z, approaches 1 for very positive z, and gives 0.5 at z=0. The default class rule is often predict 1 if p ≥ 0.5, but the threshold is a decision choice, not part of the fitted probability model.12

The model is linear in log-odds:

log[p/(1−p)] = β₀ + Σ βⱼxⱼ

Therefore a one-unit increase in xⱼ, holding other features fixed, changes log-odds by βⱼ; it multiplies the odds by e^(βⱼ). This is usually more precise than saying the probability increases by βⱼ—the probability change depends on the starting probability.

Why not least squares?

A linear-regression probability can be below 0 or above 1, and squared-error fitting does not naturally model Bernoulli outcomes. Logistic regression estimates parameters by maximum likelihood, equivalently minimizing average binary cross-entropy/log-loss:

J(β) = −(1/n)Σ [yᵢ log pᵢ + (1−yᵢ)log(1−pᵢ)]

A confident wrong prediction is penalized heavily. If regularization is used, the practical objective adds a penalty, commonly λ||β||²/2 for L2 (with implementation-specific scaling). scikit-learn’s LogisticRegression documents regularized logistic regression and solver choices.1

Decision boundary and assumptions

With threshold 0.5, the boundary is p=0.5, hence z=0: a hyperplane in feature space. Logistic regression assumes the log-odds are linear in the predictors, not that the raw probability is linear. It also assumes independent observations for the usual likelihood interpretation. Strong interactions, nonlinear effects, separation, influential points, and correlated predictors can make coefficients unstable or the boundary inadequate. Transformations, interactions, regularization, or a different model may help.

Classification quality must be judged on held-out data and with metrics appropriate to class costs. A good probability model and a good 0.5 classification rule are related but not identical goals.23

Worked numerical trace

Suppose β₀ = −2, β₁ = 0.8, and x=3.

  1. Linear score: z = −2 + 0.8(3) = 0.4.
  2. Probability: p = 1/(1+e^(−0.4)) ≈ 0.599.
  3. At threshold 0.5, predict class 1.
  4. Odds are p/(1−p) ≈ 0.599/0.401 ≈ 1.49. Since e^0.8 ≈ 2.23, one extra unit of x multiplies the odds by about 2.23, not the probability by 0.8.

If the true label is y=1, the example’s log-loss is −log(0.599) ≈ 0.513. If the model had assigned p=0.01 to the true positive, loss would be −log(0.01) ≈ 4.605, showing why confident mistakes matter.

Exercise

For z = 2 − x, calculate p and the class at x=1 using threshold 0.5. Then explain what happens when the threshold changes to 0.8.

Revealed answer

At x=1, z=1, so p=σ(1)≈0.731; threshold 0.5 predicts class 1. At threshold 0.8 the same probability is below the threshold, so it predicts class 0. Raising the threshold usually makes positive predictions more selective: typically precision rises and recall falls, but the actual trade-off must be measured.

Exam lens

Remember the chain: linear score → sigmoid probability → log-odds interpretation → thresholded class. State that training minimizes log-loss/negative log-likelihood, not “sigmoid alone.” Do not say a coefficient is a fixed probability increase.

Rapid revision checklist

Key takeaways

Sources

Footnotes

  1. scikit-learn, Logistic regression User Guide and LogisticRegression API. ↩ ↩2

  2. Hastie, Tibshirani & Friedman, The Elements of Statistical Learning, official PDF, chapters 4 and 7. ↩ ↩2

  3. Stanford CS229, course materials, classification and logistic-regression notes. ↩