Logistic Regression
On this page
2.3 — Logistic Regression
Recall first. Linear regression can output any real number. Why is that unsuitable as a probability? What function could squeeze a score into
(0,1)?
The classification idea
Logistic regression is a supervised classification model that first forms a linear score and then maps it to a probability. For binary target y ∈ {0,1}:
z = β₀ + β₁x₁ + ... + βₚxₚ
p(x) = P(y=1 | x) = σ(z) = 1/(1+e^(−z))
The sigmoid is increasing, approaches 0 for very negative z, approaches 1 for very positive z, and gives 0.5 at z=0. The default class rule is often predict 1 if p ≥ 0.5, but the threshold is a decision choice, not part of the fitted probability model.12
The model is linear in log-odds:
log[p/(1−p)] = β₀ + Σ βⱼxⱼ
Therefore a one-unit increase in xⱼ, holding other features fixed, changes log-odds by βⱼ; it multiplies the odds by e^(βⱼ). This is usually more precise than saying the probability increases by βⱼ—the probability change depends on the starting probability.
Why not least squares?
A linear-regression probability can be below 0 or above 1, and squared-error fitting does not naturally model Bernoulli outcomes. Logistic regression estimates parameters by maximum likelihood, equivalently minimizing average binary cross-entropy/log-loss:
J(β) = −(1/n)Σ [yᵢ log pᵢ + (1−yᵢ)log(1−pᵢ)]
A confident wrong prediction is penalized heavily. If regularization is used, the practical objective adds a penalty, commonly λ||β||²/2 for L2 (with implementation-specific scaling). scikit-learn’s LogisticRegression documents regularized logistic regression and solver choices.1
Decision boundary and assumptions
With threshold 0.5, the boundary is p=0.5, hence z=0: a hyperplane in feature space. Logistic regression assumes the log-odds are linear in the predictors, not that the raw probability is linear. It also assumes independent observations for the usual likelihood interpretation. Strong interactions, nonlinear effects, separation, influential points, and correlated predictors can make coefficients unstable or the boundary inadequate. Transformations, interactions, regularization, or a different model may help.
Classification quality must be judged on held-out data and with metrics appropriate to class costs. A good probability model and a good 0.5 classification rule are related but not identical goals.23
Worked numerical trace
Suppose β₀ = −2, β₁ = 0.8, and x=3.
- Linear score:
z = −2 + 0.8(3) = 0.4. - Probability:
p = 1/(1+e^(−0.4)) ≈ 0.599. - At threshold 0.5, predict class 1.
- Odds are
p/(1−p) ≈ 0.599/0.401 ≈ 1.49. Sincee^0.8 ≈ 2.23, one extra unit ofxmultiplies the odds by about 2.23, not the probability by 0.8.
If the true label is y=1, the example’s log-loss is −log(0.599) ≈ 0.513. If the model had assigned p=0.01 to the true positive, loss would be −log(0.01) ≈ 4.605, showing why confident mistakes matter.
Exercise
For z = 2 − x, calculate p and the class at x=1 using threshold 0.5. Then explain what happens when the threshold changes to 0.8.
Revealed answer
At x=1, z=1, so p=σ(1)≈0.731; threshold 0.5 predicts class 1. At threshold 0.8 the same probability is below the threshold, so it predicts class 0. Raising the threshold usually makes positive predictions more selective: typically precision rises and recall falls, but the actual trade-off must be measured.
Exam lens
Remember the chain: linear score → sigmoid probability → log-odds interpretation → thresholded class. State that training minimizes log-loss/negative log-likelihood, not “sigmoid alone.” Do not say a coefficient is a fixed probability increase.
Rapid revision checklist
- Can I write the sigmoid and log-odds equations?
- Can I calculate a probability from a linear score?
- Can I explain why log-loss punishes confident wrong predictions?
- Can I distinguish probability estimation from threshold choice?
- Can I interpret
e^βⱼas an odds multiplier?
Key takeaways
- Logistic regression models
P(y=1|x)through a sigmoid of a linear score. - It is linear in log-odds, not generally in probability.
- Log-loss trains probability quality; the classification threshold encodes decision costs.
- Coefficients change conditional log-odds, and regularization is common in implementations.
Sources
Footnotes
-
scikit-learn, Logistic regression User Guide and
LogisticRegressionAPI. ↩ ↩2 -
Hastie, Tibshirani & Friedman, The Elements of Statistical Learning, official PDF, chapters 4 and 7. ↩ ↩2
-
Stanford CS229, course materials, classification and logistic-regression notes. ↩