Logistic regression
Many questions have a yes-or-no answer. Learn logistic regression: the sigmoid that turns a line into a probability, odds and log-odds, the decision boundary, the log loss it is trained on and its gradient, and softmax for more than two classes.
Step 1 Pass or fail
Some questions have a yes-or-no answer: will a student pass, is an email spam, is a payment fraudulent. Predicting which of a few groups something belongs to is called classification, and the groups are called classes. Here there are two classes: pass, written 1, and fail, written 0.
Take 12 students, with the hours each one studied and whether they passed. A straight line fitted by least squares, as in Linear regression, treats the 0s and 1s as ordinary numbers. Here is its prediction for each student, written ŷ as in Linear regression:
| Student | Hours x | Result y | Line ŷ = −0.1662 + 0.1669 × x |
|---|---|---|---|
| 1 | 0.4 | failed (0) | −0.1662 + 0.1669 × 0.4 = −0.09944 (below 0) |
| 2 | 1 | failed (0) | −0.1662 + 0.1669 × 1 = 0.0007 |
| 3 | 1.8 | failed (0) | −0.1662 + 0.1669 × 1.8 = 0.13422 |
| 4 | 2.3 | failed (0) | −0.1662 + 0.1669 × 2.3 = 0.21767 |
| 5 | 2.9 | failed (0) | −0.1662 + 0.1669 × 2.9 = 0.31781 |
| 6 | 3.5 | passed (1) | −0.1662 + 0.1669 × 3.5 = 0.41795 |
| 7 | 4.4 | passed (1) | −0.1662 + 0.1669 × 4.4 = 0.56816 |
| 8 | 5.1 | failed (0) | −0.1662 + 0.1669 × 5.1 = 0.68499 |
| 9 | 5.6 | passed (1) | −0.1662 + 0.1669 × 5.6 = 0.76844 |
| 10 | 6.2 | passed (1) | −0.1662 + 0.1669 × 6.2 = 0.86858 |
| 11 | 7 | passed (1) | −0.1662 + 0.1669 × 7 = 1.0021 (above 1) |
| 12 | 7.7 | passed (1) | −0.1662 + 0.1669 × 7.7 = 1.11893 (above 1) |
3 of the 12 predictions are below 0 or above 1. Over the full range of hours the line runs from below 0 to above 1:
at 0 hours: −0.1662 + 0.1669 × 0 = −0.1662
at 8 hours: −0.1662 + 0.1669 × 8 = 1.169
A chance below 0 or above 1 means nothing. A straight line also keeps rising for ever, which makes it easy to pull out of shape. Add one more student, who passed after 20 hours of study, and least squares flattens the whole line to get closer to them:
slope: 0.1669 before, 0.0596 after
at 6 hours, before: −0.1662 + 0.1669 × 6 = 0.8352
at 6 hours, after: 0.227 + 0.0596 × 6 = 0.5846
Nothing changed for the students near 6 hours, but their predictions dropped. Classification needs a curve that rises from 0 to 1 and levels off at both ends, so that a student far out on a flat end hardly matters.
Try it yourself
Set w and b to fit the curve to the students yourself, then let gradient descent train it. Move the threshold to see the decision boundary shift and the number of right answers change.
chance of passing, p gap to what happened wrong prediction
Filled dots passed (1), hollow dots failed (0). The dashed line across is the threshold, and the shaded side of the boundary is predicted to pass.
p = σ(0 × hours + 0)
Log loss = 0.6931
∂L/∂w = −0.8708, ∂L/∂b = 0
Next step: w = 0 − 0.1 × (−0.8708) = 0.0871, b = 0 − 0.1 × 0 = 0
No decision boundary: with w = 0, every student gets the same chance, 0.5.
Right: 6 of 12
Best fit: w = 1.2591, b = −5.0015, log loss 0.3153
Show the calculation
| Student | x | y | z = w × x + b | p = σ(z) | Predicted | Chance given to what happened | Loss = −ln(chance) | p − y | (p − y) × x |
|---|---|---|---|---|---|---|---|---|---|
| 1 | 0.4 | 0 | 0 × 0.4 + 0 = 0 | 0.5 | pass | 1 − 0.5 = 0.5 | 0.6931 | 0.5 | 0.5 × 0.4 = 0.2 |
| 2 | 1 | 0 | 0 × 1 + 0 = 0 | 0.5 | pass | 1 − 0.5 = 0.5 | 0.6931 | 0.5 | 0.5 × 1 = 0.5 |
| 3 | 1.8 | 0 | 0 × 1.8 + 0 = 0 | 0.5 | pass | 1 − 0.5 = 0.5 | 0.6931 | 0.5 | 0.5 × 1.8 = 0.9 |
| 4 | 2.3 | 0 | 0 × 2.3 + 0 = 0 | 0.5 | pass | 1 − 0.5 = 0.5 | 0.6931 | 0.5 | 0.5 × 2.3 = 1.15 |
| 5 | 2.9 | 0 | 0 × 2.9 + 0 = 0 | 0.5 | pass | 1 − 0.5 = 0.5 | 0.6931 | 0.5 | 0.5 × 2.9 = 1.45 |
| 6 | 3.5 | 1 | 0 × 3.5 + 0 = 0 | 0.5 | pass | 0.5 | 0.6931 | −0.5 | (−0.5) × 3.5 = −1.75 |
| 7 | 4.4 | 1 | 0 × 4.4 + 0 = 0 | 0.5 | pass | 0.5 | 0.6931 | −0.5 | (−0.5) × 4.4 = −2.2 |
| 8 | 5.1 | 0 | 0 × 5.1 + 0 = 0 | 0.5 | pass | 1 − 0.5 = 0.5 | 0.6931 | 0.5 | 0.5 × 5.1 = 2.55 |
| 9 | 5.6 | 1 | 0 × 5.6 + 0 = 0 | 0.5 | pass | 0.5 | 0.6931 | −0.5 | (−0.5) × 5.6 = −2.8 |
| 10 | 6.2 | 1 | 0 × 6.2 + 0 = 0 | 0.5 | pass | 0.5 | 0.6931 | −0.5 | (−0.5) × 6.2 = −3.1 |
| 11 | 7 | 1 | 0 × 7 + 0 = 0 | 0.5 | pass | 0.5 | 0.6931 | −0.5 | (−0.5) × 7 = −3.5 |
| 12 | 7.7 | 1 | 0 × 7.7 + 0 = 0 | 0.5 | pass | 0.5 | 0.6931 | −0.5 | (−0.5) × 7.7 = −3.85 |
| Sum | 8.3178 | 0 | −10.45 |
Log loss = sum of losses ÷ n = 8.3178 ÷ 12 = 0.6931
∂L/∂w = sum of (p − y) × x ÷ n = (−10.45) ÷ 12 = −0.8708
∂L/∂b = sum of (p − y) ÷ n = 0 ÷ 12 = 0
x is hours studied, y is 1 for a pass and 0 for a fail, and n is the number of students (12). Values are rounded to 4 decimal places, or to 3 significant figures when very close to 0 or 1.
Press Run to train w and b by gradient descent, or set them yourself with the sliders.
Fit the curve by eye with the w and b sliders and see how low you can get the log loss, then press Best fit. Move the threshold and watch the boundary and the number of right answers change.
Programming exercise
Write the sigmoid, the log loss, its gradient and a training loop for logistic regression in plain Python, then use the model to classify. Save logistic.py and test_logistic.py in the same folder, fill in each function in logistic.py, and run the tests:
python test_logistic.py
"""Logistic regression: programming exercise. Write the sigmoid, the model's chance, the log loss and its gradient, atraining loop, and the predictions they lead to, in plain Python. Run thetests from this folder: python test_logistic.py Data is a list of (x, y) pairs: x is the input, such as hours studied, and yis 1 for yes (passed) and 0 for no (failed). The model's chance that y is 1at input x is p = sigmoid(w * x + b).""" import math def sigmoid(z): """Return 1 / (1 + e**(-z)), a number between 0 and 1. math.exp(-z) is e**(-z). For z below 0, return math.exp(z) / (1 + math.exp(z)) instead: it is the same number, and math.exp(-z) is too large for Python once -z passes about 709. """ raise NotImplementedError def predict_proba(w, b, x): """Return the model's chance that y is 1 at input x: sigmoid(w * x + b).""" raise NotImplementedError def log_loss(w, b, data): """Return the average over data of -ln(p) for each y = 1 and -ln(1 - p) for each y = 0. math.log(v) is ln v. """ raise NotImplementedError def gradient(w, b, data): """Return (dL/dw, dL/db): the average of (p - y) * x, and the average of (p - y), over data.""" raise NotImplementedError def fit(data, rate=0.1, steps=10000): """Start from w = 0 and b = 0, take `steps` steps of gradient descent at this rate, and return (w, b).""" raise NotImplementedError def predict(w, b, x, threshold=0.5): """Return 1 if the model's chance at x is at least threshold, otherwise 0.""" raise NotImplementedError def accuracy(w, b, data, threshold=0.5): """Return the share of data that predict gets right: a number between 0 and 1.""" raise NotImplementedError Stuck? math.exp(-z) is e to the power −z, and math.log(x) is ln x. The gradient is the average of (p − y) × x for w, and of p − y for b. The Solution tab has one way to write each function.
In practice: classifiers in production
Class imbalance. When one class is rare, say 1 fraudulent payment in 1,000, a model that always answers “not fraud” is 99.9% accurate and catches nothing. Give the rare class more weight in the loss (class_weight="balanced" in scikit-learn) or resample the training data, and judge the model with the measures in the next lesson, not with accuracy.
Perfect separation. If some boundary splits the training data with no mistakes, the log loss keeps falling as w grows, so training never settles. The weights grow without limit and every prediction heads to exactly 0 or 1. It shows up as weights that keep growing, or as a warning that training did not converge. The fix is Lesson 3’s penalty: scikit-learn’s LogisticRegression adds an L2 penalty by default, set by C, where a smaller C means a stronger penalty.
Scores, not probabilities, go into the loss. A network in PyTorch ends with raw scores z, called logits, and the loss applies the sigmoid itself: torch.nn.BCEWithLogitsLoss for two classes, torch.nn.CrossEntropyLoss (with softmax inside) for more. Working from z avoids taking ln 0 once σ(z) rounds to exactly 1. A common bug is to apply a sigmoid before BCEWithLogitsLoss. The loss then sees σ(σ(z)), which is always between 0.5 and 0.73, so it can never get close to 0. Nothing crashes, and the model trains badly.
Test your knowledge
01A model has w = 2 and b = −6. Where is its decision boundary, and what does it predict for 4 hours?Show answer
Hours = −b ÷ w = 6 ÷ 2 = 3. At 4 hours z = 2 × 4 − 6 = 2 and p = σ(2) = 0.8808, so it predicts a pass.
02A student who passed was given p = 0.2. What is their log loss, and what would it be at p = 0.9?Show answer
−ln 0.2 = 1.6094 and −ln 0.9 = 0.1054. The smaller the chance the model gave to what happened, the larger the loss.
03A model’s weight for hours studied is 0.7. What does one more hour do to the odds of passing?Show answer
It adds 0.7 to z, the log-odds, so it multiplies the odds by e to the power 0.7 = 2.0138. One more hour roughly doubles the odds.
04Why is logistic regression trained on log loss rather than mean squared error?Show answer
Through the sigmoid, squared error pushes only weakly on confident mistakes, and it is not bowl-shaped, so training can stall or stop short of the best fit. Log loss pushes hardest on confident mistakes and stays bowl-shaped.
05Why not train the model to get as many answers right as possible?Show answer
The number of right answers jumps from one whole number to the next as w and b change, so its slope is 0 almost everywhere and gradient descent has nothing to follow. Log loss changes smoothly, and a model that gives good chances usually gets most answers right too.