Sulba
000 / 100

Logistic regression

Many questions have a yes-or-no answer. Learn logistic regression: the sigmoid that turns a line into a probability, odds and log-odds, the decision boundary, the log loss it is trained on and its gradient, and softmax for more than two classes.

Lesson 8 stepsExercise Python, 7 functionsQuiz 5 questions

Step 1 Pass or fail

Hours studied by 12 students, and whether each passed (1) or failed (0).Predicting a yes or no is called classification.A straight line fitted by least squares, as in Linear regression:−0.1662 + 0.1669 × hoursAt 0 hours it says −0.17; at 8 hours, 1.17.But a chance of passing must be between 0 and 1.Classification needs a curve that rises from 0 to 1 and levels off.0246801hours studied
01/08

Some questions have a yes-or-no answer: will a student pass, is an email spam, is a payment fraudulent. Predicting which of a few groups something belongs to is called classification, and the groups are called classes. Here there are two classes: pass, written 1, and fail, written 0.

Take 12 students, with the hours each one studied and whether they passed. A straight line fitted by least squares, as in Linear regression, treats the 0s and 1s as ordinary numbers. Here is its prediction for each student, written ŷ as in Linear regression:

StudentHours xResult yLine ŷ = −0.1662 + 0.1669 × x
10.4failed (0)−0.1662 + 0.1669 × 0.4 = −0.09944 (below 0)
21failed (0)−0.1662 + 0.1669 × 1 = 0.0007
31.8failed (0)−0.1662 + 0.1669 × 1.8 = 0.13422
42.3failed (0)−0.1662 + 0.1669 × 2.3 = 0.21767
52.9failed (0)−0.1662 + 0.1669 × 2.9 = 0.31781
63.5passed (1)−0.1662 + 0.1669 × 3.5 = 0.41795
74.4passed (1)−0.1662 + 0.1669 × 4.4 = 0.56816
85.1failed (0)−0.1662 + 0.1669 × 5.1 = 0.68499
95.6passed (1)−0.1662 + 0.1669 × 5.6 = 0.76844
106.2passed (1)−0.1662 + 0.1669 × 6.2 = 0.86858
117passed (1)−0.1662 + 0.1669 × 7 = 1.0021 (above 1)
127.7passed (1)−0.1662 + 0.1669 × 7.7 = 1.11893 (above 1)

3 of the 12 predictions are below 0 or above 1. Over the full range of hours the line runs from below 0 to above 1:

at 0 hours: −0.1662 + 0.1669 × 0 = −0.1662

at 8 hours: −0.1662 + 0.1669 × 8 = 1.169

A chance below 0 or above 1 means nothing. A straight line also keeps rising for ever, which makes it easy to pull out of shape. Add one more student, who passed after 20 hours of study, and least squares flattens the whole line to get closer to them:

slope: 0.1669 before, 0.0596 after

at 6 hours, before: −0.1662 + 0.1669 × 6 = 0.8352

at 6 hours, after: 0.227 + 0.0596 × 6 = 0.5846

Nothing changed for the students near 6 hours, but their predictions dropped. Classification needs a curve that rises from 0 to 1 and levels off at both ends, so that a student far out on a flat end hardly matters.

Try it yourself

Set w and b to fit the curve to the students yourself, then let gradient descent train it. Move the threshold to see the decision boundary shift and the number of right answers change.

0246800.51hours studied

chance of passing, p gap to what happened wrong prediction

Filled dots passed (1), hollow dots failed (0). The dashed line across is the threshold, and the shaded side of the boundary is predicted to pass.

p = σ(0 × hours + 0)

Log loss = 0.6931

∂L/∂w = −0.8708, ∂L/∂b = 0

Next step: w = 0 − 0.1 × (−0.8708) = 0.0871, b = 0 − 0.1 × 0 = 0

No decision boundary: with w = 0, every student gets the same chance, 0.5.

Right: 6 of 12

Best fit: w = 1.2591, b = −5.0015, log loss 0.3153

Show the calculation
Studentxyz = w × x + bp = σ(z)PredictedChance given to what happenedLoss = −ln(chance)p − y(p − y) × x
10.400 × 0.4 + 0 = 00.5pass1 − 0.5 = 0.50.69310.50.5 × 0.4 = 0.2
2100 × 1 + 0 = 00.5pass1 − 0.5 = 0.50.69310.50.5 × 1 = 0.5
31.800 × 1.8 + 0 = 00.5pass1 − 0.5 = 0.50.69310.50.5 × 1.8 = 0.9
42.300 × 2.3 + 0 = 00.5pass1 − 0.5 = 0.50.69310.50.5 × 2.3 = 1.15
52.900 × 2.9 + 0 = 00.5pass1 − 0.5 = 0.50.69310.50.5 × 2.9 = 1.45
63.510 × 3.5 + 0 = 00.5pass0.50.6931−0.5(−0.5) × 3.5 = −1.75
74.410 × 4.4 + 0 = 00.5pass0.50.6931−0.5(−0.5) × 4.4 = −2.2
85.100 × 5.1 + 0 = 00.5pass1 − 0.5 = 0.50.69310.50.5 × 5.1 = 2.55
95.610 × 5.6 + 0 = 00.5pass0.50.6931−0.5(−0.5) × 5.6 = −2.8
106.210 × 6.2 + 0 = 00.5pass0.50.6931−0.5(−0.5) × 6.2 = −3.1
11710 × 7 + 0 = 00.5pass0.50.6931−0.5(−0.5) × 7 = −3.5
127.710 × 7.7 + 0 = 00.5pass0.50.6931−0.5(−0.5) × 7.7 = −3.85
Sum8.31780−10.45

Log loss = sum of losses ÷ n = 8.3178 ÷ 12 = 0.6931

∂L/∂w = sum of (p − y) × x ÷ n = (−10.45) ÷ 12 = −0.8708

∂L/∂b = sum of (p − y) ÷ n = 0 ÷ 12 = 0

x is hours studied, y is 1 for a pass and 0 for a fail, and n is the number of students (12). Values are rounded to 4 decimal places, or to 3 significant figures when very close to 0 or 1.

Press Run to train w and b by gradient descent, or set them yourself with the sliders.

Fit the curve by eye with the w and b sliders and see how low you can get the log loss, then press Best fit. Move the threshold and watch the boundary and the number of right answers change.

Programming exercise

Write the sigmoid, the log loss, its gradient and a training loop for logistic regression in plain Python, then use the model to classify. Save logistic.py and test_logistic.py in the same folder, fill in each function in logistic.py, and run the tests:

python test_logistic.py

"""Logistic regression: programming exercise. Write the sigmoid, the model's chance, the log loss and its gradient, atraining loop, and the predictions they lead to, in plain Python. Run thetests from this folder:     python test_logistic.py Data is a list of (x, y) pairs: x is the input, such as hours studied, and yis 1 for yes (passed) and 0 for no (failed). The model's chance that y is 1at input x is p = sigmoid(w * x + b).""" import math  def sigmoid(z):    """Return 1 / (1 + e**(-z)), a number between 0 and 1.     math.exp(-z) is e**(-z). For z below 0, return math.exp(z) / (1 + math.exp(z))    instead: it is the same number, and math.exp(-z) is too large for Python    once -z passes about 709.    """    raise NotImplementedError  def predict_proba(w, b, x):    """Return the model's chance that y is 1 at input x: sigmoid(w * x + b)."""    raise NotImplementedError  def log_loss(w, b, data):    """Return the average over data of -ln(p) for each y = 1 and -ln(1 - p) for each y = 0.     math.log(v) is ln v.    """    raise NotImplementedError  def gradient(w, b, data):    """Return (dL/dw, dL/db): the average of (p - y) * x, and the average of (p - y), over data."""    raise NotImplementedError  def fit(data, rate=0.1, steps=10000):    """Start from w = 0 and b = 0, take `steps` steps of gradient descent at this rate, and return (w, b)."""    raise NotImplementedError  def predict(w, b, x, threshold=0.5):    """Return 1 if the model's chance at x is at least threshold, otherwise 0."""    raise NotImplementedError  def accuracy(w, b, data, threshold=0.5):    """Return the share of data that predict gets right: a number between 0 and 1."""    raise NotImplementedError 

Stuck? math.exp(-z) is e to the power −z, and math.log(x) is ln x. The gradient is the average of (p − y) × x for w, and of p − y for b. The Solution tab has one way to write each function.

In practice: classifiers in production

Class imbalance. When one class is rare, say 1 fraudulent payment in 1,000, a model that always answers “not fraud” is 99.9% accurate and catches nothing. Give the rare class more weight in the loss (class_weight="balanced" in scikit-learn) or resample the training data, and judge the model with the measures in the next lesson, not with accuracy.

Perfect separation. If some boundary splits the training data with no mistakes, the log loss keeps falling as w grows, so training never settles. The weights grow without limit and every prediction heads to exactly 0 or 1. It shows up as weights that keep growing, or as a warning that training did not converge. The fix is Lesson 3’s penalty: scikit-learn’s LogisticRegression adds an L2 penalty by default, set by C, where a smaller C means a stronger penalty.

Scores, not probabilities, go into the loss. A network in PyTorch ends with raw scores z, called logits, and the loss applies the sigmoid itself: torch.nn.BCEWithLogitsLoss for two classes, torch.nn.CrossEntropyLoss (with softmax inside) for more. Working from z avoids taking ln 0 once σ(z) rounds to exactly 1. A common bug is to apply a sigmoid before BCEWithLogitsLoss. The loss then sees σ(σ(z)), which is always between 0.5 and 0.73, so it can never get close to 0. Nothing crashes, and the model trains badly.

Test your knowledge

  1. 01A model has w = 2 and b = −6. Where is its decision boundary, and what does it predict for 4 hours?Show answer

    Hours = −b ÷ w = 6 ÷ 2 = 3. At 4 hours z = 2 × 4 − 6 = 2 and p = σ(2) = 0.8808, so it predicts a pass.

  2. 02A student who passed was given p = 0.2. What is their log loss, and what would it be at p = 0.9?Show answer

    −ln 0.2 = 1.6094 and −ln 0.9 = 0.1054. The smaller the chance the model gave to what happened, the larger the loss.

  3. 03A model’s weight for hours studied is 0.7. What does one more hour do to the odds of passing?Show answer

    It adds 0.7 to z, the log-odds, so it multiplies the odds by e to the power 0.7 = 2.0138. One more hour roughly doubles the odds.

  4. 04Why is logistic regression trained on log loss rather than mean squared error?Show answer

    Through the sigmoid, squared error pushes only weakly on confident mistakes, and it is not bowl-shaped, so training can stall or stop short of the best fit. Log loss pushes hardest on confident mistakes and stays bowl-shaped.

  5. 05Why not train the model to get as many answers right as possible?Show answer

    The number of right answers jumps from one whole number to the next as w and b change, so its slope is 0 almost everywhere and gradient descent has nothing to follow. Log loss changes smoothly, and a model that gives good chances usually gets most answers right too.