0%
BuildUnder the hoodabout 25 min, 8 steps

Logistic regression, your second model

Explain binary logistic regression, probability outputs, decision thresholds, and the limits of interpreting coefficients.

Work here, beside the explanation

Despite its name, logistic regression is used for classification. Trace a raw linear score through the sigmoid, separate that probability model from a chosen decision threshold, and fit a real small example. Then compare a multiclass version on digits without pretending that a binary threshold rule applies unchanged to ten classes.

Each numbered stage below shows a complete program. Try this stage copies it into the editor beside the article, including every line it needs from earlier stages, so it works even after you reload the page. Read the program first, predict what it will print, then press Run. Loading a stage replaces what is in the editor; Undo brings your own version back.

Build with me · 1

1. Calculate a linear score

This invented model uses z = 2x - 1. The coefficient multiplies the feature and the bias shifts the result. A raw score can be negative or greater than one; it is not yet a probability. With several features, the model sums several weighted contributions. A linear score limits the shapes of decision boundaries available in the original representation, so useful features still matter even when optimisation works perfectly.

Python at this stageWorked example
PythonHover over a line to see an explanation
inputs = [0.0, 0.5, 1.0]weight = 2.0bias = -1.0for value in inputs:    score = weight * value + bias    print("Input", value, "raw score", score)

Calculate the three scores before running.

What to look for

Inputs 0, 0.5, and 1 produce raw scores -1, 0, and 1.

Make it yours

Change the bias and explain how it shifts all three scores.

Build with me · 2

2. Apply the sigmoid

The sigmoid is 1 divided by 1 plus exp(-z), where exp is the exponential function. It maps zero to 0.5, positive scores toward one, and negative scores toward zero. This transformation produces a probability model for one class in the binary setting. A value between zero and one is not itself evidence of calibration on every input population. Training must fit useful coefficients and evaluation must test the intended conditions.

Python at this stageWorked example
PythonHover over a line to see an explanation
import math for value in [0.0, 0.5, 1.0]:    score = 2.0 * value - 1.0    probability = 1.0 / (1.0 + math.exp(-score))    print("Input", value, "score", score, "positive support", round(probability, 3))

Run and distinguish the raw score column from the transformed score.

What to look for

The three positive-class probabilities round to 0.269, 0.5, and 0.731.

Make it yours

Try a larger positive and negative raw score and describe the limiting behaviour.

Build with me · 3

3. Choose a decision threshold

The explicit greater-than-or-equal rule turns a probability into a label. Changing the threshold changes decisions without retraining weights. At 0.5, a probability exactly 0.5 qualifies as positive under this rule; at 0.8 none of these three points qualifies. Positive means the class designated by the task, not a desirable or morally good outcome. Threshold choice should reflect measured error consequences using development examples.

Python at this stageWorked example
PythonHover over a line to see an explanation
import math for threshold in [0.5, 0.8]:    print("Threshold:", threshold)    for value in [0.0, 0.5, 1.0]:        score = 2.0 * value - 1.0        probability = 1.0 / (1.0 + math.exp(-score))        prediction = int(probability >= threshold)        print("Input", value, "positive support", round(probability, 3), "decision", prediction)

Run and count positive decisions under both thresholds.

What to look for

The first threshold produces decisions 0, 1, 1; the second produces 0, 0, 0.

Make it yours

Choose another threshold and predict the decisions before running.

Build with me · 4

4. Fit coefficients from labelled examples

This small training table has one numerical feature and two integer class labels. Fit chooses coefficients and bias from labelled examples under the estimator's loss and regularisation. Unlike the earlier hand-chosen formula, the numerical parameters now come from a real optimisation. Their interpretation still depends on feature units and data coverage. A simple model is easier to inspect, but simplicity does not remove the need for careful evaluation.

Python at this stageWorked example
PythonHover over a line to see an explanation
import numpy as npfrom sklearn.linear_model import LogisticRegression X_train = np.array([[0.0], [1.0], [2.0], [3.0], [7.0], [8.0], [9.0], [10.0]])y_train = np.array([0, 0, 0, 0, 1, 1, 1, 1])model = LogisticRegression(random_state=42)model.fit(X_train, y_train)new_inputs = np.array([[2.5], [5.0], [8.5]])print("Learned coefficient:", model.coef_)print("Learned bias:", model.intercept_)print("Known classes:", model.classes_)

Run and inspect the fitted coefficient's sign and the class mapping.

What to look for

The model reports one coefficient, one bias, and classes 0 and 1.

Make it yours

Explain why changing the feature's unit scale can change the coefficient's numerical size without proving a stronger real-world effect.

Build with me · 5

5. Match probability columns to labels

Predict_proba returns one column per class. The code locates the column labelled 1 rather than assuming its position. Selecting all rows in that column yields the positive-class probabilities for the three new inputs. The class mapping matters even more when labels are words or nonconsecutive integers. A correct numerical distribution with an incorrect label interpretation still produces wrong decisions.

Python at this stageWorked example
PythonHover over a line to see an explanation
import numpy as npfrom sklearn.linear_model import LogisticRegression X_train = np.array([[0.0], [1.0], [2.0], [3.0], [7.0], [8.0], [9.0], [10.0]])y_train = np.array([0, 0, 0, 0, 1, 1, 1, 1])model = LogisticRegression(random_state=42)model.fit(X_train, y_train)new_inputs = np.array([[2.5], [5.0], [8.5]])positive_column = int(np.where(model.classes_ == 1)[0][0])probabilities = model.predict_proba(new_inputs)positive_probability = probabilities[:, positive_column]print("New inputs:", new_inputs[:, 0])print("Class order:", model.classes_)print("Positive probabilities:", positive_probability)print("Row totals:", probabilities.sum(axis=1))

Run and verify that the row totals are approximately one.

What to look for

Each new input has a measured probability for class 1, with class ordering shown explicitly.

Make it yours

Describe how you would map threshold decisions back to word labels rather than relying on astype(int).

Build with me · 6

6. Compare thresholds on fixed scores

A higher threshold makes no more positive decisions on the same fixed scores. It can reduce false positives while missing additional actual positives. However, observed precision need not increase monotonically on every finite dataset: the labels and ordering of examples matter. Measure the confusion counts for candidate thresholds on development data, choose a rule based on the task, and evaluate the whole chosen procedure on a final set.

Python at this stageWorked example
PythonHover over a line to see an explanation
import numpy as npfrom sklearn.linear_model import LogisticRegression X_train = np.array([[0.0], [1.0], [2.0], [3.0], [7.0], [8.0], [9.0], [10.0]])y_train = np.array([0, 0, 0, 0, 1, 1, 1, 1])model = LogisticRegression(random_state=42)model.fit(X_train, y_train)new_inputs = np.array([[2.5], [5.0], [8.5]])positive_column = int(np.where(model.classes_ == 1)[0][0])positive_probability = model.predict_proba(new_inputs)[:, positive_column]for threshold in [0.3, 0.5, 0.7]:    decisions = (positive_probability >= threshold).astype(int)    print("Threshold", threshold, "decisions", decisions)

Run and compare which individual decisions change as the threshold rises.

What to look for

The three threshold rules are applied to the same fitted score values without retraining.

Make it yours

State what labelled evaluation data would be needed to decide which threshold is useful.

Build with me · 7

7. Fit a multiclass digit model

The multiclass classifier produces scores for ten digit classes. This is not the same as applying the binary 0.5 threshold independently to every digit score. The model selects among its class outputs using the library's multiclass procedure. Each class's coefficient row describes how the represented pixel features contribute to its learned score. Input preparation and data roles remain the same as in KNN and the network lessons.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.linear_model import LogisticRegression model = LogisticRegression(max_iter=500, random_state=42)model.fit(X_train, y_train)print("Classes:", model.classes_)print("Coefficient shape:", model.coef_.shape)print("Development accuracy:", model.score(X_dev, y_dev))print("One probability row:", np.round(model.predict_proba(X_dev[:1])[0], 3))

Run and inspect the ten-class mapping, coefficient dimensions, and measured development score.

What to look for

The coefficient matrix has ten class rows and 64 feature columns.

Make it yours

Compare this model with KNN on the same development split, keeping the metric and preparation fixed.

Build with me · 8

8. Compare two actual model families

The comparison asks how two different procedures perform on the same prepared data. KNN votes among stored examples; logistic regression learns linear class scores. A coefficient describes an association in this fitted representation, not a causal law. Correlated features and scaling can redistribute weight magnitudes. Report what the model calculates and what the held-out evidence supports without claiming a pixel or feature causes an outcome in the broader world.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier model = KNeighborsClassifier(n_neighbors=3)model.fit(X_train, y_train)from sklearn.linear_model import LogisticRegression linear = LogisticRegression(max_iter=500, random_state=42)linear.fit(X_train, y_train)print("KNN development accuracy:", model.score(X_dev, y_dev))print("Logistic regression development accuracy:", linear.score(X_dev, y_dev))print("Logistic regression parameters:", linear.coef_.size + linear.intercept_.size)print("Final test remains reserved:", len(y_final))

Run and write a limited comparison using the two actual scores.

What to look for

Both development scores are measured on identical examples; the final set is not evaluated.

Make it yours

Choose a next comparison based on the error pattern or practical cost, rather than assuming the more complex name must win.

Keep three operations separate

A numerical score comes from fitted coefficients, a probability transformation changes its numerical interpretation, and a decision rule turns outputs into labels. Training and threshold selection are different parts of the full procedure. Evaluate the complete chosen procedure, including any rejection or review policy, using data that represents the intended task.

A positive coefficient raises the model's linear score as that represented feature increases while other represented features stay fixed. Its size depends on units and correlated inputs. That inspectable arithmetic does not establish a causal explanation of the world.

The complete reference is folded away below. Compare it with your work after trying the steps; changing a personal choice such as a name, a colour, or a display threshold can produce a different valid program.

Full reference solution

This is the final complete program built in the walkthrough. All its setup is included. Personal choices may differ in your own version; model scores are measured when you run, not promises about a future dataset.

PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier model = KNeighborsClassifier(n_neighbors=3)model.fit(X_train, y_train)from sklearn.linear_model import LogisticRegression linear = LogisticRegression(max_iter=500, random_state=42)linear.fit(X_train, y_train)print("KNN development accuracy:", model.score(X_dev, y_dev))print("Logistic regression development accuracy:", linear.score(X_dev, y_dev))print("Logistic regression parameters:", linear.coef_.size + linear.intercept_.size)print("Final test remains reserved:", len(y_final))

Compare this with your version. Different names and personal choices are fine when the program follows the same logic.

Keep your progress

Sign in and every reading, quiz, and exercise you finish is saved.

Sign in