0%
BuildUnder the hoodabout 25 min, 8 steps

KNN first, because you can say it in one sentence

Work through nearest-neighbour prediction, distance, scaling, vote shares, and the choice of k.

Work here, beside the explanation

K-nearest neighbours predicts from nearby labelled examples. Work through a four-row dataset by hand, verify its vote with a library, and then inspect real neighbours of a digit image. This makes 'nearby' and 'confidence' specific numerical claims you can check, rather than explanations that merely restate the model's name.

Each numbered stage below shows a complete program. Try this stage copies it into the editor beside the article, including every line it needs from earlier stages, so it works even after you reload the page. Read the program first, predict what it will print, then press Run. Loading a stage replaces what is in the editor; Undo brings your own version back.

Build with me · 1

1. Build four labelled points

Each inner row contains one feature, preserving the two-dimensional input table expected by scikit-learn. Red and blue are arbitrary class labels for this invented example. The query has no supplied label: it is the point to classify. A two-dimensional table with one column is still different from a flat list, just as a one-image batch differs from a bare pixel vector.

Python at this stageWorked example
PythonHover over a line to see an explanation
import numpy as npfrom sklearn.neighbors import KNeighborsClassifier training_values = np.array([[1.0], [2.0], [4.0], [8.0]])training_labels = np.array(["red", "red", "blue", "blue"])query = np.array([[3.2]])for values, label in zip(training_values, training_labels):    print("Input", values[0], "label", label)print("New input:", query[0, 0])

Run and sketch the four values on a number line, then place 3.2 between two and four.

What to look for

The four training values are 1, 2, 4, and 8; the query is 3.2.

Make it yours

Choose another query value and predict which training value will be nearest.

Build with me · 2

2. Calculate every distance

With one feature, Euclidean distance is the absolute coordinate difference. The colon selects all training rows and column zero. Subtracting the query produces signed differences; absolute value makes distance nonnegative. The label does not participate in calculating distance. Only after selecting neighbours do their labels determine the vote. That separation explains why feature representation matters before class voting occurs.

Python at this stageWorked example
PythonHover over a line to see an explanation
import numpy as npfrom sklearn.neighbors import KNeighborsClassifier training_values = np.array([[1.0], [2.0], [4.0], [8.0]])training_labels = np.array(["red", "red", "blue", "blue"])query = np.array([[3.2]])distances = np.abs(training_values[:, 0] - query[0, 0])for distance, label in zip(distances, training_labels):    print("Distance", round(float(distance), 3), "label", label)

Calculate the four distances manually before running.

What to look for

The distances are 2.2, 1.2, 0.8, and 4.8 in training order.

Make it yours

Move the query closer to eight and predict the new nearest label.

Build with me · 3

3. Select one nearest example

Argmin returns the position of the smallest distance. Using that position to index both the feature table and label array preserves their pairing. With one neighbour, the prediction is that neighbour's label. A very small k follows local details closely, including possible noise or mislabels. The model is not discovering that all points beyond a universal numerical boundary are inherently blue; it is applying the representation and voting rule you supplied.

Python at this stageWorked example
PythonHover over a line to see an explanation
import numpy as npfrom sklearn.neighbors import KNeighborsClassifier training_values = np.array([[1.0], [2.0], [4.0], [8.0]])training_labels = np.array(["red", "red", "blue", "blue"])query = np.array([[3.2]])distances = np.abs(training_values[:, 0] - query[0, 0])nearest_position = int(np.argmin(distances))print("Nearest value:", training_values[nearest_position, 0])print("k=1 vote:", training_labels[nearest_position])

Run and check that the closest point is four rather than two.

What to look for

The nearest value is 4.0 and k=1 predicts blue.

Make it yours

Place the query on the other side of the midpoint between two and four and observe the changed nearest neighbour.

Build with me · 4

4. Count a three-neighbour vote

Argsort returns positions ordered from nearest to farthest. Selecting the first three gives labels blue, red, and red, so red wins. K changed the answer even though neither query nor training data changed. Larger k smooths over local details but can blur boundaries or favour common classes. There is no universally best odd number, and odd k does not prevent every tie in a multiclass task.

Python at this stageWorked example
PythonHover over a line to see an explanation
import numpy as npfrom sklearn.neighbors import KNeighborsClassifier training_values = np.array([[1.0], [2.0], [4.0], [8.0]])training_labels = np.array(["red", "red", "blue", "blue"])query = np.array([[3.2]])distances = np.abs(training_values[:, 0] - query[0, 0])positions = np.argsort(distances)[:3]labels = training_labels[positions]classes, counts = np.unique(labels, return_counts=True)winner = classes[np.argmax(counts)]print("Neighbour values:", training_values[positions, 0])print("Neighbour labels:", labels)print("Vote counts:", dict(zip(classes.tolist(), counts.tolist())))print("k=3 prediction:", winner)

Run and compare the k=3 result with the k=1 result.

What to look for

The nearest three values are 4, 2, and 1; red wins two of three votes.

Make it yours

Change the query deliberately and inspect whether one and three neighbours still disagree.

Build with me · 5

5. Check the library's class mapping

Fit stores this labelled collection. Predict performs the neighbour search and voting you just traced. Predict_proba reports vote fractions under the default uniform voting rule. Its columns follow classes_, which may not match the order labels first appeared in your source. Two thirds support for red describes this local vote; it is not proof that predictions with that score are correct two thirds of the time on future real-world inputs.

Python at this stageWorked example
PythonHover over a line to see an explanation
import numpy as npfrom sklearn.neighbors import KNeighborsClassifier training_values = np.array([[1.0], [2.0], [4.0], [8.0]])training_labels = np.array(["red", "red", "blue", "blue"])query = np.array([[3.2]])model = KNeighborsClassifier(n_neighbors=3)model.fit(training_values, training_labels)print("Prediction:", model.predict(query)[0])print("Class order:", model.classes_)print("Vote shares:", model.predict_proba(query)[0])

Run and map each vote share to its printed class label.

What to look for

The library predicts red and assigns its two-of-three vote share to the correct class column.

Make it yours

Ask for a query far from every training point and notice that KNN still returns a vote.

Build with me · 6

6. Understand multiple-feature distance

Euclidean distance squares coordinate differences, sums them, and takes a square root. Differences three and four produce distance five. Across a digit's 64 features, each pixel contributes a term. A feature measured in thousands can dominate one measured in fractions; scaling changes that balance and should reflect the task. If you learn scaling statistics such as means from data, fit them using training data only. A one-pixel shift of a whole digit changes many coordinates even if a person sees the same symbol.

Python at this stageWorked example
PythonHover over a line to see an explanation
import numpy as npa = np.array([1.0, 2.0])b = np.array([4.0, 6.0])differences = b - adistance = np.sqrt(np.sum(differences ** 2))print("Coordinate differences:", differences)print("Euclidean distance:", distance)large_units = np.array([1000.0, 1.0])print("Distance with mismatched feature units:", np.linalg.norm(large_units))

Run and connect the result five with the familiar three-four-five triangle.

What to look for

The first distance is 5.0; the mismatched-units example is dominated by the large coordinate.

Make it yours

Rescale the large coordinate and explain how this changes the definition of similarity.

Build with me · 7

7. Inspect real digit neighbours

Kneighbors returns distances and positions in the training collection, not positions in development data. Use those positions to select y_train. The query's true label is displayed only for evaluation and does not enter the neighbour search. Inspecting these values makes the prediction traceable, although a near neighbour in pixel space may still differ in meaningful handwriting structure. The chosen distance is part of the model, not an objective definition of visual similarity.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier model = KNeighborsClassifier(n_neighbors=3)model.fit(X_train, y_train)distances, positions = model.kneighbors(X_dev[:1])print("Query true label:", y_dev[0])print("Neighbour training positions:", positions[0])print("Neighbour distances:", distances[0])print("Neighbour labels:", y_train[positions[0]])print("Prediction:", model.predict(X_dev[:1])[0])

Run and count the neighbours' labels yourself before reading the prediction.

What to look for

The model exposes the three actual training neighbours used for this development input.

Make it yours

Choose another one-image slice and compare its distance pattern with the first.

Build with me · 8

8. Compare planned k values on development data

The comparison changes only k while retaining data, scale, and split. Each candidate is created and fitted separately. The declared tie rule keeps selection reproducible. Choosing k to make one known example look good would overfit that example; development evidence evaluates a broader planned collection. KNN may fit quickly but retain many examples and spend work searching at prediction time, so memory and inference cost belong beside accuracy in a practical decision.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier results = []for neighbors in [1, 3, 5, 9]:    model = KNeighborsClassifier(n_neighbors=neighbors)    model.fit(X_train, y_train)    score = float(model.score(X_dev, y_dev))    results.append((neighbors, score))    print("k", neighbors, "development accuracy", round(score, 4))best = max(results, key=lambda result: (result[1], -result[0]))print("Chosen by development accuracy, smaller k on a tie:", best)print("Final test remains reserved:", len(y_final))

Run and report all four measured values, including weaker candidates.

What to look for

A real development comparison selects one k without using final-test predictions.

Make it yours

Plan a later runtime measurement if latency matters, keeping it separate from parameter-count guesses.

What a vote explains

A nearest-neighbour explanation can identify which stored examples influenced a prediction under a chosen distance rule. It does not prove that the training collection is representative, that an unfamiliar input belongs to a known class, or that vote shares are calibrated probabilities. Inspect inputs and neighbours, compare planned settings on development data, and evaluate the selected procedure on appropriate final evidence.

The complete reference is folded away below. Compare it with your work after trying the steps; changing a personal choice such as a name, a colour, or a display threshold can produce a different valid program.

Full reference solution

This is the final complete program built in the walkthrough. All its setup is included. Personal choices may differ in your own version; model scores are measured when you run, not promises about a future dataset.

PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier results = []for neighbors in [1, 3, 5, 9]:    model = KNeighborsClassifier(n_neighbors=neighbors)    model.fit(X_train, y_train)    score = float(model.score(X_dev, y_dev))    results.append((neighbors, score))    print("k", neighbors, "development accuracy", round(score, 4))best = max(results, key=lambda result: (result[1], -result[0]))print("Chosen by development accuracy, smaller k on a tie:", best)print("Final test remains reserved:", len(y_final))

Compare this with your version. Different names and personal choices are fine when the program follows the same logic.

Worth a look

Keep your progress

Sign in and every reading, quiz, and exercise you finish is saved.

Sign in