0%
BuildUnder the hoodabout 29 min, 8 steps

The Learning Rate Is Your Step Size

See how the same gradient behaves under small, moderate, and excessive learning rates.

The update has two ingredients

This optional reading follows on from the gradient descent reading.

Gradient descent uses new parameter = old parameter - rate × gradient. Learning rate is a multiplier. The actual movement also depends on the gradient, so the parameter need not move by a fixed distance each step.

Reuse the loss (w - 6)² from the previous lesson. Start w at two, where the gradient is negative eight. With rate 0.01 the first movement is 0.08, producing w = 2.08. With rate 0.1 the movement is 0.8, producing 2.8. Both move toward six, at different speeds.

Overshooting in a case you can trace

With rate 1, the first update moves from two to ten. At ten, the gradient is eight, so the next update returns to two. Loss remains sixteen at both points. The updates bounce across the minimum instead of approaching it.

With rate 1.1, the first update reaches 10.8, farther from six than two was. Repeating the same rule drives the distance outward in this example. A training loss that grows rapidly is a signal to stop and inspect the setup, not to wait indefinitely for the model to improve.

These outcomes belong to this particular quadratic loss. Do not copy a successful numerical rate into every model. Different scales and algorithms change which rates are sensible.

Compare rates from the same starting point

For the same target-six example, record what two updates produce. Reinitialise w to two before each row's experiment.

Learning rateAfter one updateAfter two updatesLoss after two
0.012.082.158414.75789056
0.12.83.446.5536
0.5660
110216
1.110.80.2433.1776

At rate 0.5, the first update happens to land exactly on six for this particular loss. Do not generalise that success into a rule that 0.5 is always a good setting. The other rows show how a small change in a multiplier can turn convergence into bouncing or divergence.

Suppose you try a schedule: rate 0.1 for early updates, then 0.01. Record where the change occurs and continue from the current w. This is one experiment with two stages, not an independent fresh-start run at 0.01. Clear experiment notes prevent a comparison from accidentally mixing initialisation and learning-rate effects.

Run a fair comparison in blocks

Create a small numerical dataset and a line with a stated starting slope and intercept. Train for a fixed number of steps, print the loss, and record the rate. Then restart from the same initial line and repeat with a different rate.

If you continue training an already fitted line with a new rate, you are answering a different question: what happens when the rate changes midway? That can be useful, but it is not a comparison of two rates from the same starting point.

Change only the rate for this comparison. Keep examples, initial parameters, and step count fixed. If losses become non-finite or extremely large, stop and use a smaller rate before trying again.

Build with me · 1

Measure before changing anything

With both parameters zero, every prediction is zero. The squared errors are nine, twenty-five, and forty-nine; their mean is 83/3, about 27.667. This provides a baseline before any update.

Starting every comparison from the same line matters. Otherwise a setting that began close to the answer could look better simply because it had less work to do.

Blocks at this stageWorked example
make a dataset calledmeasurements
add tomeasurementsthe example1labeled3
add tomeasurementsthe example2labeled5
add tomeasurementsthe example3labeled7
make a line calledroutewith slope0and intercept0
sayInitial lossthen
the loss ofrouteonmeasurements

Create the three measurement rows and a route with zero slope and intercept. Print the loss.

What to look for

Initial loss is approximately 27.6667.

Make it yours

Calculate each squared error manually to confirm that the starting score belongs to the intended targets.

Build with me · 2

Take one measured step

The nudge block calculates how small slope and intercept changes affect mean squared error, then moves in the downhill direction. With rate 0.01, the first slope becomes about 0.2267 and intercept 0.1. Both change using information from all three rows.

The model does not try arbitrary guesses until one happens to work. The gradient is a numerical direction derived from the current errors and inputs.

Blocks at this stageWorked example
make a dataset calledmeasurements
add tomeasurementsthe example1labeled3
add tomeasurementsthe example2labeled5
add tomeasurementsthe example3labeled7
make a line calledroutewith slope0and intercept0
sayBeforethen
the loss ofrouteonmeasurements
nudgeroutetowardsmeasurementswith learning rate0.01
show lineroute
sayAfterthen
the loss ofrouteonmeasurements

Insert one nudge toward measurements with learning rate 0.01. Print the before and after loss and the new line.

What to look for

The first update reduces loss from about 27.667 to about 21.87.

Make it yours

Use the shown slope and intercept to calculate one updated prediction. It is still imperfect even though the loss improved.

Build with me · 3

Observe a sequence instead of only the endpoint

Each turn recomputes errors using the latest parameters, then takes another step. The direction can change as the line changes. Printing after each update shows whether the process is improving steadily, barely moving, or diverging.

A final low value alone hides this history. A final high value can result from different causes, such as too few sensible steps or steps large enough to overshoot.

Blocks at this stageWorked example
make a dataset calledmeasurements
add tomeasurementsthe example1labeled3
add tomeasurementsthe example2labeled5
add tomeasurementsthe example3labeled7
make a line calledroutewith slope0and intercept0
sayInitial lossthen
the loss ofrouteonmeasurements
repeat8times
nudgeroutetowardsmeasurementswith learning rate0.01
sayLoss after updatethen
the loss ofrouteonmeasurements
show lineroute

Put nudge followed by labelled loss inside repeat 8. Keep the zero-line declaration above the loop.

What to look for

At this modest rate the displayed losses decrease across the eight updates.

Make it yours

Move line creation inside the loop deliberately. Explain why every turn would repeat the same first update instead of continuing training.

Build with me · 4

Compare a much smaller update size

A very small rate moves parameters only a little on each update. The direction can be reasonable while progress is slow. Eight steps may barely change the starting loss compared with the previous rate.

Slow progress is not the same as an inability of a straight line to fit these particular points. The update size and number of steps affect how far the optimiser gets.

Blocks at this stageWorked example
make a dataset calledmeasurements
add tomeasurementsthe example1labeled3
add tomeasurementsthe example2labeled5
add tomeasurementsthe example3labeled7
make a line calledroutewith slope0and intercept0
sayInitial lossthen
the loss ofrouteonmeasurements
repeat8times
nudgeroutetowardsmeasurementswith learning rate0.0001
sayLoss after updatethen
the loss ofrouteonmeasurements
show lineroute

Set rate to 0.0001, preserve the same eight updates, and compare the trace with the earlier run.

What to look for

Loss decreases slowly and remains much higher after eight steps than with rate 0.01.

Make it yours

Increase only the number of steps. Explain the computation cost of compensating for an unnecessarily tiny rate.

Build with me · 5

Watch overshooting rather than assuming faster is better

A large rate can move past a useful region so far that the next gradient points back with even greater magnitude. Parameters and loss can grow instead of settling. The intention to move downhill does not guarantee that a finite step is small enough to reduce the loss.

The safe range depends on data scale and loss curvature. There is no universal rate that works for every dataset just because its decimal looks small.

Blocks at this stageWorked example
make a dataset calledmeasurements
add tomeasurementsthe example1labeled3
add tomeasurementsthe example2labeled5
add tomeasurementsthe example3labeled7
make a line calledroutewith slope0and intercept0
sayInitial lossthen
the loss ofrouteonmeasurements
repeat6times
nudgeroutetowardsmeasurementswith learning rate0.5
sayLoss after updatethen
the loss ofrouteonmeasurements
show lineroute

Compare rate 0.5 from the same zero line, keeping the trace short enough to inspect.

What to look for

The loss grows rapidly on these examples instead of converging.

Make it yours

Restore the modest rate and explain why reducing the rate addresses a different issue from collecting more labelled examples.

Read patterns without overclaiming

Slow, steady decrease can mean the rate is conservative, but it can also reflect a difficult objective. Oscillation or divergence can mean the rate is too large. A flat loss can mean convergence, vanishingly small changes, or a problem in the pipeline. Check values and data before assuming one explanation.

Feature scaling matters: measuring distance in metres instead of kilometres changes the gradients for a line unless the pipeline accounts for the conversion. Level 4 normalises digit brightness for the same general reason of controlling numerical scales.

Some training methods use schedules that reduce the rate later; adaptive optimisers also adjust updates using accumulated information. Neither removes the need for sensible settings and development evaluation. The goal is effective fitting with evidence, not finding a universal magic number.

Build with me · 6

Train with a sensible bounded budget

The train-line block repeats the same update rule for a specified budget. It does not guarantee exact convergence after two hundred steps. A small nonzero training loss can remain while parameters approach a fitting solution.

Slope and intercept are model parameters. The rate and step count are training settings you choose. Distinguishing these roles makes it easier to explain what training learned and what you selected.

Blocks at this stageWorked example
make a dataset calledmeasurements
add tomeasurementsthe example1labeled3
add tomeasurementsthe example2labeled5
add tomeasurementsthe example3labeled7
make a line calledroutewith slope0and intercept0
trainrouteonmeasurementsfor200steps with learning rate0.01
show lineroute
sayFinal lossthen
the loss ofrouteonmeasurements

Use train route on measurements for 200 steps with rate 0.01. Inspect the reported history, final equation, and loss.

What to look for

The bounded training run greatly reduces the loss and produces a line close to the observed trend.

Make it yours

Compare 50 and 200 steps from the same initial line. Record improvement and computation budget rather than claiming more is always better.

Build with me · 7

Use the fitted rule without another update

Prediction uses the fitted parameters but leaves them unchanged. Asking for two inputs does not perform two extra training steps. Four lies beyond the observed one-to-three range, so its output also illustrates extrapolation.

A good training fit is not a final evaluation. The formula has learned these three targets, and the wider task may have noise or omitted variables that a line cannot represent.

Blocks at this stageWorked example
make a dataset calledmeasurements
add tomeasurementsthe example1labeled3
add tomeasurementsthe example2labeled5
add tomeasurementsthe example3labeled7
make a line calledroutewith slope0and intercept0
trainrouteonmeasurementsfor200steps with learning rate0.01
sayFitted prediction at twothen
routepredicts for x =2
sayFitted prediction at fourthen
routepredicts for x =4
sayLoss after predictionsthen
the loss ofrouteonmeasurements

Place prediction outputs after the training block, with no nudge between them. Print the training loss again.

What to look for

The prediction at two is close to five. The loss remains the same as it was immediately after fitting.

Make it yours

Explain why a prediction that looks sensible at four is not a measured accuracy result without an observed answer there.

Build with me · 8

Make one controlled comparison report

This final program reports two points in one continuous run: ten updates and two hundred. The second loop adds one hundred ninety, rather than resetting and accidentally counting a different total.

Explain a training observation in terms of starting parameters, data, loss, rate, and update count. For this simple convex problem, a suitable gradient procedure can approach the best-fitting line; complex neural networks introduce further optimisation behaviour, so the small analogy has limits.

Blocks at this stageWorked example
make a dataset calledmeasurements
add tomeasurementsthe example1labeled3
add tomeasurementsthe example2labeled5
add tomeasurementsthe example3labeled7
make a line calledroutewith slope0and intercept0
repeat10times
nudgeroutetowardsmeasurementswith learning rate0.01
sayLoss after tenthen
the loss ofrouteonmeasurements
repeat190times
nudgeroutetowardsmeasurementswith learning rate0.01
sayLoss after two hundredthen
the loss ofrouteonmeasurements
show lineroute

Build the two bounded loops with labelled checkpoints and keep all initialisation above them.

What to look for

The two-hundred-update loss is lower than the ten-update loss with this rate and dataset.

Make it yours

Write a report comparing the slow, modest, and excessive rates using the observed traces. Separate optimisation progress from performance on new data.

Full reference solution

This is the complete worked program. Try building it yourself first, then use this reference to find the first place your version behaves differently. The Python below is generated from these exact blocks; helper functions are included so its behaviour can be inspected.

make a dataset calledmeasurements
add tomeasurementsthe example1labeled3
add tomeasurementsthe example2labeled5
add tomeasurementsthe example3labeled7
make a line calledroutewith slope0and intercept0
repeat10times
nudgeroutetowardsmeasurementswith learning rate0.01
sayLoss after tenthen
the loss ofrouteonmeasurements
repeat190times
nudgeroutetowardsmeasurementswith learning rate0.01
sayLoss after two hundredthen
the loss ofrouteonmeasurements
show lineroute
PythonHover over a line to see an explanation
# ---------------------------------------------------------------# Building blocks, written out in plain Python.# This part is generated for you. Your script starts further down.# ---------------------------------------------------------------  import mathimport random  class Dataset:    """A pile of labelled examples. Features can be a number, some text, or a    list of numbers. The label is whatever answer you want back."""     def __init__(self, name="dataset"):        self.name = name        self.rows = []        # Column names from the file's header row, when it came from one.        # Without these "the km of a row" has nothing to look the name up in.        self.columns = []     def add(self, features, label, fields=None):        self.rows.append(            {                "features": features,                "label": str(label),                "fields": dict(fields) if fields else {},            }        )     def size(self):        return len(self.rows)     def __len__(self):        return len(self.rows)     def __iter__(self):        """Walking a dataset gives you its rows, so "for each row in list        [testing]" reads the way it sounds."""        return iter(self.rows)     def labels(self):        seen = []        for row in self.rows:            if row["label"] not in seen:                seen.append(row["label"])        return seen     def most_common_label(self):        if not self.rows:            return ""        counts = {}        for row in self.rows:            counts[row["label"]] = counts.get(row["label"], 0) + 1        return max(counts, key=lambda label: counts[label])     def split(self, train_percent=80):        """Keeps the given percent for training and hands back the rest as a        test set. The shuffle is seeded, so you get the same split every run."""        order = list(range(len(self.rows)))        random.Random(0).shuffle(order)        cut = int(len(order) * train_percent / 100)        train = Dataset(self.name + " (train)")        test = Dataset(self.name + " (test)")        train.columns = list(self.columns)        test.columns = list(self.columns)        for position, index in enumerate(order):            row = self.rows[index]            target = train if position < cut else test            target.add(row["features"], row["label"], row.get("fields"))        return train, test     def show(self, limit=10):        print(self.name + ": " + str(len(self.rows)) + " examples")        for row in self.rows[:limit]:            print("  " + str(row["features"]) + "  ->  " + row["label"])        if len(self.rows) > limit:            print("  ... and " + str(len(self.rows) - limit) + " more")  def new_dataset(name="dataset"):    return Dataset(name)  def row_field(row, name):    """One named piece of a row: "the km of this row", "the label of it".     Names come from the header line of the csv. "label" always works, even on    a file with no header, because every row has one."""    wanted = str(name).strip()    if not isinstance(row, dict):        print("That is not a row. Use this inside a for each over a dataset.")        return ""    if wanted.lower() == "label":        return row.get("label", "")    fields = row.get("fields") or {}    if wanted in fields:        return fields[wanted]    # Header names are matched loosely, so "Rain" finds the "rain" column.    for key in fields:        if str(key).strip().lower() == wanted.lower():            return fields[key]    known = ", ".join([str(k) for k in fields]) if fields else "none"    print(        "No column called "        + wanted        + " in this row. Columns here: "        + known        + "."    )    return ""  def load_csv(dataset, path, label_column=-1, has_header=True):    """Reads a comma separated file into a dataset.     Everything except the label column becomes the features, and anything that    looks like a number is turned into one. This is how a file you uploaded    becomes something you can train on."""    try:        with open(path) as handle:            rows = [line.rstrip("\n").rstrip("\r") for line in handle]    except OSError:        print("Could not find " + path + ". Check the name in the Files panel.")        return dataset     rows = [row for row in rows if row.strip()]     # A file with no label column is a perfectly normal thing to load: it is    # what a test set looks like before you have predicted anything.    labelled = str(label_column).strip().lower() not in ("none", "", "no", "-")     header = []    if has_header and rows:        header = [cell.strip() for cell in rows[0].split(",")]        rows = rows[1:]     added = 0    for row in rows:        cells = [cell.strip() for cell in row.split(",")]        if not cells or (labelled and len(cells) < 2):            continue         index = -1        if labelled:            index = int(label_column)            if index < 0:                index = len(cells) + index            if index < 0 or index >= len(cells):                continue         label = cells[index] if labelled else ""        keep = [i for i in range(len(cells)) if i != index]         typed = []        for i in keep:            try:                typed.append(float(cells[i]))            except ValueError:                typed.append(cells[i])         # Every kept column gets its header name, so "the km of a row" works.        fields = {}        for position, i in enumerate(keep):            if i < len(header) and header[i]:                fields[header[i]] = typed[position]        if header and not dataset.columns:            dataset.columns = [header[i] for i in keep if i < len(header)]         dataset.add(typed[0] if len(typed) == 1 else typed, label, fields)        added += 1     print("Loaded " + str(added) + " rows from " + path + ".")    return dataset  def save_submission(predictions, dataset, path="submission.csv"):    """Prints your predictions in the exact two column format a round is    scored in: a header, then one id and one label per line.     It is printed rather than saved to a file because the console is the one    place you can copy it from. Paste it into a new file in the Files panel,    or straight into the upload box."""    labels = list(predictions or [])    rows = list(dataset) if dataset is not None else []     if len(labels) != len(rows):        print(            "You have "            + str(len(labels))            + " predictions for "            + str(len(rows))            + " rows. Those have to match before this means anything."        )        return     lines = ["id,label"]    for position, row in enumerate(rows):        # Looked up directly rather than through row_field, which would        # complain on every row of a file that simply has no id column.        fields = row.get("fields") or {}        row_id = ""        for key in fields:            if str(key).strip().lower() == "id":                row_id = fields[key]                break        if not str(row_id).strip():            row_id = "r" + str(position + 1).zfill(2)        if isinstance(row_id, float) and row_id == int(row_id):            row_id = int(row_id)        lines.append(str(row_id) + "," + str(labels[position]).strip())     print("--- submission.csv, copy from here ---")    for line in lines:        print(line)    print("--- to here, " + str(len(lines) - 1) + " rows ---")  class Line:    """y = slope * x + intercept. Two knobs, and learning is turning them."""     def __init__(self, slope=1.0, intercept=0.0, name="line"):        self.slope = float(slope)        self.intercept = float(intercept)        self.name = name     def predict(self, x):        return self.slope * float(x) + self.intercept     def _pairs(self, dataset):        pairs = []        for row in dataset.rows:            features = row["features"]            if isinstance(features, (list, tuple)):                features = features[0]            pairs.append((float(features), float(row["label"])))        return pairs     def loss(self, dataset):        """Mean squared error: the average of how wrong it is, squared."""        pairs = self._pairs(dataset)        if not pairs:            return 0.0        total = 0.0        for x, y in pairs:            error = self.predict(x) - y            # e * e rather than e ** 2. On a run that is blowing up, the power            # operator raises OverflowError and kills the script, while the            # multiply quietly reaches inf so the chart can report it.            total += error * error        average = total / len(pairs)        if average != average or average in (float("inf"), float("-inf")):            return float("inf")        return round(average, 4)     def step(self, dataset, rate=0.01):        """One nudge downhill. Look at how wrong the line is, then move both        knobs a small amount in the direction that makes it less wrong."""        pairs = self._pairs(dataset)        if not pairs:            return        slope_push = 0.0        intercept_push = 0.0        for x, y in pairs:            error = self.predict(x) - y            slope_push += 2 * error * x            intercept_push += 2 * error        self.slope -= rate * slope_push / len(pairs)        self.intercept -= rate * intercept_push / len(pairs)     def train(self, dataset, steps=200, rate=0.01):        """Many nudges in a row, with a chart of the loss as it goes.         The chart is the point. A run that bounces, a run that crawls and a        run that drops and flattens are three different problems, and the        final number alone tells them apart badly."""        count = int(steps)        history = [self.loss(dataset)]        for _ in range(count):            self.step(dataset, rate)            history.append(self.loss(dataset))         self._chart(history)        print(            "Trained "            + self.name            + " for "            + str(count)            + " steps. Loss is now "            + str(self.loss(dataset))        )     def _chart(self, history):        """A loss chart in text: one row per sampled step, bar length in        proportion to the loss at that point."""        if len(history) < 2:            return         # At most a dozen rows. The first few steps are always shown, because        # that is where a loss curve does nearly all of its moving, and a run        # that has settled by step 3 would otherwise look like a flat line.        wanted = 12        if len(history) <= wanted:            picks = list(range(len(history)))        else:            early = list(range(min(4, len(history))))            spread = wanted - len(early)            first = early[-1]            gap = (len(history) - 1 - first) / float(spread)            rest = [int(round(first + (i + 1) * gap)) for i in range(spread)]            picks = sorted(set(early + rest))            picks[-1] = len(history) - 1         finite = [history[i] for i in picks if history[i] < float("inf")]        top = max(finite) if finite else 0.0        width = len(str(len(history) - 1))         # Bars are drawn on a log scale. A first loss of 500 next to a final        # loss of 0.2 would otherwise flatten every interesting bar to nothing,        # and the shape of the drop is the whole reason to look at this.        scale = math.log(1.0 + top) if top > 0 else 0.0         print("Loss per step:")        for i in picks:            value = history[i]            step_label = str(i).rjust(width)            if value == float("inf"):                print("  step " + step_label + "  " + "off the chart")                continue            bars = 0            if scale > 0:                bars = int(round(30 * math.log(1.0 + value) / scale))            print(                "  step "                + step_label                + "  "                + ("#" * bars).ljust(30)                + " "                + str(round(value, 4))            )     def show(self):        print(            self.name            + ": y = "            + str(round(self.slope, 4))            + " * x + "            + str(round(self.intercept, 4))        )  def new_line(slope=1.0, intercept=0.0, name="line"):    return Line(slope, intercept, name)  # ---------------------------------------------------------------# Your script# ---------------------------------------------------------------  measurements = new_dataset("measurements")measurements.add(1, "3")measurements.add(2, "5")measurements.add(3, "7")route = new_line(0, 0, "route")for _ in range(int(10)):    route.step(measurements, 0.01)print("Loss after ten", route.loss(measurements))for _ in range(int(190)):    route.step(measurements, 0.01)print("Loss after two hundred", route.loss(measurements))route.show()

Compare this with your version. Different names and personal choices are fine when the program follows the same logic.

Keep your progress

Sign in and every reading, quiz, and exercise you finish is saved.

Sign in