Gradient Descent Is Hot and Cold
Understand an update by following the loss slope, first with one parameter and then with a fitted line.
Start with one adjustable number
This optional reading shows how a training procedure improves a parameter step by step.
Imagine a model that always predicts a number w, and one training target equal to 6. Its squared loss is (w - 6)². At w = 2, the error is -4 and the loss is 16. At w = 3, loss is 9. Moving upward helped.
You could search by trying nearby values and keeping better ones. That is a useful intuition for the goal. Gradient descent takes a more specific route: it calculates the local slope of the loss with respect to the parameter.
For this loss, the gradient is 2 × (w - 6). You do not need calculus to use the worked example, but the formula tells you both the direction and how steep the loss is locally.
Follow one update
At w = 2, the gradient is -8. With learning rate 0.1, the update is:
new w = old w - learning rate × gradientnew w = 2 - 0.1 × (-8)new w = 2.8Subtracting a negative number increases w, moving toward six. The new loss is (2.8 - 6)² = 10.24, lower than 16. The next update uses the gradient at 2.8, not the old gradient at two.
When w is above six, the gradient is positive, so subtracting it moves w down. At exactly six the gradient is zero. This smooth one-parameter example has an uncomplicated minimum; large neural networks have more complex loss surfaces.
Continue the updates, using the new value each time
Keep the target six and learning rate 0.1. The starting parameter is two. Each row below evaluates the current parameter before computing its replacement.
| Current w | Gradient 2 × (w - 6) | Next w | Loss at next w |
|---|---|---|---|
| 2 | -8 | 2.8 | 10.24 |
| 2.8 | -6.4 | 3.44 | 6.5536 |
| 3.44 | -5.12 | 3.952 | 4.194304 |
For the second update, subtract 0.1 × (-6.4) from 2.8. That adds 0.64, producing 3.44. Then calculate its loss: (3.44 - 6)² = 6.5536. The remaining gap is smaller, so the next gradient and update are smaller too.
This is why “step size” is shorthand: with one fixed learning rate, the actual movement can shrink as the gradient shrinks. Reusing the initial gradient negative eight on every iteration would implement a different procedure.
The loss curve plots these scores in execution order. It does not tell you how many examples were used unless the experiment records that separately. Keep the data, initial value, update rule, rate, and number of updates with the curve so someone else can reproduce it.
Update a line's two parameters
For prediction = m × x + b, let each residual be prediction minus target. For mean squared error, the slope gradient averages 2 × residual × x; the intercept gradient averages 2 × residual. These formulas measure how changing each parameter would change the current loss.
The Lines block nudge [line] towards [examples] with learning rate [...] performs one update using those averages. Put it in a repeat loop and print the loss at intervals to watch fitting progress. The train-line block performs several updates and prints a compact history.
Start the line at known slope and intercept values. Reinitialise to those same values when comparing different learning rates, otherwise the second experiment starts from the first experiment's fitted state.
Build with me · 1
Measure before changing anything
With both parameters zero, every prediction is zero. The squared errors are nine, twenty-five, and forty-nine; their mean is 83/3, about 27.667. This provides a baseline before any update.
Starting every comparison from the same line matters. Otherwise a setting that began close to the answer could look better simply because it had less work to do.
Create the three measurement rows and a route with zero slope and intercept. Print the loss.
What to look for
Initial loss is approximately 27.6667.
Make it yours
Calculate each squared error manually to confirm that the starting score belongs to the intended targets.
Build with me · 2
Take one measured step
The nudge block calculates how small slope and intercept changes affect mean squared error, then moves in the downhill direction. With rate 0.01, the first slope becomes about 0.2267 and intercept 0.1. Both change using information from all three rows.
The model does not try arbitrary guesses until one happens to work. The gradient is a numerical direction derived from the current errors and inputs.
Insert one nudge toward measurements with learning rate 0.01. Print the before and after loss and the new line.
What to look for
The first update reduces loss from about 27.667 to about 21.87.
Make it yours
Use the shown slope and intercept to calculate one updated prediction. It is still imperfect even though the loss improved.
Build with me · 3
Observe a sequence instead of only the endpoint
Each turn recomputes errors using the latest parameters, then takes another step. The direction can change as the line changes. Printing after each update shows whether the process is improving steadily, barely moving, or diverging.
A final low value alone hides this history. A final high value can result from different causes, such as too few sensible steps or steps large enough to overshoot.
Put nudge followed by labelled loss inside repeat 8. Keep the zero-line declaration above the loop.
What to look for
At this modest rate the displayed losses decrease across the eight updates.
Make it yours
Move line creation inside the loop deliberately. Explain why every turn would repeat the same first update instead of continuing training.
What the curve can tell you
A decreasing loss means the updates are improving the training objective on those examples. A flat curve might mean the fit is near a minimum, the step size is too small, or the model cannot represent the pattern. A rising or exploding loss may indicate excessive step size or numerical problems.
Ordinary gradient descent does not necessarily test a tentative update and undo it when loss rises. It follows its update rule. You diagnose the resulting curve and change the setup. The "hot and cold" analogy describes feedback, not an exact account of every algorithm.
Keep evaluation separate. It is possible for training loss to fall while development performance worsens. That is why a successful fitting loop is only one part of building a useful model.
Build with me · 4
Compare a much smaller update size
A very small rate moves parameters only a little on each update. The direction can be reasonable while progress is slow. Eight steps may barely change the starting loss compared with the previous rate.
Slow progress is not the same as an inability of a straight line to fit these particular points. The update size and number of steps affect how far the optimiser gets.
Set rate to 0.0001, preserve the same eight updates, and compare the trace with the earlier run.
What to look for
Loss decreases slowly and remains much higher after eight steps than with rate 0.01.
Make it yours
Increase only the number of steps. Explain the computation cost of compensating for an unnecessarily tiny rate.
Build with me · 5
Watch overshooting rather than assuming faster is better
A large rate can move past a useful region so far that the next gradient points back with even greater magnitude. Parameters and loss can grow instead of settling. The intention to move downhill does not guarantee that a finite step is small enough to reduce the loss.
The safe range depends on data scale and loss curvature. There is no universal rate that works for every dataset just because its decimal looks small.
Compare rate 0.5 from the same zero line, keeping the trace short enough to inspect.
What to look for
The loss grows rapidly on these examples instead of converging.
Make it yours
Restore the modest rate and explain why reducing the rate addresses a different issue from collecting more labelled examples.
Build with me · 6
Train with a sensible bounded budget
The train-line block repeats the same update rule for a specified budget. It does not guarantee exact convergence after two hundred steps. A small nonzero training loss can remain while parameters approach a fitting solution.
Slope and intercept are model parameters. The rate and step count are training settings you choose. Distinguishing these roles makes it easier to explain what training learned and what you selected.
Use train route on measurements for 200 steps with rate 0.01. Inspect the reported history, final equation, and loss.
What to look for
The bounded training run greatly reduces the loss and produces a line close to the observed trend.
Make it yours
Compare 50 and 200 steps from the same initial line. Record improvement and computation budget rather than claiming more is always better.
Build with me · 7
Use the fitted rule without another update
Prediction uses the fitted parameters but leaves them unchanged. Asking for two inputs does not perform two extra training steps. Four lies beyond the observed one-to-three range, so its output also illustrates extrapolation.
A good training fit is not a final evaluation. The formula has learned these three targets, and the wider task may have noise or omitted variables that a line cannot represent.
Place prediction outputs after the training block, with no nudge between them. Print the training loss again.
What to look for
The prediction at two is close to five. The loss remains the same as it was immediately after fitting.
Make it yours
Explain why a prediction that looks sensible at four is not a measured accuracy result without an observed answer there.
Build with me · 8
Make one controlled comparison report
This final program reports two points in one continuous run: ten updates and two hundred. The second loop adds one hundred ninety, rather than resetting and accidentally counting a different total.
Explain a training observation in terms of starting parameters, data, loss, rate, and update count. For this simple convex problem, a suitable gradient procedure can approach the best-fitting line; complex neural networks introduce further optimisation behaviour, so the small analogy has limits.
Build the two bounded loops with labelled checkpoints and keep all initialisation above them.
What to look for
The two-hundred-update loss is lower than the ten-update loss with this rate and dataset.
Make it yours
Write a report comparing the slow, modest, and excessive rates using the observed traces. Separate optimisation progress from performance on new data.
Full reference solution
This is the complete worked program. Try building it yourself first, then use this reference to find the first place your version behaves differently. The Python below is generated from these exact blocks; helper functions are included so its behaviour can be inspected.
# ---------------------------------------------------------------# Building blocks, written out in plain Python.# This part is generated for you. Your script starts further down.# --------------------------------------------------------------- import mathimport random class Dataset: """A pile of labelled examples. Features can be a number, some text, or a list of numbers. The label is whatever answer you want back.""" def __init__(self, name="dataset"): self.name = name self.rows = [] # Column names from the file's header row, when it came from one. # Without these "the km of a row" has nothing to look the name up in. self.columns = [] def add(self, features, label, fields=None): self.rows.append( { "features": features, "label": str(label), "fields": dict(fields) if fields else {}, } ) def size(self): return len(self.rows) def __len__(self): return len(self.rows) def __iter__(self): """Walking a dataset gives you its rows, so "for each row in list [testing]" reads the way it sounds.""" return iter(self.rows) def labels(self): seen = [] for row in self.rows: if row["label"] not in seen: seen.append(row["label"]) return seen def most_common_label(self): if not self.rows: return "" counts = {} for row in self.rows: counts[row["label"]] = counts.get(row["label"], 0) + 1 return max(counts, key=lambda label: counts[label]) def split(self, train_percent=80): """Keeps the given percent for training and hands back the rest as a test set. The shuffle is seeded, so you get the same split every run.""" order = list(range(len(self.rows))) random.Random(0).shuffle(order) cut = int(len(order) * train_percent / 100) train = Dataset(self.name + " (train)") test = Dataset(self.name + " (test)") train.columns = list(self.columns) test.columns = list(self.columns) for position, index in enumerate(order): row = self.rows[index] target = train if position < cut else test target.add(row["features"], row["label"], row.get("fields")) return train, test def show(self, limit=10): print(self.name + ": " + str(len(self.rows)) + " examples") for row in self.rows[:limit]: print(" " + str(row["features"]) + " -> " + row["label"]) if len(self.rows) > limit: print(" ... and " + str(len(self.rows) - limit) + " more") def new_dataset(name="dataset"): return Dataset(name) def row_field(row, name): """One named piece of a row: "the km of this row", "the label of it". Names come from the header line of the csv. "label" always works, even on a file with no header, because every row has one.""" wanted = str(name).strip() if not isinstance(row, dict): print("That is not a row. Use this inside a for each over a dataset.") return "" if wanted.lower() == "label": return row.get("label", "") fields = row.get("fields") or {} if wanted in fields: return fields[wanted] # Header names are matched loosely, so "Rain" finds the "rain" column. for key in fields: if str(key).strip().lower() == wanted.lower(): return fields[key] known = ", ".join([str(k) for k in fields]) if fields else "none" print( "No column called " + wanted + " in this row. Columns here: " + known + "." ) return "" def load_csv(dataset, path, label_column=-1, has_header=True): """Reads a comma separated file into a dataset. Everything except the label column becomes the features, and anything that looks like a number is turned into one. This is how a file you uploaded becomes something you can train on.""" try: with open(path) as handle: rows = [line.rstrip("\n").rstrip("\r") for line in handle] except OSError: print("Could not find " + path + ". Check the name in the Files panel.") return dataset rows = [row for row in rows if row.strip()] # A file with no label column is a perfectly normal thing to load: it is # what a test set looks like before you have predicted anything. labelled = str(label_column).strip().lower() not in ("none", "", "no", "-") header = [] if has_header and rows: header = [cell.strip() for cell in rows[0].split(",")] rows = rows[1:] added = 0 for row in rows: cells = [cell.strip() for cell in row.split(",")] if not cells or (labelled and len(cells) < 2): continue index = -1 if labelled: index = int(label_column) if index < 0: index = len(cells) + index if index < 0 or index >= len(cells): continue label = cells[index] if labelled else "" keep = [i for i in range(len(cells)) if i != index] typed = [] for i in keep: try: typed.append(float(cells[i])) except ValueError: typed.append(cells[i]) # Every kept column gets its header name, so "the km of a row" works. fields = {} for position, i in enumerate(keep): if i < len(header) and header[i]: fields[header[i]] = typed[position] if header and not dataset.columns: dataset.columns = [header[i] for i in keep if i < len(header)] dataset.add(typed[0] if len(typed) == 1 else typed, label, fields) added += 1 print("Loaded " + str(added) + " rows from " + path + ".") return dataset def save_submission(predictions, dataset, path="submission.csv"): """Prints your predictions in the exact two column format a round is scored in: a header, then one id and one label per line. It is printed rather than saved to a file because the console is the one place you can copy it from. Paste it into a new file in the Files panel, or straight into the upload box.""" labels = list(predictions or []) rows = list(dataset) if dataset is not None else [] if len(labels) != len(rows): print( "You have " + str(len(labels)) + " predictions for " + str(len(rows)) + " rows. Those have to match before this means anything." ) return lines = ["id,label"] for position, row in enumerate(rows): # Looked up directly rather than through row_field, which would # complain on every row of a file that simply has no id column. fields = row.get("fields") or {} row_id = "" for key in fields: if str(key).strip().lower() == "id": row_id = fields[key] break if not str(row_id).strip(): row_id = "r" + str(position + 1).zfill(2) if isinstance(row_id, float) and row_id == int(row_id): row_id = int(row_id) lines.append(str(row_id) + "," + str(labels[position]).strip()) print("--- submission.csv, copy from here ---") for line in lines: print(line) print("--- to here, " + str(len(lines) - 1) + " rows ---") class Line: """y = slope * x + intercept. Two knobs, and learning is turning them.""" def __init__(self, slope=1.0, intercept=0.0, name="line"): self.slope = float(slope) self.intercept = float(intercept) self.name = name def predict(self, x): return self.slope * float(x) + self.intercept def _pairs(self, dataset): pairs = [] for row in dataset.rows: features = row["features"] if isinstance(features, (list, tuple)): features = features[0] pairs.append((float(features), float(row["label"]))) return pairs def loss(self, dataset): """Mean squared error: the average of how wrong it is, squared.""" pairs = self._pairs(dataset) if not pairs: return 0.0 total = 0.0 for x, y in pairs: error = self.predict(x) - y # e * e rather than e ** 2. On a run that is blowing up, the power # operator raises OverflowError and kills the script, while the # multiply quietly reaches inf so the chart can report it. total += error * error average = total / len(pairs) if average != average or average in (float("inf"), float("-inf")): return float("inf") return round(average, 4) def step(self, dataset, rate=0.01): """One nudge downhill. Look at how wrong the line is, then move both knobs a small amount in the direction that makes it less wrong.""" pairs = self._pairs(dataset) if not pairs: return slope_push = 0.0 intercept_push = 0.0 for x, y in pairs: error = self.predict(x) - y slope_push += 2 * error * x intercept_push += 2 * error self.slope -= rate * slope_push / len(pairs) self.intercept -= rate * intercept_push / len(pairs) def train(self, dataset, steps=200, rate=0.01): """Many nudges in a row, with a chart of the loss as it goes. The chart is the point. A run that bounces, a run that crawls and a run that drops and flattens are three different problems, and the final number alone tells them apart badly.""" count = int(steps) history = [self.loss(dataset)] for _ in range(count): self.step(dataset, rate) history.append(self.loss(dataset)) self._chart(history) print( "Trained " + self.name + " for " + str(count) + " steps. Loss is now " + str(self.loss(dataset)) ) def _chart(self, history): """A loss chart in text: one row per sampled step, bar length in proportion to the loss at that point.""" if len(history) < 2: return # At most a dozen rows. The first few steps are always shown, because # that is where a loss curve does nearly all of its moving, and a run # that has settled by step 3 would otherwise look like a flat line. wanted = 12 if len(history) <= wanted: picks = list(range(len(history))) else: early = list(range(min(4, len(history)))) spread = wanted - len(early) first = early[-1] gap = (len(history) - 1 - first) / float(spread) rest = [int(round(first + (i + 1) * gap)) for i in range(spread)] picks = sorted(set(early + rest)) picks[-1] = len(history) - 1 finite = [history[i] for i in picks if history[i] < float("inf")] top = max(finite) if finite else 0.0 width = len(str(len(history) - 1)) # Bars are drawn on a log scale. A first loss of 500 next to a final # loss of 0.2 would otherwise flatten every interesting bar to nothing, # and the shape of the drop is the whole reason to look at this. scale = math.log(1.0 + top) if top > 0 else 0.0 print("Loss per step:") for i in picks: value = history[i] step_label = str(i).rjust(width) if value == float("inf"): print(" step " + step_label + " " + "off the chart") continue bars = 0 if scale > 0: bars = int(round(30 * math.log(1.0 + value) / scale)) print( " step " + step_label + " " + ("#" * bars).ljust(30) + " " + str(round(value, 4)) ) def show(self): print( self.name + ": y = " + str(round(self.slope, 4)) + " * x + " + str(round(self.intercept, 4)) ) def new_line(slope=1.0, intercept=0.0, name="line"): return Line(slope, intercept, name) # ---------------------------------------------------------------# Your script# --------------------------------------------------------------- measurements = new_dataset("measurements")measurements.add(1, "3")measurements.add(2, "5")measurements.add(3, "7")route = new_line(0, 0, "route")for _ in range(int(10)): route.step(measurements, 0.01)print("Loss after ten", route.loss(measurements))for _ in range(int(190)): route.step(measurements, 0.01)print("Loss after two hundred", route.loss(measurements))route.show()Compare this with your version. Different names and personal choices are fine when the program follows the same logic.
Keep your progress
Sign in and every reading, quiz, and exercise you finish is saved.