A Model Is a Rule With Knobs
Use a two-parameter line to understand models, fitting, parameters, and the limits of the knob analogy.
A prediction you can calculate by hand
This optional reading uses a straight line to show what a model's adjustable numbers are. The stages use the Lines blocks: in this article's editor, choose the Lines group in Add blocks.
Suppose a small delivery service estimates travel time from distance. A proposed model is:
predicted minutes = 4 × distance in kilometres + 8For a three-kilometre trip, multiply 4 by 3 to get 12, then add 8 to get 20 minutes. The input is distance, the output is predicted time, and the calculation is the model's form.
The numbers 4 and 8 are parameters. You can think of them as knobs: changing either changes predictions while keeping the multiply-then-add form. Four controls time added per kilometre; eight supplies a starting amount independent of distance.
Build with me · 1
Create a specific rule
Slope four and intercept eight define one line. The fixed form is multiply input by slope, then add intercept. The two numbers are parameters, while distance is an input supplied for each prediction.
We chose these parameters by hand. No training has occurred yet, and a plausible equation is not proof that distance explains every delivery delay.
Find Lines in Add blocks. Create route with slope 4 and intercept 8, then add show line route.
What to look for
The equation shows a slope of 4 and intercept of 8.
Make it yours
Name the units: minutes per kilometre for slope and minutes for intercept.
Build with me · 2
Substitute a distance into the rule
For three kilometres, four times three is twelve, plus eight is twenty. The rounded prediction value computes a number but does not display it until placed inside an output statement.
Predicting uses the current parameters without changing them. This is the forward calculation, whether the parameters were chosen manually or learned from examples.
Put route predicts for x = 3 in a labelled say socket below the declaration.
What to look for
Minutes for 3 km is 20.0.
Make it yours
Choose a distance between one and eight and calculate its result before running.
Distinguish the roles
Training examples contain distances and observed times. Fitting chooses parameter values that reduce a defined error across those examples. Prediction substitutes a new distance into the fitted calculation. Evaluation compares predictions with observed times from examples reserved for evaluation.
The training algorithm also has settings, such as how large each update should be. These are often called hyperparameters. A slope learned from examples is a model parameter. A learning rate you choose for the training algorithm is a training setting. Both are numbers, but they play different roles.
In a neural network, ordinary training usually adjusts weights and biases within a chosen architecture. In other model families, fitting can construct a tree or retain examples instead. The phrase "rule with knobs" is useful for lines and neural networks; it is not a complete definition of every machine-learning algorithm.
Build with me · 3
Check what the intercept means
When input is zero, the slope term is zero, so the prediction equals the intercept. At one, the prediction includes one slope-sized increment. The values are eight and twelve in this example.
You may interpret eight as a preparation-time approximation, but that is a hypothesis about the world. If no zero-distance journey was observed, the intercept is an implied model value rather than a measured fact.
Print route at zero and one without changing its parameters.
What to look for
At zero is 8.0 and At one is 12.0.
Make it yours
Explain the difference between the equation's starting value and a claim about an actual preparation process.
Build with me · 4
Measure a change in output per input unit
Going from two to five kilometres adds three kilometres. At four minutes per kilometre, the predicted time grows by twelve minutes. The intercept is present in both predictions and cancels when we subtract them.
Slope describes change, not the height of a point on the graph. A high intercept can create a high prediction even with a zero slope.
Print the two predictions and nest their difference in another output block.
What to look for
The predictions are 16.0 and 28.0; their difference is 12.0.
Make it yours
Change slope to zero. Predict both outputs and explain why they become equal.
Turn one knob at a time
The next two stages change the parameters by hand. No training happens here: you choose the numbers. To fit a line from data, you also need examples and a loss, one number saying how wrong the line is, so that the training procedure knows what counts as improvement.
Build with me · 5
Shift every prediction by the same amount
Changing intercept from eight to ten adds two to every prediction. It does not change how much the model grows per kilometre. This can address a consistent offset in a controlled example.
The set-knob statement changes the existing line after the before outputs. Reading the stack in order lets you compare versions within one run rather than remembering results from separate runs.
Insert set intercept of route to 10 between the before and after outputs. Leave slope fixed.
What to look for
At one changes from 12 to 14; at five changes from 28 to 30.
Make it yours
Try intercept six and predict the common shift. Restore eight for the next stage.
Build with me · 6
Make the effect depend on distance
Increasing slope by one adds one extra minute for each kilometre. Compared with the original rule, the one-kilometre prediction grows by one and the five-kilometre prediction by five.
A slope adjustment and an intercept adjustment therefore have different signatures across the dataset. Inspect multiple points rather than choosing a parameter change from one lucky example.
Set slope to 5 while keeping intercept 8, then print predictions at one and five.
What to look for
At one is 13.0 and At five is 33.0.
Make it yours
Choose a negative slope and explain the resulting direction of change. A computable line is not necessarily plausible for the intended task.
Training is the procedure that chooses these numbers from examples. Prediction is the calculation using the numbers already chosen. Neither operation decides whether distance alone is an adequate explanation of delivery time. That is a modelling choice you must evaluate.
If rain and traffic change the answer substantially, a beautifully fitted distance-only line can still miss them. More training steps cannot insert an input your program never supplies. Begin debugging by distinguishing the rule's form, the available inputs, the parameter values, and the evidence used to judge them.
Build with me · 7
Preserve the rule when units change
Two kilometres is two thousand metres. To preserve the same physical predictions, slope becomes 0.004 minutes per metre instead of four minutes per kilometre. Intercept stays eight minutes.
Changing the input unit without changing the matching parameter creates a different model. Units are part of the meaning of a number, even though the computer can multiply mismatched numbers without recognising your mistake.
Create the metres version with slope 0.004 and feed it 2000 and 5000.
What to look for
The predictions remain 16.0 and 28.0, matching two and five kilometres.
Make it yours
Calculate what the incorrect slope four would produce for 2000, then explain why the huge value is a units error.
Explain a model's limits
A straight line assumes a constant increase in predicted time for each extra kilometre. Real journeys can involve traffic, routes, weather, and stops that this one-feature model does not represent. A good fit to a few observations does not establish that the line captures the true cause of travel time.
Before using a model, state its inputs, output, adjustable parameters, and intended range. Then ask what information it cannot use. In this example, two routes of equal distance get the same prediction even if one has many traffic lights. That limitation follows directly from the input representation.
Build with me · 8
Keep observations beside the prediction rule
The observed distances span one to eight kilometres. A prediction at 4.5 lies between observed inputs; twenty lies beyond them. Both can be calculated, but the second extends the trend into a region without supporting observations.
The loss compares this hand-chosen line with all eight observed answers. It does not prove causation or cover omitted features such as traffic. State the input range, units, parameters, and limitations together.
Create the eight measurements, show them, and compare the two predictions with their position relative to the observed range.
What to look for
At 4.5 km the model predicts 26 minutes; at 20 km it predicts 88. Training MSE for the supplied observations is 0.5.
Make it yours
Explain why the farther prediction does not become trustworthy merely because the arithmetic is exact.
Full reference solution
This is the complete worked program. Try building it yourself first, then use this reference to find the first place your version behaves differently. The Python below is generated from these exact blocks; helper functions are included so its behaviour can be inspected.
# ---------------------------------------------------------------# Building blocks, written out in plain Python.# This part is generated for you. Your script starts further down.# --------------------------------------------------------------- import mathimport random class Dataset: """A pile of labelled examples. Features can be a number, some text, or a list of numbers. The label is whatever answer you want back.""" def __init__(self, name="dataset"): self.name = name self.rows = [] # Column names from the file's header row, when it came from one. # Without these "the km of a row" has nothing to look the name up in. self.columns = [] def add(self, features, label, fields=None): self.rows.append( { "features": features, "label": str(label), "fields": dict(fields) if fields else {}, } ) def size(self): return len(self.rows) def __len__(self): return len(self.rows) def __iter__(self): """Walking a dataset gives you its rows, so "for each row in list [testing]" reads the way it sounds.""" return iter(self.rows) def labels(self): seen = [] for row in self.rows: if row["label"] not in seen: seen.append(row["label"]) return seen def most_common_label(self): if not self.rows: return "" counts = {} for row in self.rows: counts[row["label"]] = counts.get(row["label"], 0) + 1 return max(counts, key=lambda label: counts[label]) def split(self, train_percent=80): """Keeps the given percent for training and hands back the rest as a test set. The shuffle is seeded, so you get the same split every run.""" order = list(range(len(self.rows))) random.Random(0).shuffle(order) cut = int(len(order) * train_percent / 100) train = Dataset(self.name + " (train)") test = Dataset(self.name + " (test)") train.columns = list(self.columns) test.columns = list(self.columns) for position, index in enumerate(order): row = self.rows[index] target = train if position < cut else test target.add(row["features"], row["label"], row.get("fields")) return train, test def show(self, limit=10): print(self.name + ": " + str(len(self.rows)) + " examples") for row in self.rows[:limit]: print(" " + str(row["features"]) + " -> " + row["label"]) if len(self.rows) > limit: print(" ... and " + str(len(self.rows) - limit) + " more") def new_dataset(name="dataset"): return Dataset(name) def row_field(row, name): """One named piece of a row: "the km of this row", "the label of it". Names come from the header line of the csv. "label" always works, even on a file with no header, because every row has one.""" wanted = str(name).strip() if not isinstance(row, dict): print("That is not a row. Use this inside a for each over a dataset.") return "" if wanted.lower() == "label": return row.get("label", "") fields = row.get("fields") or {} if wanted in fields: return fields[wanted] # Header names are matched loosely, so "Rain" finds the "rain" column. for key in fields: if str(key).strip().lower() == wanted.lower(): return fields[key] known = ", ".join([str(k) for k in fields]) if fields else "none" print( "No column called " + wanted + " in this row. Columns here: " + known + "." ) return "" def load_csv(dataset, path, label_column=-1, has_header=True): """Reads a comma separated file into a dataset. Everything except the label column becomes the features, and anything that looks like a number is turned into one. This is how a file you uploaded becomes something you can train on.""" try: with open(path) as handle: rows = [line.rstrip("\n").rstrip("\r") for line in handle] except OSError: print("Could not find " + path + ". Check the name in the Files panel.") return dataset rows = [row for row in rows if row.strip()] # A file with no label column is a perfectly normal thing to load: it is # what a test set looks like before you have predicted anything. labelled = str(label_column).strip().lower() not in ("none", "", "no", "-") header = [] if has_header and rows: header = [cell.strip() for cell in rows[0].split(",")] rows = rows[1:] added = 0 for row in rows: cells = [cell.strip() for cell in row.split(",")] if not cells or (labelled and len(cells) < 2): continue index = -1 if labelled: index = int(label_column) if index < 0: index = len(cells) + index if index < 0 or index >= len(cells): continue label = cells[index] if labelled else "" keep = [i for i in range(len(cells)) if i != index] typed = [] for i in keep: try: typed.append(float(cells[i])) except ValueError: typed.append(cells[i]) # Every kept column gets its header name, so "the km of a row" works. fields = {} for position, i in enumerate(keep): if i < len(header) and header[i]: fields[header[i]] = typed[position] if header and not dataset.columns: dataset.columns = [header[i] for i in keep if i < len(header)] dataset.add(typed[0] if len(typed) == 1 else typed, label, fields) added += 1 print("Loaded " + str(added) + " rows from " + path + ".") return dataset def save_submission(predictions, dataset, path="submission.csv"): """Prints your predictions in the exact two column format a round is scored in: a header, then one id and one label per line. It is printed rather than saved to a file because the console is the one place you can copy it from. Paste it into a new file in the Files panel, or straight into the upload box.""" labels = list(predictions or []) rows = list(dataset) if dataset is not None else [] if len(labels) != len(rows): print( "You have " + str(len(labels)) + " predictions for " + str(len(rows)) + " rows. Those have to match before this means anything." ) return lines = ["id,label"] for position, row in enumerate(rows): # Looked up directly rather than through row_field, which would # complain on every row of a file that simply has no id column. fields = row.get("fields") or {} row_id = "" for key in fields: if str(key).strip().lower() == "id": row_id = fields[key] break if not str(row_id).strip(): row_id = "r" + str(position + 1).zfill(2) if isinstance(row_id, float) and row_id == int(row_id): row_id = int(row_id) lines.append(str(row_id) + "," + str(labels[position]).strip()) print("--- submission.csv, copy from here ---") for line in lines: print(line) print("--- to here, " + str(len(lines) - 1) + " rows ---") class Line: """y = slope * x + intercept. Two knobs, and learning is turning them.""" def __init__(self, slope=1.0, intercept=0.0, name="line"): self.slope = float(slope) self.intercept = float(intercept) self.name = name def predict(self, x): return self.slope * float(x) + self.intercept def _pairs(self, dataset): pairs = [] for row in dataset.rows: features = row["features"] if isinstance(features, (list, tuple)): features = features[0] pairs.append((float(features), float(row["label"]))) return pairs def loss(self, dataset): """Mean squared error: the average of how wrong it is, squared.""" pairs = self._pairs(dataset) if not pairs: return 0.0 total = 0.0 for x, y in pairs: error = self.predict(x) - y # e * e rather than e ** 2. On a run that is blowing up, the power # operator raises OverflowError and kills the script, while the # multiply quietly reaches inf so the chart can report it. total += error * error average = total / len(pairs) if average != average or average in (float("inf"), float("-inf")): return float("inf") return round(average, 4) def step(self, dataset, rate=0.01): """One nudge downhill. Look at how wrong the line is, then move both knobs a small amount in the direction that makes it less wrong.""" pairs = self._pairs(dataset) if not pairs: return slope_push = 0.0 intercept_push = 0.0 for x, y in pairs: error = self.predict(x) - y slope_push += 2 * error * x intercept_push += 2 * error self.slope -= rate * slope_push / len(pairs) self.intercept -= rate * intercept_push / len(pairs) def train(self, dataset, steps=200, rate=0.01): """Many nudges in a row, with a chart of the loss as it goes. The chart is the point. A run that bounces, a run that crawls and a run that drops and flattens are three different problems, and the final number alone tells them apart badly.""" count = int(steps) history = [self.loss(dataset)] for _ in range(count): self.step(dataset, rate) history.append(self.loss(dataset)) self._chart(history) print( "Trained " + self.name + " for " + str(count) + " steps. Loss is now " + str(self.loss(dataset)) ) def _chart(self, history): """A loss chart in text: one row per sampled step, bar length in proportion to the loss at that point.""" if len(history) < 2: return # At most a dozen rows. The first few steps are always shown, because # that is where a loss curve does nearly all of its moving, and a run # that has settled by step 3 would otherwise look like a flat line. wanted = 12 if len(history) <= wanted: picks = list(range(len(history))) else: early = list(range(min(4, len(history)))) spread = wanted - len(early) first = early[-1] gap = (len(history) - 1 - first) / float(spread) rest = [int(round(first + (i + 1) * gap)) for i in range(spread)] picks = sorted(set(early + rest)) picks[-1] = len(history) - 1 finite = [history[i] for i in picks if history[i] < float("inf")] top = max(finite) if finite else 0.0 width = len(str(len(history) - 1)) # Bars are drawn on a log scale. A first loss of 500 next to a final # loss of 0.2 would otherwise flatten every interesting bar to nothing, # and the shape of the drop is the whole reason to look at this. scale = math.log(1.0 + top) if top > 0 else 0.0 print("Loss per step:") for i in picks: value = history[i] step_label = str(i).rjust(width) if value == float("inf"): print(" step " + step_label + " " + "off the chart") continue bars = 0 if scale > 0: bars = int(round(30 * math.log(1.0 + value) / scale)) print( " step " + step_label + " " + ("#" * bars).ljust(30) + " " + str(round(value, 4)) ) def show(self): print( self.name + ": y = " + str(round(self.slope, 4)) + " * x + " + str(round(self.intercept, 4)) ) def new_line(slope=1.0, intercept=0.0, name="line"): return Line(slope, intercept, name) # ---------------------------------------------------------------# Your script# --------------------------------------------------------------- measurements = new_dataset("measurements")measurements.add(1, "12")measurements.add(2, "16")measurements.add(3, "21")measurements.add(4, "24")measurements.add(5, "29")measurements.add(6, "33")measurements.add(7, "36")measurements.add(8, "41")route = new_line(4, 8, "route")measurements.show()print("Within range at 4.5 km", route.predict(4.5))print("Beyond range at 20 km", route.predict(20))print("Training MSE", route.loss(measurements))Compare this with your version. Different names and personal choices are fine when the program follows the same logic.
Keep your progress
Sign in and every reading, quiz, and exercise you finish is saved.