Predicting From a Trend
Make numerical predictions from a fitted line while separating interpolation, extrapolation, and uncertainty.
Evaluate the fitted rule
This optional reading is about using a fitted line to predict. The stages use the Lines blocks: in this article's editor, choose the Lines group in Add blocks.
Assume a training process selected predicted minutes = 4 × km + 8. For a trip of 6.5 kilometres, multiply four by 6.5 to get 26, then add eight to predict 34 minutes. The prediction can be a decimal even if your training table used whole numbers.
Using a rule you already have on a new input is prediction, also called inference. You have not changed the slope or intercept, and you have not supplied the actual time. The model can predict before the delivery happens. Once the actual time is known, you can calculate an error.
If the delivery took 37 minutes, prediction minus actual is 34 - 37 = -3. The negative sign means an underestimate under this convention. State your convention because actual-minus-prediction would reverse the sign.
Build with me · 1
Create a specific rule
Slope four and intercept eight define one line. The fixed form is multiply input by slope, then add intercept. The two numbers are parameters, while distance is an input supplied for each prediction.
We chose these parameters by hand. No training has occurred yet, and a plausible equation is not proof that distance explains every delivery delay.
Find Lines in Add blocks. Create route with slope 4 and intercept 8, then add show line route.
What to look for
The equation shows a slope of 4 and intercept of 8.
Make it yours
Name the units: minutes per kilometre for slope and minutes for intercept.
Build with me · 2
Substitute a distance into the rule
For three kilometres, four times three is twelve, plus eight is twenty. The rounded prediction value computes a number but does not display it until placed inside an output statement.
Predicting uses the current parameters without changing them. This is the forward calculation, whether the parameters were chosen manually or learned from examples.
Put route predicts for x = 3 in a labelled say socket below the declaration.
What to look for
Minutes for 3 km is 20.0.
Make it yours
Choose a distance between one and eight and calculate its result before running.
Inside and outside the data range
Suppose training distances covered one to eight kilometres. Predicting at 6.5 is interpolation: the input lies within that observed range. Predicting at twenty is extrapolation: it lies outside.
Interpolation often asks less of the model, but it is not automatically trustworthy. The training points may be sparse near 6.5, labels may be noisy, or an important condition may have changed. A new road closure is not represented by distance alone.
At twenty kilometres, the line predicts 88 minutes. The arithmetic is valid, but whether the estimate is useful depends on whether the same relationship extends that far. Longer routes may use faster roads or have different preparation costs. Correct arithmetic does not validate a modelling assumption.
Build with me · 3
Check what the intercept means
When input is zero, the slope term is zero, so the prediction equals the intercept. At one, the prediction includes one slope-sized increment. The values are eight and twelve in this example.
You may interpret eight as a preparation-time approximation, but that is a hypothesis about the world. If no zero-distance journey was observed, the intercept is an implied model value rather than a measured fact.
Print route at zero and one without changing its parameters.
What to look for
At zero is 8.0 and At one is 12.0.
Make it yours
Explain the difference between the equation's starting value and a claim about an actual preparation process.
Build with me · 4
Measure a change in output per input unit
Going from two to five kilometres adds three kilometres. At four minutes per kilometre, the predicted time grows by twelve minutes. The intercept is present in both predictions and cancels when we subtract them.
Slope describes change, not the height of a point on the graph. A high intercept can create a high prediction even with a zero slope.
Print the two predictions and nest their difference in another output block.
What to look for
The predictions are 16.0 and 28.0; their difference is 12.0.
Make it yours
Change slope to zero. Predict both outputs and explain why they become equal.
Avoid false precision
A model might return 34.017 minutes because its fitted parameters have many decimal places. That does not mean it knows the arrival time to a thousandth of a minute. Numerical precision and predictive certainty are different.
Inspect held-out errors to understand the scale of mistakes under conditions you tested. If comparable predictions often miss by several minutes, report a point estimate with that context. Do not invent a guaranteed interval by adding a convenient percentage around the prediction.
Separate association from intervention
A fitted relationship says how inputs and outcomes were associated in the collected data. It does not prove changing one input causes exactly the predicted output change. Longer routes might also differ in traffic, driver experience, or scheduling.
For a forecasting demo, association can still be useful. For a decision such as redesigning routes, you need evidence about the consequences of the intervention. State what your model was built to do.
The next stages turn the line's knobs by hand. They are a reminder that the numbers in a fitted rule came from particular data: different examples would have produced a different slope and intercept, and so different predictions.
Build with me · 5
Shift every prediction by the same amount
Changing intercept from eight to ten adds two to every prediction. It does not change how much the model grows per kilometre. This can address a consistent offset in a controlled example.
The set-knob statement changes the existing line after the before outputs. Reading the stack in order lets you compare versions within one run rather than remembering results from separate runs.
Insert set intercept of route to 10 between the before and after outputs. Leave slope fixed.
What to look for
At one changes from 12 to 14; at five changes from 28 to 30.
Make it yours
Try intercept six and predict the common shift. Restore eight for the next stage.
Build with me · 6
Make the effect depend on distance
Increasing slope by one adds one extra minute for each kilometre. Compared with the original rule, the one-kilometre prediction grows by one and the five-kilometre prediction by five.
A slope adjustment and an intercept adjustment therefore have different signatures across the dataset. Inspect multiple points rather than choosing a parameter change from one lucky example.
Set slope to 5 while keeping intercept 8, then print predictions at one and five.
What to look for
At one is 13.0 and At five is 33.0.
Make it yours
Choose a negative slope and explain the resulting direction of change. A computable line is not necessarily plausible for the intended task.
Build with me · 7
Preserve the rule when units change
Two kilometres is two thousand metres. To preserve the same physical predictions, slope becomes 0.004 minutes per metre instead of four minutes per kilometre. Intercept stays eight minutes.
Changing the input unit without changing the matching parameter creates a different model. Units are part of the meaning of a number, even though the computer can multiply mismatched numbers without recognising your mistake.
Create the metres version with slope 0.004 and feed it 2000 and 5000.
What to look for
The predictions remain 16.0 and 28.0, matching two and five kilometres.
Make it yours
Calculate what the incorrect slope four would produce for 2000, then explain why the huge value is a units error.
A useful prediction statement
A negative distance would also produce a number if you inserted it into the formula. That does not make the input meaningful. A surrounding program should reject impossible values or ask for clarification before calling the model.
For a trip of 3.5 kilometres, 4 × 3.5 + 8 = 22, so say “The model predicts twenty-two minutes, using a line fitted to deliveries between one and eight kilometres.” Do not invent a plus-or-minus range. To estimate uncertainty, you need an appropriate method and evidence about errors on relevant new examples. Reporting the model's scope is already more useful than presenting an unexplained number as a promise.
A clear report is: "The line predicts 34 minutes for 6.5 kilometres. It was fitted on trips from one to eight kilometres, using distance alone. Weather and route differences are not included."
That explains the number, range, and missing information. Practise making the same statement for a prediction beyond eight kilometres and explicitly call it extrapolation. You should be able to explain why the calculation still runs even when the evidence for using it becomes weaker.
Build with me · 8
Keep observations beside the prediction rule
The observed distances span one to eight kilometres. A prediction at 4.5 lies between observed inputs; twenty lies beyond them. Both can be calculated, but the second extends the trend into a region without supporting observations.
The loss compares this hand-chosen line with all eight observed answers. It does not prove causation or cover omitted features such as traffic. State the input range, units, parameters, and limitations together.
Create the eight measurements, show them, and compare the two predictions with their position relative to the observed range.
What to look for
At 4.5 km the model predicts 26 minutes; at 20 km it predicts 88. Training MSE for the supplied observations is 0.5.
Make it yours
Explain why the farther prediction does not become trustworthy merely because the arithmetic is exact.
Full reference solution
This is the complete worked program. Try building it yourself first, then use this reference to find the first place your version behaves differently. The Python below is generated from these exact blocks; helper functions are included so its behaviour can be inspected.
# ---------------------------------------------------------------# Building blocks, written out in plain Python.# This part is generated for you. Your script starts further down.# --------------------------------------------------------------- import mathimport random class Dataset: """A pile of labelled examples. Features can be a number, some text, or a list of numbers. The label is whatever answer you want back.""" def __init__(self, name="dataset"): self.name = name self.rows = [] # Column names from the file's header row, when it came from one. # Without these "the km of a row" has nothing to look the name up in. self.columns = [] def add(self, features, label, fields=None): self.rows.append( { "features": features, "label": str(label), "fields": dict(fields) if fields else {}, } ) def size(self): return len(self.rows) def __len__(self): return len(self.rows) def __iter__(self): """Walking a dataset gives you its rows, so "for each row in list [testing]" reads the way it sounds.""" return iter(self.rows) def labels(self): seen = [] for row in self.rows: if row["label"] not in seen: seen.append(row["label"]) return seen def most_common_label(self): if not self.rows: return "" counts = {} for row in self.rows: counts[row["label"]] = counts.get(row["label"], 0) + 1 return max(counts, key=lambda label: counts[label]) def split(self, train_percent=80): """Keeps the given percent for training and hands back the rest as a test set. The shuffle is seeded, so you get the same split every run.""" order = list(range(len(self.rows))) random.Random(0).shuffle(order) cut = int(len(order) * train_percent / 100) train = Dataset(self.name + " (train)") test = Dataset(self.name + " (test)") train.columns = list(self.columns) test.columns = list(self.columns) for position, index in enumerate(order): row = self.rows[index] target = train if position < cut else test target.add(row["features"], row["label"], row.get("fields")) return train, test def show(self, limit=10): print(self.name + ": " + str(len(self.rows)) + " examples") for row in self.rows[:limit]: print(" " + str(row["features"]) + " -> " + row["label"]) if len(self.rows) > limit: print(" ... and " + str(len(self.rows) - limit) + " more") def new_dataset(name="dataset"): return Dataset(name) def row_field(row, name): """One named piece of a row: "the km of this row", "the label of it". Names come from the header line of the csv. "label" always works, even on a file with no header, because every row has one.""" wanted = str(name).strip() if not isinstance(row, dict): print("That is not a row. Use this inside a for each over a dataset.") return "" if wanted.lower() == "label": return row.get("label", "") fields = row.get("fields") or {} if wanted in fields: return fields[wanted] # Header names are matched loosely, so "Rain" finds the "rain" column. for key in fields: if str(key).strip().lower() == wanted.lower(): return fields[key] known = ", ".join([str(k) for k in fields]) if fields else "none" print( "No column called " + wanted + " in this row. Columns here: " + known + "." ) return "" def load_csv(dataset, path, label_column=-1, has_header=True): """Reads a comma separated file into a dataset. Everything except the label column becomes the features, and anything that looks like a number is turned into one. This is how a file you uploaded becomes something you can train on.""" try: with open(path) as handle: rows = [line.rstrip("\n").rstrip("\r") for line in handle] except OSError: print("Could not find " + path + ". Check the name in the Files panel.") return dataset rows = [row for row in rows if row.strip()] # A file with no label column is a perfectly normal thing to load: it is # what a test set looks like before you have predicted anything. labelled = str(label_column).strip().lower() not in ("none", "", "no", "-") header = [] if has_header and rows: header = [cell.strip() for cell in rows[0].split(",")] rows = rows[1:] added = 0 for row in rows: cells = [cell.strip() for cell in row.split(",")] if not cells or (labelled and len(cells) < 2): continue index = -1 if labelled: index = int(label_column) if index < 0: index = len(cells) + index if index < 0 or index >= len(cells): continue label = cells[index] if labelled else "" keep = [i for i in range(len(cells)) if i != index] typed = [] for i in keep: try: typed.append(float(cells[i])) except ValueError: typed.append(cells[i]) # Every kept column gets its header name, so "the km of a row" works. fields = {} for position, i in enumerate(keep): if i < len(header) and header[i]: fields[header[i]] = typed[position] if header and not dataset.columns: dataset.columns = [header[i] for i in keep if i < len(header)] dataset.add(typed[0] if len(typed) == 1 else typed, label, fields) added += 1 print("Loaded " + str(added) + " rows from " + path + ".") return dataset def save_submission(predictions, dataset, path="submission.csv"): """Prints your predictions in the exact two column format a round is scored in: a header, then one id and one label per line. It is printed rather than saved to a file because the console is the one place you can copy it from. Paste it into a new file in the Files panel, or straight into the upload box.""" labels = list(predictions or []) rows = list(dataset) if dataset is not None else [] if len(labels) != len(rows): print( "You have " + str(len(labels)) + " predictions for " + str(len(rows)) + " rows. Those have to match before this means anything." ) return lines = ["id,label"] for position, row in enumerate(rows): # Looked up directly rather than through row_field, which would # complain on every row of a file that simply has no id column. fields = row.get("fields") or {} row_id = "" for key in fields: if str(key).strip().lower() == "id": row_id = fields[key] break if not str(row_id).strip(): row_id = "r" + str(position + 1).zfill(2) if isinstance(row_id, float) and row_id == int(row_id): row_id = int(row_id) lines.append(str(row_id) + "," + str(labels[position]).strip()) print("--- submission.csv, copy from here ---") for line in lines: print(line) print("--- to here, " + str(len(lines) - 1) + " rows ---") class Line: """y = slope * x + intercept. Two knobs, and learning is turning them.""" def __init__(self, slope=1.0, intercept=0.0, name="line"): self.slope = float(slope) self.intercept = float(intercept) self.name = name def predict(self, x): return self.slope * float(x) + self.intercept def _pairs(self, dataset): pairs = [] for row in dataset.rows: features = row["features"] if isinstance(features, (list, tuple)): features = features[0] pairs.append((float(features), float(row["label"]))) return pairs def loss(self, dataset): """Mean squared error: the average of how wrong it is, squared.""" pairs = self._pairs(dataset) if not pairs: return 0.0 total = 0.0 for x, y in pairs: error = self.predict(x) - y # e * e rather than e ** 2. On a run that is blowing up, the power # operator raises OverflowError and kills the script, while the # multiply quietly reaches inf so the chart can report it. total += error * error average = total / len(pairs) if average != average or average in (float("inf"), float("-inf")): return float("inf") return round(average, 4) def step(self, dataset, rate=0.01): """One nudge downhill. Look at how wrong the line is, then move both knobs a small amount in the direction that makes it less wrong.""" pairs = self._pairs(dataset) if not pairs: return slope_push = 0.0 intercept_push = 0.0 for x, y in pairs: error = self.predict(x) - y slope_push += 2 * error * x intercept_push += 2 * error self.slope -= rate * slope_push / len(pairs) self.intercept -= rate * intercept_push / len(pairs) def train(self, dataset, steps=200, rate=0.01): """Many nudges in a row, with a chart of the loss as it goes. The chart is the point. A run that bounces, a run that crawls and a run that drops and flattens are three different problems, and the final number alone tells them apart badly.""" count = int(steps) history = [self.loss(dataset)] for _ in range(count): self.step(dataset, rate) history.append(self.loss(dataset)) self._chart(history) print( "Trained " + self.name + " for " + str(count) + " steps. Loss is now " + str(self.loss(dataset)) ) def _chart(self, history): """A loss chart in text: one row per sampled step, bar length in proportion to the loss at that point.""" if len(history) < 2: return # At most a dozen rows. The first few steps are always shown, because # that is where a loss curve does nearly all of its moving, and a run # that has settled by step 3 would otherwise look like a flat line. wanted = 12 if len(history) <= wanted: picks = list(range(len(history))) else: early = list(range(min(4, len(history)))) spread = wanted - len(early) first = early[-1] gap = (len(history) - 1 - first) / float(spread) rest = [int(round(first + (i + 1) * gap)) for i in range(spread)] picks = sorted(set(early + rest)) picks[-1] = len(history) - 1 finite = [history[i] for i in picks if history[i] < float("inf")] top = max(finite) if finite else 0.0 width = len(str(len(history) - 1)) # Bars are drawn on a log scale. A first loss of 500 next to a final # loss of 0.2 would otherwise flatten every interesting bar to nothing, # and the shape of the drop is the whole reason to look at this. scale = math.log(1.0 + top) if top > 0 else 0.0 print("Loss per step:") for i in picks: value = history[i] step_label = str(i).rjust(width) if value == float("inf"): print(" step " + step_label + " " + "off the chart") continue bars = 0 if scale > 0: bars = int(round(30 * math.log(1.0 + value) / scale)) print( " step " + step_label + " " + ("#" * bars).ljust(30) + " " + str(round(value, 4)) ) def show(self): print( self.name + ": y = " + str(round(self.slope, 4)) + " * x + " + str(round(self.intercept, 4)) ) def new_line(slope=1.0, intercept=0.0, name="line"): return Line(slope, intercept, name) # ---------------------------------------------------------------# Your script# --------------------------------------------------------------- measurements = new_dataset("measurements")measurements.add(1, "12")measurements.add(2, "16")measurements.add(3, "21")measurements.add(4, "24")measurements.add(5, "29")measurements.add(6, "33")measurements.add(7, "36")measurements.add(8, "41")route = new_line(4, 8, "route")measurements.show()print("Within range at 4.5 km", route.predict(4.5))print("Beyond range at 20 km", route.predict(20))print("Training MSE", route.loss(measurements))Compare this with your version. Different names and personal choices are fine when the program follows the same logic.
Keep your progress
Sign in and every reading, quiz, and exercise you finish is saved.