0%
BuildUnder the hoodabout 28 min, 8 steps

Scoring a Wrong Line

Calculate residuals, squared errors, and their mean to compare two lines on the same examples.

Why count error at all?

This optional reading builds the error score that fitting a line depends on. The stages use the Lines blocks in this article's editor.

A line with one set of parameters makes different predictions from another line. A loss function turns the mistakes into one quantity the fitting algorithm can minimise. It does not choose the task for you; you must decide what error measure is appropriate.

Use three teaching examples: x values 1, 2, 3 with true y values 3, 5, 7. Compare a first candidate prediction = 2 × x. Define error as prediction minus actual.

xActualPredictionErrorSquared error
132-11
254-11
376-11

The mean squared error, abbreviated MSE, is (1 + 1 + 1) / 3 = 1. Mean means average: add the values and divide by their count. Squaring means multiplying a number by itself.

Build with me · 1

Create the three numerical targets

Each row pairs an input with an observed answer. The label values are stored by the teaching dataset and converted to numbers by the line-loss calculation. Their job is different from a category name such as late.

Keep all candidates on these same three observations so a lower loss reflects a different rule rather than an easier evaluation set.

Blocks at this stageWorked example
make a dataset calledmeasurements
add tomeasurementsthe example1labeled3
add tomeasurementsthe example2labeled5
add tomeasurementsthe example3labeled7
show datasetmeasurements

Add inputs 1, 2, 3 with targets 3, 5, 7 to measurements and show them.

What to look for

Three rows print with those exact pairs.

Make it yours

Calculate the simple rule that matches all three, but keep it aside while inspecting imperfect candidates.

Build with me · 2

Inspect predictions before summarising

Candidate A uses slope two and intercept zero. It predicts two, four, six, which are each one below the observed answers. Read every row before compressing the errors into a single score.

A scalar loss is useful for comparing and fitting, but intermediate predictions reveal whether a neat number describes the intended task.

Blocks at this stageWorked example
make a dataset calledmeasurements
add tomeasurementsthe example1labeled3
add tomeasurementsthe example2labeled5
add tomeasurementsthe example3labeled7
make a line calledroutewith slope2and intercept0
sayAt onethen
routepredicts for x =1
sayAt twothen
routepredicts for x =2
sayAt threethen
routepredicts for x =3

Create route with slope 2 and intercept 0, then print all three predictions.

What to look for

The outputs are 2, 4, 6.

Make it yours

Describe the shared offset. Which parameter could correct it without changing the per-unit increase?

Build with me · 3

Define which direction subtraction means

A residual is prediction minus actual in this lesson. Two minus three is negative one, so a negative residual indicates an underestimate. Another source may choose the opposite sign; define the convention before comparing results.

Squaring later removes the sign, but the signed value remains useful for finding systematic under- or over-prediction.

Blocks at this stageWorked example
make a dataset calledmeasurements
add tomeasurementsthe example1labeled3
add tomeasurementsthe example2labeled5
add tomeasurementsthe example3labeled7
make a line calledroutewith slope2and intercept0
seterrorto
routepredicts for x =1
-3
sayPrediction minus actualthen
error

Store route at one minus actual three under error and print it.

What to look for

Prediction minus actual is -1.0.

Make it yours

Reverse the subtraction and explain which interpretation of positive and negative must change.

Build with me · 4

Prevent opposite errors from cancelling

Squaring multiplies a value by itself. Negative one times negative one is positive one. An error of two would contribute four, while an error of ten contributes one hundred.

Large mistakes therefore receive more influence in mean squared error. That is a choice of loss, not a universal statement that every task should value errors this way.

Blocks at this stageWorked example
make a dataset calledmeasurements
add tomeasurementsthe example1labeled3
add tomeasurementsthe example2labeled5
add tomeasurementsthe example3labeled7
make a line calledroutewith slope2and intercept0
seterrorto
routepredicts for x =1
-3
sayResidualthen
error
saySquared residualthen
error
x
error

Put the error variable in both sockets of multiplication and print the result below the signed residual.

What to look for

Residual is -1 and Squared residual is 1.

Make it yours

Try a hypothetical residual of -2 and compare its square with the -1 case.

Build with me · 5

Add squared errors and divide by their count

All three squared residuals are one. Their sum is three, and dividing by three observations gives MSE one. Mean means average: the division is part of the definition.

A total would tend to grow with more rows even if their typical errors stayed the same. Compare candidates on the same rows, with the same units and error definition.

Blocks at this stageWorked example
make a dataset calledmeasurements
add tomeasurementsthe example1labeled3
add tomeasurementsthe example2labeled5
add tomeasurementsthe example3labeled7
make a line calledroutewith slope2and intercept0
sayManual MSEthen
1+1
+1
/3
sayLoss blockthen
the loss ofrouteonmeasurements

Build the nested sum and division, then print the line-loss block beside it as an independent check.

What to look for

Manual MSE and Loss block both show 1.

Make it yours

Explain what a reported value of 3 would suggest someone forgot in this calculation.

Work through a second candidate

A second candidate shows why the squaring step matters, and a third shows what a perfect training score does and does not prove.

Build with me · 6

Compare a second candidate carefully

Candidate B predicts two, five, eight. Its residuals are negative one, zero, positive one. Their signed average is zero because under- and over-prediction cancel. That does not mean every prediction is correct.

The squares are one, zero, one, whose mean is two thirds. Under MSE this candidate is better than A, despite the misleading zero signed average.

Blocks at this stageWorked example
make a dataset calledmeasurements
add tomeasurementsthe example1labeled3
add tomeasurementsthe example2labeled5
add tomeasurementsthe example3labeled7
make a line calledroutewith slope3and intercept-1
sayAt onethen
routepredicts for x =1
sayAt twothen
routepredicts for x =2
sayAt threethen
routepredicts for x =3
sayMean signed errorthen
-1+0
+1
/3
sayMSEthen
the loss ofrouteonmeasurements

Set slope 3 and intercept -1. Inspect predictions and compare signed mean with the loss block.

What to look for

Mean signed error is 0; MSE is approximately 0.6667.

Make it yours

Explain why zero average signed error and zero MSE are very different claims.

Build with me · 7

Recognise what zero training loss proves

Slope two and intercept one match the three targets exactly, so every squared residual is zero. The new input four produces nine by extending that rule.

Zero training loss proves the rule matches these supplied observations under this loss. It does not supply the unknown actual answer at four or prove that every future relationship is linear.

Blocks at this stageWorked example
make a dataset calledmeasurements
add tomeasurementsthe example1labeled3
add tomeasurementsthe example2labeled5
add tomeasurementsthe example3labeled7
make a line calledroutewith slope2and intercept1
sayTraining MSEthen
the loss ofrouteonmeasurements
sayNew input at fourthen
routepredicts for x =4

Change the intercept to 1, keep slope 2, and compare the loss with a new-input prediction.

What to look for

Training MSE is 0 and the new prediction is 9.

Make it yours

Invent a future observation that would disagree with nine. Explain why the old zero training loss would still have been correctly calculated.

What squaring changes

An error of two contributes four; an error of ten contributes one hundred. Large errors therefore affect MSE more strongly than small ones. That can be useful when large misses matter, but it can also make a few unusual or incorrectly labelled examples dominate the fit.

If outputs are minutes, squared errors are in minutes squared. The square root of MSE, called root mean squared error, returns to minutes. Do not interpret MSE 9 as an average signed error of nine minutes; its square root is three minutes, and even that is not the mean absolute error.

A loss of zero is possible only when every prediction matches its label under this measure. A low loss is relative to the output scale and task. Converting minutes to seconds changes the squared loss scale substantially without changing the physical predictions.

Calculate a whole score without skipping the averaging step

Try a separate two-example problem with true answers ten and ten. Model A predicts eight and twelve. Its errors are negative two and positive two, so the average signed error is zero. Its squared errors are four and four, and MSE is (4 + 4) / 2 = four.

Model B predicts ten and thirteen. Its squared errors are zero and nine, so MSE is 9 / 2 = 4.5. Under this loss, Model A wins despite getting neither answer exactly right. MSE cares about the size of numerical mistakes, not just a count of exact matches.

If you add a third example, divide the new squared-error total by three. Comparing totals from sets of different sizes is misleading: a longer set can accumulate more error simply because it has more rows. Even averages should be compared on the same evaluation examples when choosing between models.

You can audit a spreadsheet or block output using this tiny case. If it prints zero for Model A, it probably averaged signed errors. If it prints eight, it probably summed the squared errors without taking their mean. Intermediate values tell you which operation was missed.

Use it while fitting

The loss block in these stages, the loss of [line] on [examples], does this whole calculation for a line and a dataset. Change slope or intercept and compare the loss on the same examples, as the last stage does.

Training loss guides parameter updates. Development loss helps choose the model and training settings. Final-test loss estimates the selected model on reserved data. The formula is the same in each role, but the meaning of the evidence differs.

Before trusting a displayed average, calculate one small example by hand and check the count. A missing row, wrong label column, or accidental unit conversion can produce a neat loss number for the wrong task.

Build with me · 8

Make the comparison explicit in one run

This reference changes parameters while leaving the observations fixed. The order of the labelled outputs documents which rule each loss belongs to. This makes a comparison reproducible instead of relying on memory of several edited runs.

If outputs represent minutes, the MSE unit is minutes squared. A square root would return to minutes, but neither statistic becomes percentage accuracy. Name the metric and units when reporting it.

Blocks at this stageWorked example
make a dataset calledmeasurements
add tomeasurementsthe example1labeled3
add tomeasurementsthe example2labeled5
add tomeasurementsthe example3labeled7
make a line calledroutewith slope2and intercept0
sayA: slope 2, intercept 0then
the loss ofrouteonmeasurements
set theslopeofrouteto3
set theinterceptofrouteto-1
sayB: slope 3, intercept -1then
the loss ofrouteonmeasurements
set theslopeofrouteto2
set theinterceptofrouteto1
sayC: slope 2, intercept 1then
the loss ofrouteonmeasurements

Run all three candidates and check each labelled loss against the manual calculations.

What to look for

A is 1, B is about 0.6667, and C is 0 on the same three rows.

Make it yours

Choose a new imperfect line and calculate at least one residual by hand before trusting its displayed loss.

Full reference solution

This is the complete worked program. Try building it yourself first, then use this reference to find the first place your version behaves differently. The Python below is generated from these exact blocks; helper functions are included so its behaviour can be inspected.

make a dataset calledmeasurements
add tomeasurementsthe example1labeled3
add tomeasurementsthe example2labeled5
add tomeasurementsthe example3labeled7
make a line calledroutewith slope2and intercept0
sayA: slope 2, intercept 0then
the loss ofrouteonmeasurements
set theslopeofrouteto3
set theinterceptofrouteto-1
sayB: slope 3, intercept -1then
the loss ofrouteonmeasurements
set theslopeofrouteto2
set theinterceptofrouteto1
sayC: slope 2, intercept 1then
the loss ofrouteonmeasurements
PythonHover over a line to see an explanation
# ---------------------------------------------------------------# Building blocks, written out in plain Python.# This part is generated for you. Your script starts further down.# ---------------------------------------------------------------  import mathimport random  class Dataset:    """A pile of labelled examples. Features can be a number, some text, or a    list of numbers. The label is whatever answer you want back."""     def __init__(self, name="dataset"):        self.name = name        self.rows = []        # Column names from the file's header row, when it came from one.        # Without these "the km of a row" has nothing to look the name up in.        self.columns = []     def add(self, features, label, fields=None):        self.rows.append(            {                "features": features,                "label": str(label),                "fields": dict(fields) if fields else {},            }        )     def size(self):        return len(self.rows)     def __len__(self):        return len(self.rows)     def __iter__(self):        """Walking a dataset gives you its rows, so "for each row in list        [testing]" reads the way it sounds."""        return iter(self.rows)     def labels(self):        seen = []        for row in self.rows:            if row["label"] not in seen:                seen.append(row["label"])        return seen     def most_common_label(self):        if not self.rows:            return ""        counts = {}        for row in self.rows:            counts[row["label"]] = counts.get(row["label"], 0) + 1        return max(counts, key=lambda label: counts[label])     def split(self, train_percent=80):        """Keeps the given percent for training and hands back the rest as a        test set. The shuffle is seeded, so you get the same split every run."""        order = list(range(len(self.rows)))        random.Random(0).shuffle(order)        cut = int(len(order) * train_percent / 100)        train = Dataset(self.name + " (train)")        test = Dataset(self.name + " (test)")        train.columns = list(self.columns)        test.columns = list(self.columns)        for position, index in enumerate(order):            row = self.rows[index]            target = train if position < cut else test            target.add(row["features"], row["label"], row.get("fields"))        return train, test     def show(self, limit=10):        print(self.name + ": " + str(len(self.rows)) + " examples")        for row in self.rows[:limit]:            print("  " + str(row["features"]) + "  ->  " + row["label"])        if len(self.rows) > limit:            print("  ... and " + str(len(self.rows) - limit) + " more")  def new_dataset(name="dataset"):    return Dataset(name)  def row_field(row, name):    """One named piece of a row: "the km of this row", "the label of it".     Names come from the header line of the csv. "label" always works, even on    a file with no header, because every row has one."""    wanted = str(name).strip()    if not isinstance(row, dict):        print("That is not a row. Use this inside a for each over a dataset.")        return ""    if wanted.lower() == "label":        return row.get("label", "")    fields = row.get("fields") or {}    if wanted in fields:        return fields[wanted]    # Header names are matched loosely, so "Rain" finds the "rain" column.    for key in fields:        if str(key).strip().lower() == wanted.lower():            return fields[key]    known = ", ".join([str(k) for k in fields]) if fields else "none"    print(        "No column called "        + wanted        + " in this row. Columns here: "        + known        + "."    )    return ""  def load_csv(dataset, path, label_column=-1, has_header=True):    """Reads a comma separated file into a dataset.     Everything except the label column becomes the features, and anything that    looks like a number is turned into one. This is how a file you uploaded    becomes something you can train on."""    try:        with open(path) as handle:            rows = [line.rstrip("\n").rstrip("\r") for line in handle]    except OSError:        print("Could not find " + path + ". Check the name in the Files panel.")        return dataset     rows = [row for row in rows if row.strip()]     # A file with no label column is a perfectly normal thing to load: it is    # what a test set looks like before you have predicted anything.    labelled = str(label_column).strip().lower() not in ("none", "", "no", "-")     header = []    if has_header and rows:        header = [cell.strip() for cell in rows[0].split(",")]        rows = rows[1:]     added = 0    for row in rows:        cells = [cell.strip() for cell in row.split(",")]        if not cells or (labelled and len(cells) < 2):            continue         index = -1        if labelled:            index = int(label_column)            if index < 0:                index = len(cells) + index            if index < 0 or index >= len(cells):                continue         label = cells[index] if labelled else ""        keep = [i for i in range(len(cells)) if i != index]         typed = []        for i in keep:            try:                typed.append(float(cells[i]))            except ValueError:                typed.append(cells[i])         # Every kept column gets its header name, so "the km of a row" works.        fields = {}        for position, i in enumerate(keep):            if i < len(header) and header[i]:                fields[header[i]] = typed[position]        if header and not dataset.columns:            dataset.columns = [header[i] for i in keep if i < len(header)]         dataset.add(typed[0] if len(typed) == 1 else typed, label, fields)        added += 1     print("Loaded " + str(added) + " rows from " + path + ".")    return dataset  def save_submission(predictions, dataset, path="submission.csv"):    """Prints your predictions in the exact two column format a round is    scored in: a header, then one id and one label per line.     It is printed rather than saved to a file because the console is the one    place you can copy it from. Paste it into a new file in the Files panel,    or straight into the upload box."""    labels = list(predictions or [])    rows = list(dataset) if dataset is not None else []     if len(labels) != len(rows):        print(            "You have "            + str(len(labels))            + " predictions for "            + str(len(rows))            + " rows. Those have to match before this means anything."        )        return     lines = ["id,label"]    for position, row in enumerate(rows):        # Looked up directly rather than through row_field, which would        # complain on every row of a file that simply has no id column.        fields = row.get("fields") or {}        row_id = ""        for key in fields:            if str(key).strip().lower() == "id":                row_id = fields[key]                break        if not str(row_id).strip():            row_id = "r" + str(position + 1).zfill(2)        if isinstance(row_id, float) and row_id == int(row_id):            row_id = int(row_id)        lines.append(str(row_id) + "," + str(labels[position]).strip())     print("--- submission.csv, copy from here ---")    for line in lines:        print(line)    print("--- to here, " + str(len(lines) - 1) + " rows ---")  class Line:    """y = slope * x + intercept. Two knobs, and learning is turning them."""     def __init__(self, slope=1.0, intercept=0.0, name="line"):        self.slope = float(slope)        self.intercept = float(intercept)        self.name = name     def predict(self, x):        return self.slope * float(x) + self.intercept     def _pairs(self, dataset):        pairs = []        for row in dataset.rows:            features = row["features"]            if isinstance(features, (list, tuple)):                features = features[0]            pairs.append((float(features), float(row["label"])))        return pairs     def loss(self, dataset):        """Mean squared error: the average of how wrong it is, squared."""        pairs = self._pairs(dataset)        if not pairs:            return 0.0        total = 0.0        for x, y in pairs:            error = self.predict(x) - y            # e * e rather than e ** 2. On a run that is blowing up, the power            # operator raises OverflowError and kills the script, while the            # multiply quietly reaches inf so the chart can report it.            total += error * error        average = total / len(pairs)        if average != average or average in (float("inf"), float("-inf")):            return float("inf")        return round(average, 4)     def step(self, dataset, rate=0.01):        """One nudge downhill. Look at how wrong the line is, then move both        knobs a small amount in the direction that makes it less wrong."""        pairs = self._pairs(dataset)        if not pairs:            return        slope_push = 0.0        intercept_push = 0.0        for x, y in pairs:            error = self.predict(x) - y            slope_push += 2 * error * x            intercept_push += 2 * error        self.slope -= rate * slope_push / len(pairs)        self.intercept -= rate * intercept_push / len(pairs)     def train(self, dataset, steps=200, rate=0.01):        """Many nudges in a row, with a chart of the loss as it goes.         The chart is the point. A run that bounces, a run that crawls and a        run that drops and flattens are three different problems, and the        final number alone tells them apart badly."""        count = int(steps)        history = [self.loss(dataset)]        for _ in range(count):            self.step(dataset, rate)            history.append(self.loss(dataset))         self._chart(history)        print(            "Trained "            + self.name            + " for "            + str(count)            + " steps. Loss is now "            + str(self.loss(dataset))        )     def _chart(self, history):        """A loss chart in text: one row per sampled step, bar length in        proportion to the loss at that point."""        if len(history) < 2:            return         # At most a dozen rows. The first few steps are always shown, because        # that is where a loss curve does nearly all of its moving, and a run        # that has settled by step 3 would otherwise look like a flat line.        wanted = 12        if len(history) <= wanted:            picks = list(range(len(history)))        else:            early = list(range(min(4, len(history))))            spread = wanted - len(early)            first = early[-1]            gap = (len(history) - 1 - first) / float(spread)            rest = [int(round(first + (i + 1) * gap)) for i in range(spread)]            picks = sorted(set(early + rest))            picks[-1] = len(history) - 1         finite = [history[i] for i in picks if history[i] < float("inf")]        top = max(finite) if finite else 0.0        width = len(str(len(history) - 1))         # Bars are drawn on a log scale. A first loss of 500 next to a final        # loss of 0.2 would otherwise flatten every interesting bar to nothing,        # and the shape of the drop is the whole reason to look at this.        scale = math.log(1.0 + top) if top > 0 else 0.0         print("Loss per step:")        for i in picks:            value = history[i]            step_label = str(i).rjust(width)            if value == float("inf"):                print("  step " + step_label + "  " + "off the chart")                continue            bars = 0            if scale > 0:                bars = int(round(30 * math.log(1.0 + value) / scale))            print(                "  step "                + step_label                + "  "                + ("#" * bars).ljust(30)                + " "                + str(round(value, 4))            )     def show(self):        print(            self.name            + ": y = "            + str(round(self.slope, 4))            + " * x + "            + str(round(self.intercept, 4))        )  def new_line(slope=1.0, intercept=0.0, name="line"):    return Line(slope, intercept, name)  # ---------------------------------------------------------------# Your script# ---------------------------------------------------------------  measurements = new_dataset("measurements")measurements.add(1, "3")measurements.add(2, "5")measurements.add(3, "7")route = new_line(2, 0, "route")print("A: slope 2, intercept 0", route.loss(measurements))route.slope = float(3)route.intercept = float(-1)print("B: slope 3, intercept -1", route.loss(measurements))route.slope = float(2)route.intercept = float(1)print("C: slope 2, intercept 1", route.loss(measurements))

Compare this with your version. Different names and personal choices are fine when the program follows the same logic.

Keep your progress

Sign in and every reading, quiz, and exercise you finish is saved.

Sign in