0%
BuildUnder the hoodabout 29 min, 8 steps

Your First Spreadsheet Model

Use a small table and scatter chart to understand a numerical trend before relying on a fitted line.

Build the table here, then recognise spreadsheet notation

The practice editor beside this article performs the complete calculation in your browser. You do not need to open a spreadsheet application. We also explain the equivalent cell formulas so that you can transfer the same reasoning to a table tool later. A cell address combines a column letter and row number: A1 means column A, row one.

Our eight paired observations are distances 1, 2, 3, 4, 5, 6, 7, 8 kilometres and times 12, 16, 21, 24, 29, 33, 36, 41 minutes. A row describes one delivery; its two values must stay together when reordered. In a spreadsheet, km in A1 and minutes in B1 are headers, so the first observation occupies A2:B2 and the last A9:B9.

Build with me · 1

Keep each measured pair together

Each row pairs kilometres with observed minutes. The eight distances run from one to eight, while measured times are twelve, sixteen, twenty-one, twenty-four, twenty-nine, thirty-three, thirty-six, and forty-one.

Moving only a distance column would break those pairings. In a spreadsheet, sort the whole table together. In this program, the paired dataset gives us a way to check the separate lists used for the visible row-by-row calculation.

Blocks at this stageWorked example
make a dataset calledmeasurements
add tomeasurementsthe example1labeled12
add tomeasurementsthe example2labeled16
add tomeasurementsthe example3labeled21
add tomeasurementsthe example4labeled24
add tomeasurementsthe example5labeled29
add tomeasurementsthe example6labeled33
add tomeasurementsthe example7labeled36
add tomeasurementsthe example8labeled41
show datasetmeasurements

Inspect the dataset before adding a line. Read the third pair aloud as three kilometres and twenty-one minutes.

What to look for

Eight paired observations print.

Make it yours

Explain why sorting one list independently from the other would create false observations.

Read the trend without confusing a chart with evidence

A scatter chart positions each pair using its numeric x and y values. Distance belongs on the horizontal axis and minutes on the vertical axis. The observation (3, 21) belongs at x = 3, y = 21. A category chart that spaces every label equally is a different representation. Axis titles and units prevent a plausible-looking chart from hiding swapped meanings.

The observations rise approximately along a line, but not all points lie on it. A trendline is a fitted model, not a proof that distance causes every change in delivery time. Traffic, weather, routes, and stops are missing from this one-input example. A chart's visual angle also changes when its axis scales change; read numerical values rather than estimating slope only by appearance.

Create a prediction column

The hand-chosen rule is predicted minutes = 4 × kilometres + 8. In the editor, a line block and a loop calculate it for every row. In spreadsheet row two, the equivalent formula is =4*A2+8.

A copied spreadsheet formula normally moves a relative reference: A2 becomes A3 for the next observation. The equals sign begins a formula and the asterisk means multiplication. If a formula displays as text, a text-only cell format or leading apostrophe may be responsible. If all predictions are identical, inspect whether the input reference was fixed accidentally.

Build with me · 2

Calculate the proposed line for every input

The proposed line is four times distance plus eight. Its predictions are twelve, sixteen, twenty, twenty-four, twenty-eight, thirty-two, thirty-six, and forty. These form the prediction column of a table.

In a spreadsheet with distance in A2, the same first-row formula is =4*A2+8. The block loop performs the same repeated calculation while the current distance changes.

Blocks at this stageWorked example
make a line calledroutewith slope4and intercept8
make an empty list calleddistances
add1todistances
add2todistances
add3todistances
add4todistances
add5todistances
add6todistances
add7todistances
add8todistances
for eachdistancein listdistances
sayDistancethen
distance
sayPredicted minutesthen
routepredicts for x =
distance

Create route and loop over the distance list, printing distance beside its prediction.

What to look for

Eight predictions match the stated sequence.

Make it yours

Inspect the third prediction, twenty, beside the actual third value, twenty-one. Name the one-minute underestimate.

Rather than hard-code four and eight into every formula, give the parameters fixed locations. If slope is G2 and intercept is H2, the formula becomes =$G$2*A2+$H$2. The dollar signs keep those parameter references fixed while A2 changes by row. The line object's slope and intercept play the same role in blocks: change one stored parameter and later predictions use the new value.

Build with me · 3

Change one stored parameter for every row

A model parameter should have one authoritative location. Changing intercept from eight to nine updates every later prediction from route. In spreadsheet notation, fixed parameter cells such as $G$2 and $H$2 serve the same purpose.

Dollar signs keep a reference fixed when a formula is copied down. The current row's input reference changes while the slope and intercept locations stay the same.

Blocks at this stageWorked example
make a line calledroutewith slope4and intercept8
sayBefore firstthen
routepredicts for x =1
sayBefore lastthen
routepredicts for x =8
set theinterceptofrouteto9
sayAfter firstthen
routepredicts for x =1
sayAfter lastthen
routepredicts for x =8

Compare predictions before and after one set-intercept statement. Keep the slope unchanged.

What to look for

The first and last predictions each increase by one, from 12 and 40 to 13 and 41.

Make it yours

Explain why editing eight hard-coded formulas separately is more error-prone than changing one parameter location.

Residuals and squares, in cells and blocks

Define residual as prediction minus actual. At distance three, prediction twenty minus actual twenty-one gives negative one. Squaring means multiplying that residual by itself, producing positive one. A negative residual identifies an underestimate; a positive residual identifies an overestimate under this sign convention.

In a spreadsheet, C2 could contain prediction, D2 =C2-B2, and E2 =D2^2. The caret means exponentiation in that formula language. Copy the formulas through row nine while keeping the paired input and actual columns aligned. In the block program, separate variables expose the same intermediate values before adding each squared residual to a running total.

Build with me · 4

Trace prediction, residual, and square

For distance three, prediction is twenty and actual is twenty-one. Residual means prediction minus actual, so it is negative one. Its square is one. Keeping those as separate columns or variables reveals where a mistake first appears.

In spreadsheet row four, because row one is a header, these operations could be C4 for prediction, D4 = C4-B4, and E4 = D4^2. The notation differs, while the numerical steps stay identical.

Blocks at this stageWorked example
make a line calledroutewith slope4and intercept8
setactualto21
setpredictedto
routepredicts for x =3
seterrorto
predicted
-
actual
sayPredictionthen
predicted
sayResidualthen
error
saySquarethen
error
x
error

Store the three intermediate values and print them in order. Read the subtraction direction carefully.

What to look for

Prediction is 20, Residual is -1, and Square is 1.

Make it yours

Imagine a prediction of twenty-three instead. Calculate its residual and square before changing the source.

Build with me · 5

Repeat the calculation without losing alignment

The loop walks distances while row selects the matching actual value. Both begin with their first observation. Every turn calculates the prediction, residual, and square, then adds that square to a running total.

Create row and squared_total before the loop. Resetting the total inside would keep only the current row's error. Increment row after each calculation so the next distance reads the next actual value.

Blocks at this stageWorked example
make a dataset calledmeasurements
add tomeasurementsthe example1labeled12
add tomeasurementsthe example2labeled16
add tomeasurementsthe example3labeled21
add tomeasurementsthe example4labeled24
add tomeasurementsthe example5labeled29
add tomeasurementsthe example6labeled33
add tomeasurementsthe example7labeled36
add tomeasurementsthe example8labeled41
make a line calledroutewith slope4and intercept8
make an empty list calleddistances
add1todistances
add2todistances
add3todistances
add4todistances
add5todistances
add6todistances
add7todistances
add8todistances
make an empty list calledactuals
add12toactuals
add16toactuals
add21toactuals
add24toactuals
add29toactuals
add33toactuals
add36toactuals
add41toactuals
setrowto1
setsquared_totalto0
for eachdistancein listdistances
setactualto
item
row
ofactuals
setpredictedto
routepredicts for x =
distance
seterrorto
predicted
-
actual
setsquaredto
error
x
error
changesquared_totalby
squared
sayDistancethen
distance
sayActualthen
actual
sayPredictionthen
predicted
sayResidualthen
round
error
to3places
saySquared residualthen
round
squared
to3places
changerowby1

Build the indexed actual lookup and arithmetic inside for each, with both initial values above it. Inspect the printed groups.

What to look for

Residuals are 0, 0, -1, 0, -1, -1, 0, -1, and their squared sum is four.

Make it yours

Choose one middle row and match all five printed values with its source pair.

In a spreadsheet, the average of the squared-error column is =AVERAGE(E2:E9), where the colon includes every cell from E2 through E9. Do not include the header, leave out the last row, or average the signed residuals instead.

Build with me · 6

Average the completed squared-error column

The squared total is four across eight observations. Dividing four by eight gives MSE 0.5. This is a training loss in minutes squared, not 50% accuracy or a mean signed error.

The independent line-loss block calculates the same metric from the paired dataset. Matching the two implementations is a useful check that row alignment and the denominator were correct.

Blocks at this stageWorked example
make a dataset calledmeasurements
add tomeasurementsthe example1labeled12
add tomeasurementsthe example2labeled16
add tomeasurementsthe example3labeled21
add tomeasurementsthe example4labeled24
add tomeasurementsthe example5labeled29
add tomeasurementsthe example6labeled33
add tomeasurementsthe example7labeled36
add tomeasurementsthe example8labeled41
make a line calledroutewith slope4and intercept8
make an empty list calleddistances
add1todistances
add2todistances
add3todistances
add4todistances
add5todistances
add6todistances
add7todistances
add8todistances
make an empty list calledactuals
add12toactuals
add16toactuals
add21toactuals
add24toactuals
add29toactuals
add33toactuals
add36toactuals
add41toactuals
setrowto1
setsquared_totalto0
for eachdistancein listdistances
setactualto
item
row
ofactuals
setpredictedto
routepredicts for x =
distance
seterrorto
predicted
-
actual
setsquaredto
error
x
error
changesquared_totalby
squared
sayDistancethen
distance
sayActualthen
actual
sayPredictionthen
predicted
sayResidualthen
round
error
to3places
saySquared residualthen
round
squared
to3places
changerowby1
sayManual mean squared errorthen
round
squared_total
/8
to3places
sayLine loss checkthen
the loss ofrouteonmeasurements

Print squared_total divided by eight below the loop, then print route's loss on measurements.

What to look for

Both the manual average and line-loss check are 0.5.

Make it yours

Explain what forgetting the division would produce and what averaging signed residuals instead would measure.

Compare one changed parameter on the same rows

The next stage changes only the slope, from 4 to 4.1, and compares the loss on the same eight rows. Check one middle row by hand before trusting the summary.

Build with me · 7

Inspect why a new setting improves this loss

With slope 4.1, predictions are 12.1, 16.2, 20.3, 24.4, 28.5, 32.6, 36.7, and 40.8. The squared errors sum to 1.64, so the MSE is 1.64 divided by eight, or 0.205.

Small floating-point display differences are normal. Round for readability after calculating, not by replacing each intermediate residual with a whole number, which would alter the loss itself.

Blocks at this stageWorked example
make a dataset calledmeasurements
add tomeasurementsthe example1labeled12
add tomeasurementsthe example2labeled16
add tomeasurementsthe example3labeled21
add tomeasurementsthe example4labeled24
add tomeasurementsthe example5labeled29
add tomeasurementsthe example6labeled33
add tomeasurementsthe example7labeled36
add tomeasurementsthe example8labeled41
make a line calledroutewith slope4.1and intercept8
make an empty list calleddistances
add1todistances
add2todistances
add3todistances
add4todistances
add5todistances
add6todistances
add7todistances
add8todistances
make an empty list calledactuals
add12toactuals
add16toactuals
add21toactuals
add24toactuals
add29toactuals
add33toactuals
add36toactuals
add41toactuals
setrowto1
setsquared_totalto0
for eachdistancein listdistances
setactualto
item
row
ofactuals
setpredictedto
routepredicts for x =
distance
seterrorto
predicted
-
actual
setsquaredto
error
x
error
changesquared_totalby
squared
sayDistancethen
distance
sayActualthen
actual
sayPredictionthen
predicted
sayResidualthen
round
error
to3places
saySquared residualthen
round
squared
to3places
changerowby1
sayManual mean squared errorthen
round
squared_total
/8
to3places
sayLine loss checkthen
the loss ofrouteonmeasurements

Change only the slope to 4.1 and inspect every row before comparing the mean.

What to look for

Both MSE calculations show 0.205, lower than 0.5 on these same observations.

Make it yours

Check the third row: prediction 20.3, residual -0.7, square 0.49. Use that path to diagnose a different total.

Restore slope four before the separate intercept experiment. Changing intercept from eight to nine then gives MSE 0.5 again on these rows. Equal losses do not mean identical predictions: some rows improve while others worsen. Always identify both parameter values when comparing results; a score without its model version is ambiguous.

Keep the range and evidence with the saved work

The observed distance range is one to eight kilometres. Predictions beyond that range extend the linear assumption without supplying new evidence. A trendline's displayed rounded equation may also produce slightly different results from full-precision internal coefficients; that is a precision issue, not failed multiplication.

Keep the source observations, units, parameter values, formulas or blocks, and version notes together. These are training comparisons of hand-selected rules. If you use a validation set to select settings, reserve a separate final test for reporting the selected model. The calculation can be identical across those sets while the evidential role is different.

Build with me · 8

Keep source, parameters, calculations, and claim together

The complete report preserves the observed pairs, hand-chosen parameters, intermediate calculations, and alternative loss. A scatter plot would place distance on the horizontal axis and time on the vertical axis; its numeric axes must preserve the actual distances between values.

A fitted or hand-chosen line summarises these examples. It does not establish causation, account for omitted conditions, or justify extrapolation far beyond eight kilometres. If you use validation to select a setting, keep final evaluation separate.

Blocks at this stageWorked example
make a dataset calledmeasurements
add tomeasurementsthe example1labeled12
add tomeasurementsthe example2labeled16
add tomeasurementsthe example3labeled21
add tomeasurementsthe example4labeled24
add tomeasurementsthe example5labeled29
add tomeasurementsthe example6labeled33
add tomeasurementsthe example7labeled36
add tomeasurementsthe example8labeled41
make a line calledroutewith slope4and intercept8
make an empty list calleddistances
add1todistances
add2todistances
add3todistances
add4todistances
add5todistances
add6todistances
add7todistances
add8todistances
make an empty list calledactuals
add12toactuals
add16toactuals
add21toactuals
add24toactuals
add29toactuals
add33toactuals
add36toactuals
add41toactuals
setrowto1
setsquared_totalto0
for eachdistancein listdistances
setactualto
item
row
ofactuals
setpredictedto
routepredicts for x =
distance
seterrorto
predicted
-
actual
setsquaredto
error
x
error
changesquared_totalby
squared
sayDistancethen
distance
sayActualthen
actual
sayPredictionthen
predicted
sayResidualthen
round
error
to3places
saySquared residualthen
round
squared
to3places
changerowby1
sayManual mean squared errorthen
round
squared_total
/8
to3places
sayLine loss checkthen
the loss ofrouteonmeasurements
set theslopeofrouteto4.1
sayAlternative slope 4.1 MSEthen
the loss ofrouteonmeasurements
sayObserved distances: 1 to 8 km. These are training comparisons.

Run the reference after constructing your own table calculation. Explain the difference between the original model's row outputs and the alternative model's final summary.

What to look for

The original MSE is 0.5 and the alternative is 0.205 on the same eight training rows.

Make it yours

Describe how the same sequence would look in spreadsheet columns, then point to each corresponding block operation.

Full reference solution

This is the complete worked program. Try building it yourself first, then use this reference to find the first place your version behaves differently. The Python below is generated from these exact blocks; helper functions are included so its behaviour can be inspected.

make a dataset calledmeasurements
add tomeasurementsthe example1labeled12
add tomeasurementsthe example2labeled16
add tomeasurementsthe example3labeled21
add tomeasurementsthe example4labeled24
add tomeasurementsthe example5labeled29
add tomeasurementsthe example6labeled33
add tomeasurementsthe example7labeled36
add tomeasurementsthe example8labeled41
make a line calledroutewith slope4and intercept8
make an empty list calleddistances
add1todistances
add2todistances
add3todistances
add4todistances
add5todistances
add6todistances
add7todistances
add8todistances
make an empty list calledactuals
add12toactuals
add16toactuals
add21toactuals
add24toactuals
add29toactuals
add33toactuals
add36toactuals
add41toactuals
setrowto1
setsquared_totalto0
for eachdistancein listdistances
setactualto
item
row
ofactuals
setpredictedto
routepredicts for x =
distance
seterrorto
predicted
-
actual
setsquaredto
error
x
error
changesquared_totalby
squared
sayDistancethen
distance
sayActualthen
actual
sayPredictionthen
predicted
sayResidualthen
round
error
to3places
saySquared residualthen
round
squared
to3places
changerowby1
sayManual mean squared errorthen
round
squared_total
/8
to3places
sayLine loss checkthen
the loss ofrouteonmeasurements
set theslopeofrouteto4.1
sayAlternative slope 4.1 MSEthen
the loss ofrouteonmeasurements
sayObserved distances: 1 to 8 km. These are training comparisons.
PythonHover over a line to see an explanation
# ---------------------------------------------------------------# Building blocks, written out in plain Python.# This part is generated for you. Your script starts further down.# ---------------------------------------------------------------  import mathimport random  class Dataset:    """A pile of labelled examples. Features can be a number, some text, or a    list of numbers. The label is whatever answer you want back."""     def __init__(self, name="dataset"):        self.name = name        self.rows = []        # Column names from the file's header row, when it came from one.        # Without these "the km of a row" has nothing to look the name up in.        self.columns = []     def add(self, features, label, fields=None):        self.rows.append(            {                "features": features,                "label": str(label),                "fields": dict(fields) if fields else {},            }        )     def size(self):        return len(self.rows)     def __len__(self):        return len(self.rows)     def __iter__(self):        """Walking a dataset gives you its rows, so "for each row in list        [testing]" reads the way it sounds."""        return iter(self.rows)     def labels(self):        seen = []        for row in self.rows:            if row["label"] not in seen:                seen.append(row["label"])        return seen     def most_common_label(self):        if not self.rows:            return ""        counts = {}        for row in self.rows:            counts[row["label"]] = counts.get(row["label"], 0) + 1        return max(counts, key=lambda label: counts[label])     def split(self, train_percent=80):        """Keeps the given percent for training and hands back the rest as a        test set. The shuffle is seeded, so you get the same split every run."""        order = list(range(len(self.rows)))        random.Random(0).shuffle(order)        cut = int(len(order) * train_percent / 100)        train = Dataset(self.name + " (train)")        test = Dataset(self.name + " (test)")        train.columns = list(self.columns)        test.columns = list(self.columns)        for position, index in enumerate(order):            row = self.rows[index]            target = train if position < cut else test            target.add(row["features"], row["label"], row.get("fields"))        return train, test     def show(self, limit=10):        print(self.name + ": " + str(len(self.rows)) + " examples")        for row in self.rows[:limit]:            print("  " + str(row["features"]) + "  ->  " + row["label"])        if len(self.rows) > limit:            print("  ... and " + str(len(self.rows) - limit) + " more")  def new_dataset(name="dataset"):    return Dataset(name)  def row_field(row, name):    """One named piece of a row: "the km of this row", "the label of it".     Names come from the header line of the csv. "label" always works, even on    a file with no header, because every row has one."""    wanted = str(name).strip()    if not isinstance(row, dict):        print("That is not a row. Use this inside a for each over a dataset.")        return ""    if wanted.lower() == "label":        return row.get("label", "")    fields = row.get("fields") or {}    if wanted in fields:        return fields[wanted]    # Header names are matched loosely, so "Rain" finds the "rain" column.    for key in fields:        if str(key).strip().lower() == wanted.lower():            return fields[key]    known = ", ".join([str(k) for k in fields]) if fields else "none"    print(        "No column called "        + wanted        + " in this row. Columns here: "        + known        + "."    )    return ""  def load_csv(dataset, path, label_column=-1, has_header=True):    """Reads a comma separated file into a dataset.     Everything except the label column becomes the features, and anything that    looks like a number is turned into one. This is how a file you uploaded    becomes something you can train on."""    try:        with open(path) as handle:            rows = [line.rstrip("\n").rstrip("\r") for line in handle]    except OSError:        print("Could not find " + path + ". Check the name in the Files panel.")        return dataset     rows = [row for row in rows if row.strip()]     # A file with no label column is a perfectly normal thing to load: it is    # what a test set looks like before you have predicted anything.    labelled = str(label_column).strip().lower() not in ("none", "", "no", "-")     header = []    if has_header and rows:        header = [cell.strip() for cell in rows[0].split(",")]        rows = rows[1:]     added = 0    for row in rows:        cells = [cell.strip() for cell in row.split(",")]        if not cells or (labelled and len(cells) < 2):            continue         index = -1        if labelled:            index = int(label_column)            if index < 0:                index = len(cells) + index            if index < 0 or index >= len(cells):                continue         label = cells[index] if labelled else ""        keep = [i for i in range(len(cells)) if i != index]         typed = []        for i in keep:            try:                typed.append(float(cells[i]))            except ValueError:                typed.append(cells[i])         # Every kept column gets its header name, so "the km of a row" works.        fields = {}        for position, i in enumerate(keep):            if i < len(header) and header[i]:                fields[header[i]] = typed[position]        if header and not dataset.columns:            dataset.columns = [header[i] for i in keep if i < len(header)]         dataset.add(typed[0] if len(typed) == 1 else typed, label, fields)        added += 1     print("Loaded " + str(added) + " rows from " + path + ".")    return dataset  def save_submission(predictions, dataset, path="submission.csv"):    """Prints your predictions in the exact two column format a round is    scored in: a header, then one id and one label per line.     It is printed rather than saved to a file because the console is the one    place you can copy it from. Paste it into a new file in the Files panel,    or straight into the upload box."""    labels = list(predictions or [])    rows = list(dataset) if dataset is not None else []     if len(labels) != len(rows):        print(            "You have "            + str(len(labels))            + " predictions for "            + str(len(rows))            + " rows. Those have to match before this means anything."        )        return     lines = ["id,label"]    for position, row in enumerate(rows):        # Looked up directly rather than through row_field, which would        # complain on every row of a file that simply has no id column.        fields = row.get("fields") or {}        row_id = ""        for key in fields:            if str(key).strip().lower() == "id":                row_id = fields[key]                break        if not str(row_id).strip():            row_id = "r" + str(position + 1).zfill(2)        if isinstance(row_id, float) and row_id == int(row_id):            row_id = int(row_id)        lines.append(str(row_id) + "," + str(labels[position]).strip())     print("--- submission.csv, copy from here ---")    for line in lines:        print(line)    print("--- to here, " + str(len(lines) - 1) + " rows ---")  class Line:    """y = slope * x + intercept. Two knobs, and learning is turning them."""     def __init__(self, slope=1.0, intercept=0.0, name="line"):        self.slope = float(slope)        self.intercept = float(intercept)        self.name = name     def predict(self, x):        return self.slope * float(x) + self.intercept     def _pairs(self, dataset):        pairs = []        for row in dataset.rows:            features = row["features"]            if isinstance(features, (list, tuple)):                features = features[0]            pairs.append((float(features), float(row["label"])))        return pairs     def loss(self, dataset):        """Mean squared error: the average of how wrong it is, squared."""        pairs = self._pairs(dataset)        if not pairs:            return 0.0        total = 0.0        for x, y in pairs:            error = self.predict(x) - y            # e * e rather than e ** 2. On a run that is blowing up, the power            # operator raises OverflowError and kills the script, while the            # multiply quietly reaches inf so the chart can report it.            total += error * error        average = total / len(pairs)        if average != average or average in (float("inf"), float("-inf")):            return float("inf")        return round(average, 4)     def step(self, dataset, rate=0.01):        """One nudge downhill. Look at how wrong the line is, then move both        knobs a small amount in the direction that makes it less wrong."""        pairs = self._pairs(dataset)        if not pairs:            return        slope_push = 0.0        intercept_push = 0.0        for x, y in pairs:            error = self.predict(x) - y            slope_push += 2 * error * x            intercept_push += 2 * error        self.slope -= rate * slope_push / len(pairs)        self.intercept -= rate * intercept_push / len(pairs)     def train(self, dataset, steps=200, rate=0.01):        """Many nudges in a row, with a chart of the loss as it goes.         The chart is the point. A run that bounces, a run that crawls and a        run that drops and flattens are three different problems, and the        final number alone tells them apart badly."""        count = int(steps)        history = [self.loss(dataset)]        for _ in range(count):            self.step(dataset, rate)            history.append(self.loss(dataset))         self._chart(history)        print(            "Trained "            + self.name            + " for "            + str(count)            + " steps. Loss is now "            + str(self.loss(dataset))        )     def _chart(self, history):        """A loss chart in text: one row per sampled step, bar length in        proportion to the loss at that point."""        if len(history) < 2:            return         # At most a dozen rows. The first few steps are always shown, because        # that is where a loss curve does nearly all of its moving, and a run        # that has settled by step 3 would otherwise look like a flat line.        wanted = 12        if len(history) <= wanted:            picks = list(range(len(history)))        else:            early = list(range(min(4, len(history))))            spread = wanted - len(early)            first = early[-1]            gap = (len(history) - 1 - first) / float(spread)            rest = [int(round(first + (i + 1) * gap)) for i in range(spread)]            picks = sorted(set(early + rest))            picks[-1] = len(history) - 1         finite = [history[i] for i in picks if history[i] < float("inf")]        top = max(finite) if finite else 0.0        width = len(str(len(history) - 1))         # Bars are drawn on a log scale. A first loss of 500 next to a final        # loss of 0.2 would otherwise flatten every interesting bar to nothing,        # and the shape of the drop is the whole reason to look at this.        scale = math.log(1.0 + top) if top > 0 else 0.0         print("Loss per step:")        for i in picks:            value = history[i]            step_label = str(i).rjust(width)            if value == float("inf"):                print("  step " + step_label + "  " + "off the chart")                continue            bars = 0            if scale > 0:                bars = int(round(30 * math.log(1.0 + value) / scale))            print(                "  step "                + step_label                + "  "                + ("#" * bars).ljust(30)                + " "                + str(round(value, 4))            )     def show(self):        print(            self.name            + ": y = "            + str(round(self.slope, 4))            + " * x + "            + str(round(self.intercept, 4))        )  def new_line(slope=1.0, intercept=0.0, name="line"):    return Line(slope, intercept, name)  # ---------------------------------------------------------------# Your script# ---------------------------------------------------------------  measurements = new_dataset("measurements")measurements.add(1, "12")measurements.add(2, "16")measurements.add(3, "21")measurements.add(4, "24")measurements.add(5, "29")measurements.add(6, "33")measurements.add(7, "36")measurements.add(8, "41")route = new_line(4, 8, "route")distances = []distances.append(1)distances.append(2)distances.append(3)distances.append(4)distances.append(5)distances.append(6)distances.append(7)distances.append(8)actuals = []actuals.append(12)actuals.append(16)actuals.append(21)actuals.append(24)actuals.append(29)actuals.append(33)actuals.append(36)actuals.append(41)row = 1squared_total = 0for distance in distances:    actual = actuals[int(row) - 1]    predicted = route.predict(distance)    error = (predicted - actual)    squared = (error * error)    squared_total = squared_total + squared    print("Distance", distance)    print("Actual", actual)    print("Prediction", predicted)    print("Residual", round(error, 3))    print("Squared residual", round(squared, 3))    row = row + 1print("Manual mean squared error", round((squared_total / 8), 3))print("Line loss check", route.loss(measurements))route.slope = float(4.1)print("Alternative slope 4.1 MSE", route.loss(measurements))print("Observed distances: 1 to 8 km. These are training comparisons.")

Compare this with your version. Different names and personal choices are fine when the program follows the same logic.

Worth a look

Keep your progress

Sign in and every reading, quiz, and exercise you finish is saved.

Sign in