0%
BuildUnder the hoodabout 28 min, 8 steps

Train, Validate, Test

Design evaluation so that parameter fitting, model selection, and final reporting each have their own evidence.

Assign a job to each set

This optional reading adds a third role, validation, to the training and development split you have used so far.

Training examples fit parameters. Validation examples, also called development examples, help you choose model types, features, thresholds, and settings. Final-test examples estimate performance after those choices are finished.

A score on training data is useful: it can reveal a broken fit or show how well the model optimises its training objective. It is weak evidence of performance on new inputs because fitting already used those examples. Calling training accuracy meaningless loses this useful diagnostic role; calling it a final test overstates it.

Split the right units

Randomly splitting rows is appropriate only when rows are sufficiently independent for the intended task. Keep copies and augmented versions of an image with the same original. Keep adjacent frames of one recording together. If the goal is new-person recognition, consider keeping people separate across the split.

For forecasting future events, training on later examples and testing on earlier ones can leak future information. A chronological split may better represent the real job. The split method follows the deployment question, not a universal recipe.

Preprocessing can leak information too. If you learn scaling statistics, feature selection, or imputation values from the entire dataset before splitting, the evaluation may influence the pipeline. Fit those transformations on training data, then apply them unchanged to validation and test.

Build with me · 1

Inspect the source before splitting

A data split assigns roles to observations. It does not automatically make observations independent. Duplicate frames, related people, or future information can cross a random row split and leak information even when the dataset names look correct.

These invented delivery rows are a compact bookkeeping example. In a real project, choose split units that match the question, such as new people or future dates.

Blocks at this stageWorked example
make a dataset calleddeliveries
loadcouriers-practice.csvintodeliveriesusing column-1as the label, andskip the first row
sayAll rowsthen
how many examples indeliveries
show datasetdeliveries

Inspect the included CSV and identify what one row represents before running any model.

What to look for

All rows is 12.

Make it yours

Describe how you would split a collection containing several photos of each person if the goal were performance on new people.

Build with me · 2

Separate fitting data before fitting

The first split keeps sixty percent for training, rounded down to seven rows by this implementation. Five remain reserved. At this point neither a candidate nor a baseline has been trained.

Reserving rows after inspecting training success would not undo earlier use of their answers. The timing and actual information flow determine whether evidence was held out.

Blocks at this stageWorked example
make a dataset calleddeliveries
loadcouriers-practice.csvintodeliveriesusing column-1as the label, andskip the first row
splitdeliveriesintotrainingandreservedkeeping60percent for training
sayTrainingthen
how many examples intraining
sayReservedthen
how many examples inreserved

Split deliveries into training and reserved, with percentage 60. Print both counts.

What to look for

Training is 7 and Reserved is 5.

Make it yours

Explain why these rounded counts do not match a perfect 60/40 percentage exactly for twelve rows.

Build with me · 3

Give the reserved rows two distinct purposes

Split the five reserved rows again: three for validation and two for final testing. Training fits the model, validation selects choices, and the final set judges the selected version after those choices.

There is no universal best percentage. This example illustrates separate roles with tiny numbers. Two final rows provide very little evidence about the wider task.

Blocks at this stageWorked example
make a dataset calleddeliveries
loadcouriers-practice.csvintodeliveriesusing column-1as the label, andskip the first row
splitdeliveriesintotrainingandreservedkeeping60percent for training
splitreservedintovalidationandfinal_testkeeping60percent for training
sayTrainingthen
how many examples intraining
sayValidationthen
how many examples invalidation
sayFinal testthen
how many examples infinal_test

Add a second split of reserved into validation and final_test, keeping 60 percent for validation.

What to look for

The counts are 7 training, 3 validation, and 2 final-test rows; they add to 12.

Make it yours

For one hundred independent examples, describe a 60/20/20 allocation and the job of each group.

Build with me · 4

Fit both candidates on the same training set

The closest-example candidate and constant-label baseline see the same seven fitting rows. Neither train block names validation or final_test. Their different procedures therefore start from the same permitted evidence.

If preprocessing learned scaling or replacement values, it too would need to fit on training data and apply unchanged to the other sets. Splitting after learning those values from everything would leak evaluation information.

Blocks at this stageWorked example
make a dataset calleddeliveries
loadcouriers-practice.csvintodeliveriesusing column-1as the label, andskip the first row
splitdeliveriesintotrainingandreservedkeeping60percent for training
splitreservedintovalidationandfinal_testkeeping60percent for training
make aclosest examplemodel calledcandidate
traincandidateontraining
make aalways the most common labelmodel calledbaseline
trainbaselineontraining
show whatcandidatelearned
show whatbaselinelearned

Create and train both model types on training only. Inspect the fit counts.

What to look for

Both model reports show seven training examples.

Make it yours

Point to the dataset menu in each train block and explain why selecting deliveries would invalidate the intended split.

Work through a selection decision

Suppose you compare two learning rates while keeping the same training and validation sets. Rate A gives lower validation loss, so you choose it. That choice used validation evidence, even though the validation examples did not directly update the weights during fitting.

After many such choices, you can over-adapt to the validation set. Repeatedly trying changes until a particular small set looks excellent does not make it an untouched estimate. Keep the search limited and purposeful, obtain more representative validation data when possible, and report the selection process.

Evaluate the chosen version on the final test. If its result disappoints you, record it. You may start another development cycle, but if you use those test errors to guide changes, that set is now part of development. A new independent final test is needed for a fresh estimate.

Build with me · 5

Compare choices before touching the final set

Validation results guide the choice between methods. Even though validation rows do not directly fit parameters, their answers influence selection. That is why validation and final testing are different roles.

A three-row validation score is coarse: one result changes accuracy by a third. Trying hundreds of versions until this small set looks perfect can over-adapt to it.

Blocks at this stageWorked example
make a dataset calleddeliveries
loadcouriers-practice.csvintodeliveriesusing column-1as the label, andskip the first row
splitdeliveriesintotrainingandreservedkeeping60percent for training
splitreservedintovalidationandfinal_testkeeping60percent for training
make aclosest examplemodel calledcandidate
traincandidateontraining
make aalways the most common labelmodel calledbaseline
trainbaselineontraining
sayCandidate validationthen
out of 100, the accuracy ofcandidateonvalidation
sayBaseline validationthen
out of 100, the accuracy ofbaselineonvalidation

Print each model's accuracy on validation. Do not add final-test outputs yet.

What to look for

Two measured percentages compare the same three validation rows.

Make it yours

State a selection rule before examining more results, including how you will handle a tie.

Build with me · 6

Make the selection rule visible

This rule selects higher validation accuracy, with a tie assigned to the closest-example candidate for this demonstration. A real task may instead prefer the simpler baseline on a tie; the important point is to state the choice rather than hide it after seeing results.

Selection changes which model the application uses. It does not magically turn validation into untouched evidence. The rule explicitly depends on those labels.

Blocks at this stageWorked example
make a dataset calleddeliveries
loadcouriers-practice.csvintodeliveriesusing column-1as the label, andskip the first row
splitdeliveriesintotrainingandreservedkeeping60percent for training
splitreservedintovalidationandfinal_testkeeping60percent for training
make aclosest examplemodel calledcandidate
traincandidateontraining
make aalways the most common labelmodel calledbaseline
trainbaselineontraining
if
out of 100, the accuracy ofcandidateonvalidation
>=
out of 100, the accuracy ofbaselineonvalidation
then
saySelect the closest-example candidate
otherwise
saySelect the constant-label baseline

Put the two validation accuracy values in a comparison, then choose the model name through if/else messages.

What to look for

Exactly one selected-method message prints according to the measured validation scores.

Make it yours

Change only the documented tie rule and explain when that change could affect selection.

Build with me · 7

Evaluate the chosen method once

The final-test accuracy is calculated only inside the branch for the selected method. We do not inspect both final scores and then pick whichever looks best. That would use final-test answers for selection.

If the final result disappoints you, report it. You may start another development cycle, but using these failures to guide changes makes this set development evidence for that new cycle.

Blocks at this stageWorked example
make a dataset calleddeliveries
loadcouriers-practice.csvintodeliveriesusing column-1as the label, andskip the first row
splitdeliveriesintotrainingandreservedkeeping60percent for training
splitreservedintovalidationandfinal_testkeeping60percent for training
make aclosest examplemodel calledcandidate
traincandidateontraining
make aalways the most common labelmodel calledbaseline
trainbaselineontraining
if
out of 100, the accuracy ofcandidateonvalidation
>=
out of 100, the accuracy ofbaselineonvalidation
then
sayChosen by validation: closest-example candidate
sayFinal-test accuracythen
out of 100, the accuracy ofcandidateonfinal_test
otherwise
sayChosen by validation: constant-label baseline
sayFinal-test accuracythen
out of 100, the accuracy ofbaselineonfinal_test

Add final-test scoring inside each corresponding selection branch. Read the order: fit, select using validation, then score the selected model.

What to look for

One method name and one final-test accuracy print, measured on two rows.

Make it yours

Explain what should change in your evaluation plan if you use these two final errors to redesign the model.

Allocate the examples before choosing a model

Imagine a collection of one hundred independently captured labelled images. For a teaching example, reserve sixty for training, twenty for validation, and twenty for final testing. These counts illustrate the roles; they are not a universal recommended split.

Train two candidates on the same sixty examples. Candidate A gets sixteen of twenty validation examples right, while B gets seventeen. If validation accuracy is your stated selection rule, you choose B, then evaluate B on the untouched final twenty.

Suppose B gets fifteen final-test examples right. Report 15/20, or 75%, along with the small sample size. Do not replace that result with the better validation score. The final set was different evidence, and differences between small sets are unsurprising.

If you inspect its five errors, change the model, and try again on those same twenty examples, those images have now influenced development. You may still use them to learn, but stop calling them an untouched final test. A new independent set is needed for a fresh final claim. Split bookkeeping is about how information was used, not the variable names assigned to arrays.

Interpret gaps cautiously

Very good training performance with poorer validation performance is consistent with overfitting. It can also involve distribution changes, labelling problems, or sampling noise. Poor scores on both sets can reflect underfitting, weak inputs, or a faulty pipeline. Look at examples and intermediate values before declaring one cause.

Report counts as well as percentages. Ten thousand duplicate frames are not ten thousand independent scenes. Include class balance, split units, and important conditions so a reader can judge what the test covers.

The final question is operational: if you received a genuinely new input tomorrow, which saved model and preprocessing would you use, and what evidence supports its expected performance? A well-designed split helps answer that question honestly.

Build with me · 8

Keep counts and roles with every score

The report records the partition sizes, validation comparison, selection rule, and chosen final result. Percentages without counts can make this tiny example appear much stronger than it is. A 50-point final-score change here is only one row.

Generalisation depends on whether future inputs resemble the intended evaluation population. Good bookkeeping protects against one class of misleading result; representative independent data and careful task definition remain necessary.

Blocks at this stageWorked example
make a dataset calleddeliveries
loadcouriers-practice.csvintodeliveriesusing column-1as the label, andskip the first row
splitdeliveriesintotrainingandreservedkeeping60percent for training
splitreservedintovalidationandfinal_testkeeping60percent for training
make aclosest examplemodel calledcandidate
traincandidateontraining
make aalways the most common labelmodel calledbaseline
trainbaselineontraining
sayTraining countthen
how many examples intraining
sayValidation countthen
how many examples invalidation
sayFinal-test countthen
how many examples infinal_test
sayCandidate validationthen
out of 100, the accuracy ofcandidateonvalidation
sayBaseline validationthen
out of 100, the accuracy ofbaselineonvalidation
if
out of 100, the accuracy ofcandidateonvalidation
>=
out of 100, the accuracy ofbaselineonvalidation
then
sayChosen by validation: closest-example candidate
sayFinal-test accuracythen
out of 100, the accuracy ofcandidateonfinal_test
otherwise
sayChosen by validation: constant-label baseline
sayFinal-test accuracythen
out of 100, the accuracy ofbaselineonfinal_test
sayTwo final rows are a demonstration, not a dependable estimate

Run the complete report and explain each output's role. Keep the included files with the source if you save the experiment.

What to look for

The counts remain 7, 3, and 2, with an actual validation-selected final result.

Make it yours

Write a better collection and split plan for a real project, including duplicates, groups, and time if relevant.

Full reference solution

This is the complete worked program. Try building it yourself first, then use this reference to find the first place your version behaves differently. The Python below is generated from these exact blocks; helper functions are included so its behaviour can be inspected.

make a dataset calleddeliveries
loadcouriers-practice.csvintodeliveriesusing column-1as the label, andskip the first row
splitdeliveriesintotrainingandreservedkeeping60percent for training
splitreservedintovalidationandfinal_testkeeping60percent for training
make aclosest examplemodel calledcandidate
traincandidateontraining
make aalways the most common labelmodel calledbaseline
trainbaselineontraining
sayTraining countthen
how many examples intraining
sayValidation countthen
how many examples invalidation
sayFinal-test countthen
how many examples infinal_test
sayCandidate validationthen
out of 100, the accuracy ofcandidateonvalidation
sayBaseline validationthen
out of 100, the accuracy ofbaselineonvalidation
if
out of 100, the accuracy ofcandidateonvalidation
>=
out of 100, the accuracy ofbaselineonvalidation
then
sayChosen by validation: closest-example candidate
sayFinal-test accuracythen
out of 100, the accuracy ofcandidateonfinal_test
otherwise
sayChosen by validation: constant-label baseline
sayFinal-test accuracythen
out of 100, the accuracy ofbaselineonfinal_test
sayTwo final rows are a demonstration, not a dependable estimate
PythonHover over a line to see an explanation
# ---------------------------------------------------------------# Building blocks, written out in plain Python.# This part is generated for you. Your script starts further down.# ---------------------------------------------------------------  import mathimport random  class Dataset:    """A pile of labelled examples. Features can be a number, some text, or a    list of numbers. The label is whatever answer you want back."""     def __init__(self, name="dataset"):        self.name = name        self.rows = []        # Column names from the file's header row, when it came from one.        # Without these "the km of a row" has nothing to look the name up in.        self.columns = []     def add(self, features, label, fields=None):        self.rows.append(            {                "features": features,                "label": str(label),                "fields": dict(fields) if fields else {},            }        )     def size(self):        return len(self.rows)     def __len__(self):        return len(self.rows)     def __iter__(self):        """Walking a dataset gives you its rows, so "for each row in list        [testing]" reads the way it sounds."""        return iter(self.rows)     def labels(self):        seen = []        for row in self.rows:            if row["label"] not in seen:                seen.append(row["label"])        return seen     def most_common_label(self):        if not self.rows:            return ""        counts = {}        for row in self.rows:            counts[row["label"]] = counts.get(row["label"], 0) + 1        return max(counts, key=lambda label: counts[label])     def split(self, train_percent=80):        """Keeps the given percent for training and hands back the rest as a        test set. The shuffle is seeded, so you get the same split every run."""        order = list(range(len(self.rows)))        random.Random(0).shuffle(order)        cut = int(len(order) * train_percent / 100)        train = Dataset(self.name + " (train)")        test = Dataset(self.name + " (test)")        train.columns = list(self.columns)        test.columns = list(self.columns)        for position, index in enumerate(order):            row = self.rows[index]            target = train if position < cut else test            target.add(row["features"], row["label"], row.get("fields"))        return train, test     def show(self, limit=10):        print(self.name + ": " + str(len(self.rows)) + " examples")        for row in self.rows[:limit]:            print("  " + str(row["features"]) + "  ->  " + row["label"])        if len(self.rows) > limit:            print("  ... and " + str(len(self.rows) - limit) + " more")  def new_dataset(name="dataset"):    return Dataset(name)  def row_field(row, name):    """One named piece of a row: "the km of this row", "the label of it".     Names come from the header line of the csv. "label" always works, even on    a file with no header, because every row has one."""    wanted = str(name).strip()    if not isinstance(row, dict):        print("That is not a row. Use this inside a for each over a dataset.")        return ""    if wanted.lower() == "label":        return row.get("label", "")    fields = row.get("fields") or {}    if wanted in fields:        return fields[wanted]    # Header names are matched loosely, so "Rain" finds the "rain" column.    for key in fields:        if str(key).strip().lower() == wanted.lower():            return fields[key]    known = ", ".join([str(k) for k in fields]) if fields else "none"    print(        "No column called "        + wanted        + " in this row. Columns here: "        + known        + "."    )    return ""  def load_csv(dataset, path, label_column=-1, has_header=True):    """Reads a comma separated file into a dataset.     Everything except the label column becomes the features, and anything that    looks like a number is turned into one. This is how a file you uploaded    becomes something you can train on."""    try:        with open(path) as handle:            rows = [line.rstrip("\n").rstrip("\r") for line in handle]    except OSError:        print("Could not find " + path + ". Check the name in the Files panel.")        return dataset     rows = [row for row in rows if row.strip()]     # A file with no label column is a perfectly normal thing to load: it is    # what a test set looks like before you have predicted anything.    labelled = str(label_column).strip().lower() not in ("none", "", "no", "-")     header = []    if has_header and rows:        header = [cell.strip() for cell in rows[0].split(",")]        rows = rows[1:]     added = 0    for row in rows:        cells = [cell.strip() for cell in row.split(",")]        if not cells or (labelled and len(cells) < 2):            continue         index = -1        if labelled:            index = int(label_column)            if index < 0:                index = len(cells) + index            if index < 0 or index >= len(cells):                continue         label = cells[index] if labelled else ""        keep = [i for i in range(len(cells)) if i != index]         typed = []        for i in keep:            try:                typed.append(float(cells[i]))            except ValueError:                typed.append(cells[i])         # Every kept column gets its header name, so "the km of a row" works.        fields = {}        for position, i in enumerate(keep):            if i < len(header) and header[i]:                fields[header[i]] = typed[position]        if header and not dataset.columns:            dataset.columns = [header[i] for i in keep if i < len(header)]         dataset.add(typed[0] if len(typed) == 1 else typed, label, fields)        added += 1     print("Loaded " + str(added) + " rows from " + path + ".")    return dataset  def save_submission(predictions, dataset, path="submission.csv"):    """Prints your predictions in the exact two column format a round is    scored in: a header, then one id and one label per line.     It is printed rather than saved to a file because the console is the one    place you can copy it from. Paste it into a new file in the Files panel,    or straight into the upload box."""    labels = list(predictions or [])    rows = list(dataset) if dataset is not None else []     if len(labels) != len(rows):        print(            "You have "            + str(len(labels))            + " predictions for "            + str(len(rows))            + " rows. Those have to match before this means anything."        )        return     lines = ["id,label"]    for position, row in enumerate(rows):        # Looked up directly rather than through row_field, which would        # complain on every row of a file that simply has no id column.        fields = row.get("fields") or {}        row_id = ""        for key in fields:            if str(key).strip().lower() == "id":                row_id = fields[key]                break        if not str(row_id).strip():            row_id = "r" + str(position + 1).zfill(2)        if isinstance(row_id, float) and row_id == int(row_id):            row_id = int(row_id)        lines.append(str(row_id) + "," + str(labels[position]).strip())     print("--- submission.csv, copy from here ---")    for line in lines:        print(line)    print("--- to here, " + str(len(lines) - 1) + " rows ---")  def _words_in(value):    letters = []    for character in str(value).lower():        letters.append(character if character.isalnum() else " ")    return [word for word in "".join(letters).split() if word]  # Words that turn up in every kind of sentence carry no signal about the# label, so the word matching model looks past them._EVERYDAY_WORDS = set(    "a an and are as at be been but by can could did do for from had has have "    "he her his i if in is it its me my not of on or our she so than that the "    "their them then there they this to too us was we were what when which "    "who will with would you your".split())  def _content_words(value):    words = _words_in(value)    kept = [word for word in words if word not in _EVERYDAY_WORDS]    return kept if kept else words  def _as_numbers(value):    if isinstance(value, (list, tuple)):        return [float(item) for item in value]    return [float(value)]  def _is_numeric(value):    try:        _as_numbers(value)        return True    except (TypeError, ValueError):        return False  def _distance(left, right):    """How far apart two examples are. Numbers use straight line distance,    text uses how many words the two do not share."""    if _is_numeric(left) and _is_numeric(right):        a = _as_numbers(left)        b = _as_numbers(right)        while len(a) < len(b):            a.append(0.0)        while len(b) < len(a):            b.append(0.0)        total = 0.0        for i in range(len(a)):            total += (a[i] - b[i]) ** 2        return math.sqrt(total)    a = set(_words_in(left))    b = set(_words_in(right))    if not a and not b:        return 0.0    shared = len(a & b)    return 1.0 - (shared / float(len(a | b)))  class Model:    """Three small classifiers behind one name.     nearest  looks for the closest example it was trained on    words    scores how many words the input shares with each label    common   always answers with the most common label, the baseline to beat    """     def __init__(self, kind="nearest", name="model"):        self.kind = kind        self.name = name        self.data = None        self.word_scores = {}     def train(self, dataset):        self.data = dataset        self.word_scores = {}        if self.kind == "words":            for row in dataset.rows:                bucket = self.word_scores.setdefault(row["label"], {})                for word in _content_words(row["features"]):                    bucket[word] = bucket.get(word, 0) + 1        print("Trained " + self.name + " on " + str(dataset.size()) + " examples.")     def _scores(self, features):        if self.data is None or not self.data.rows:            return {}        if self.kind == "common":            counts = {}            for row in self.data.rows:                counts[row["label"]] = counts.get(row["label"], 0) + 1            return counts        if self.kind == "words":            # Score each label by how much of its training vocabulary shows up            # in the input, divided by how much text that label was trained on            # so a label with more examples cannot win on volume alone.            labels = self.data.labels()            scores = {}            for label in labels:                bucket = self.word_scores.get(label, {})                seen = sum(bucket.values()) or 1                running = 0.0                for word in _content_words(features):                    running += bucket.get(word, 0) / float(seen)                scores[label] = round(running, 6)            if sum(scores.values()) == 0:                # Nothing in the input was ever seen in training. Say so by                # splitting the vote evenly, which reads as low confidence.                return dict((label, 1) for label in labels)            return scores        ranked = sorted(self.data.rows, key=lambda row: _distance(features, row["features"]))        neighbours = ranked[: min(3, len(ranked))]        scores = {}        for row in neighbours:            scores[row["label"]] = scores.get(row["label"], 0) + 1        return scores     def predict(self, features):        scores = self._scores(features)        if not scores:            return ""        return max(scores, key=lambda label: scores[label])     def confidence(self, features):        """How much of the vote the winning label took, out of 100."""        scores = self._scores(features)        total = sum(scores.values())        if not scores or total == 0:            return 0.0        best = max(scores.values())        return round(100.0 * best / float(total), 1)     def accuracy(self, dataset):        if not dataset.rows:            return 0.0        right = 0        for row in dataset.rows:            if self.predict(row["features"]) == row["label"]:                right += 1        return round(100.0 * right / float(len(dataset.rows)), 1)     def show(self):        print(self.name + " is a " + self.kind + " model.")        if self.data is None:            print("  It has not been trained yet.")            return        print("  Trained on " + str(self.data.size()) + " examples.")        print("  Labels it can answer with: " + ", ".join(self.data.labels()))  def new_model(kind="nearest", name="model"):    return Model(kind, name)  # ---------------------------------------------------------------# Your script# ---------------------------------------------------------------  deliveries = new_dataset("deliveries")load_csv(deliveries, "couriers-practice.csv", -1, True)training, reserved = deliveries.split(60)validation, final_test = reserved.split(60)candidate = new_model("nearest", "candidate")candidate.train(training)baseline = new_model("common", "baseline")baseline.train(training)print("Training count", training.size())print("Validation count", validation.size())print("Final-test count", final_test.size())print("Candidate validation", candidate.accuracy(validation))print("Baseline validation", baseline.accuracy(validation))if (candidate.accuracy(validation) >= baseline.accuracy(validation)):    print("Chosen by validation: closest-example candidate")    print("Final-test accuracy", candidate.accuracy(final_test))else:    print("Chosen by validation: constant-label baseline")    print("Final-test accuracy", baseline.accuracy(final_test))print("Two final rows are a demonstration, not a dependable estimate")

Compare this with your version. Different names and personal choices are fine when the program follows the same logic.

Keep your progress

Sign in and every reading, quiz, and exercise you finish is saved.

Sign in