Train, Validate, Test
Design evaluation so that parameter fitting, model selection, and final reporting each have their own evidence.
Assign a job to each set
This optional reading adds a third role, validation, to the training and development split you have used so far.
Training examples fit parameters. Validation examples, also called development examples, help you choose model types, features, thresholds, and settings. Final-test examples estimate performance after those choices are finished.
A score on training data is useful: it can reveal a broken fit or show how well the model optimises its training objective. It is weak evidence of performance on new inputs because fitting already used those examples. Calling training accuracy meaningless loses this useful diagnostic role; calling it a final test overstates it.
Split the right units
Randomly splitting rows is appropriate only when rows are sufficiently independent for the intended task. Keep copies and augmented versions of an image with the same original. Keep adjacent frames of one recording together. If the goal is new-person recognition, consider keeping people separate across the split.
For forecasting future events, training on later examples and testing on earlier ones can leak future information. A chronological split may better represent the real job. The split method follows the deployment question, not a universal recipe.
Preprocessing can leak information too. If you learn scaling statistics, feature selection, or imputation values from the entire dataset before splitting, the evaluation may influence the pipeline. Fit those transformations on training data, then apply them unchanged to validation and test.
Build with me · 1
Inspect the source before splitting
A data split assigns roles to observations. It does not automatically make observations independent. Duplicate frames, related people, or future information can cross a random row split and leak information even when the dataset names look correct.
These invented delivery rows are a compact bookkeeping example. In a real project, choose split units that match the question, such as new people or future dates.
Inspect the included CSV and identify what one row represents before running any model.
What to look for
All rows is 12.
Make it yours
Describe how you would split a collection containing several photos of each person if the goal were performance on new people.
Build with me · 2
Separate fitting data before fitting
The first split keeps sixty percent for training, rounded down to seven rows by this implementation. Five remain reserved. At this point neither a candidate nor a baseline has been trained.
Reserving rows after inspecting training success would not undo earlier use of their answers. The timing and actual information flow determine whether evidence was held out.
Split deliveries into training and reserved, with percentage 60. Print both counts.
What to look for
Training is 7 and Reserved is 5.
Make it yours
Explain why these rounded counts do not match a perfect 60/40 percentage exactly for twelve rows.
Build with me · 3
Give the reserved rows two distinct purposes
Split the five reserved rows again: three for validation and two for final testing. Training fits the model, validation selects choices, and the final set judges the selected version after those choices.
There is no universal best percentage. This example illustrates separate roles with tiny numbers. Two final rows provide very little evidence about the wider task.
Add a second split of reserved into validation and final_test, keeping 60 percent for validation.
What to look for
The counts are 7 training, 3 validation, and 2 final-test rows; they add to 12.
Make it yours
For one hundred independent examples, describe a 60/20/20 allocation and the job of each group.
Build with me · 4
Fit both candidates on the same training set
The closest-example candidate and constant-label baseline see the same seven fitting rows. Neither train block names validation or final_test. Their different procedures therefore start from the same permitted evidence.
If preprocessing learned scaling or replacement values, it too would need to fit on training data and apply unchanged to the other sets. Splitting after learning those values from everything would leak evaluation information.
Create and train both model types on training only. Inspect the fit counts.
What to look for
Both model reports show seven training examples.
Make it yours
Point to the dataset menu in each train block and explain why selecting deliveries would invalidate the intended split.
Work through a selection decision
Suppose you compare two learning rates while keeping the same training and validation sets. Rate A gives lower validation loss, so you choose it. That choice used validation evidence, even though the validation examples did not directly update the weights during fitting.
After many such choices, you can over-adapt to the validation set. Repeatedly trying changes until a particular small set looks excellent does not make it an untouched estimate. Keep the search limited and purposeful, obtain more representative validation data when possible, and report the selection process.
Evaluate the chosen version on the final test. If its result disappoints you, record it. You may start another development cycle, but if you use those test errors to guide changes, that set is now part of development. A new independent final test is needed for a fresh estimate.
Build with me · 5
Compare choices before touching the final set
Validation results guide the choice between methods. Even though validation rows do not directly fit parameters, their answers influence selection. That is why validation and final testing are different roles.
A three-row validation score is coarse: one result changes accuracy by a third. Trying hundreds of versions until this small set looks perfect can over-adapt to it.
Print each model's accuracy on validation. Do not add final-test outputs yet.
What to look for
Two measured percentages compare the same three validation rows.
Make it yours
State a selection rule before examining more results, including how you will handle a tie.
Build with me · 6
Make the selection rule visible
This rule selects higher validation accuracy, with a tie assigned to the closest-example candidate for this demonstration. A real task may instead prefer the simpler baseline on a tie; the important point is to state the choice rather than hide it after seeing results.
Selection changes which model the application uses. It does not magically turn validation into untouched evidence. The rule explicitly depends on those labels.
Put the two validation accuracy values in a comparison, then choose the model name through if/else messages.
What to look for
Exactly one selected-method message prints according to the measured validation scores.
Make it yours
Change only the documented tie rule and explain when that change could affect selection.
Build with me · 7
Evaluate the chosen method once
The final-test accuracy is calculated only inside the branch for the selected method. We do not inspect both final scores and then pick whichever looks best. That would use final-test answers for selection.
If the final result disappoints you, report it. You may start another development cycle, but using these failures to guide changes makes this set development evidence for that new cycle.
Add final-test scoring inside each corresponding selection branch. Read the order: fit, select using validation, then score the selected model.
What to look for
One method name and one final-test accuracy print, measured on two rows.
Make it yours
Explain what should change in your evaluation plan if you use these two final errors to redesign the model.
Allocate the examples before choosing a model
Imagine a collection of one hundred independently captured labelled images. For a teaching example, reserve sixty for training, twenty for validation, and twenty for final testing. These counts illustrate the roles; they are not a universal recommended split.
Train two candidates on the same sixty examples. Candidate A gets sixteen of twenty validation examples right, while B gets seventeen. If validation accuracy is your stated selection rule, you choose B, then evaluate B on the untouched final twenty.
Suppose B gets fifteen final-test examples right. Report 15/20, or 75%, along with the small sample size. Do not replace that result with the better validation score. The final set was different evidence, and differences between small sets are unsurprising.
If you inspect its five errors, change the model, and try again on those same twenty examples, those images have now influenced development. You may still use them to learn, but stop calling them an untouched final test. A new independent set is needed for a fresh final claim. Split bookkeeping is about how information was used, not the variable names assigned to arrays.
Interpret gaps cautiously
Very good training performance with poorer validation performance is consistent with overfitting. It can also involve distribution changes, labelling problems, or sampling noise. Poor scores on both sets can reflect underfitting, weak inputs, or a faulty pipeline. Look at examples and intermediate values before declaring one cause.
Report counts as well as percentages. Ten thousand duplicate frames are not ten thousand independent scenes. Include class balance, split units, and important conditions so a reader can judge what the test covers.
The final question is operational: if you received a genuinely new input tomorrow, which saved model and preprocessing would you use, and what evidence supports its expected performance? A well-designed split helps answer that question honestly.
Build with me · 8
Keep counts and roles with every score
The report records the partition sizes, validation comparison, selection rule, and chosen final result. Percentages without counts can make this tiny example appear much stronger than it is. A 50-point final-score change here is only one row.
Generalisation depends on whether future inputs resemble the intended evaluation population. Good bookkeeping protects against one class of misleading result; representative independent data and careful task definition remain necessary.
Run the complete report and explain each output's role. Keep the included files with the source if you save the experiment.
What to look for
The counts remain 7, 3, and 2, with an actual validation-selected final result.
Make it yours
Write a better collection and split plan for a real project, including duplicates, groups, and time if relevant.
Full reference solution
This is the complete worked program. Try building it yourself first, then use this reference to find the first place your version behaves differently. The Python below is generated from these exact blocks; helper functions are included so its behaviour can be inspected.
# ---------------------------------------------------------------# Building blocks, written out in plain Python.# This part is generated for you. Your script starts further down.# --------------------------------------------------------------- import mathimport random class Dataset: """A pile of labelled examples. Features can be a number, some text, or a list of numbers. The label is whatever answer you want back.""" def __init__(self, name="dataset"): self.name = name self.rows = [] # Column names from the file's header row, when it came from one. # Without these "the km of a row" has nothing to look the name up in. self.columns = [] def add(self, features, label, fields=None): self.rows.append( { "features": features, "label": str(label), "fields": dict(fields) if fields else {}, } ) def size(self): return len(self.rows) def __len__(self): return len(self.rows) def __iter__(self): """Walking a dataset gives you its rows, so "for each row in list [testing]" reads the way it sounds.""" return iter(self.rows) def labels(self): seen = [] for row in self.rows: if row["label"] not in seen: seen.append(row["label"]) return seen def most_common_label(self): if not self.rows: return "" counts = {} for row in self.rows: counts[row["label"]] = counts.get(row["label"], 0) + 1 return max(counts, key=lambda label: counts[label]) def split(self, train_percent=80): """Keeps the given percent for training and hands back the rest as a test set. The shuffle is seeded, so you get the same split every run.""" order = list(range(len(self.rows))) random.Random(0).shuffle(order) cut = int(len(order) * train_percent / 100) train = Dataset(self.name + " (train)") test = Dataset(self.name + " (test)") train.columns = list(self.columns) test.columns = list(self.columns) for position, index in enumerate(order): row = self.rows[index] target = train if position < cut else test target.add(row["features"], row["label"], row.get("fields")) return train, test def show(self, limit=10): print(self.name + ": " + str(len(self.rows)) + " examples") for row in self.rows[:limit]: print(" " + str(row["features"]) + " -> " + row["label"]) if len(self.rows) > limit: print(" ... and " + str(len(self.rows) - limit) + " more") def new_dataset(name="dataset"): return Dataset(name) def row_field(row, name): """One named piece of a row: "the km of this row", "the label of it". Names come from the header line of the csv. "label" always works, even on a file with no header, because every row has one.""" wanted = str(name).strip() if not isinstance(row, dict): print("That is not a row. Use this inside a for each over a dataset.") return "" if wanted.lower() == "label": return row.get("label", "") fields = row.get("fields") or {} if wanted in fields: return fields[wanted] # Header names are matched loosely, so "Rain" finds the "rain" column. for key in fields: if str(key).strip().lower() == wanted.lower(): return fields[key] known = ", ".join([str(k) for k in fields]) if fields else "none" print( "No column called " + wanted + " in this row. Columns here: " + known + "." ) return "" def load_csv(dataset, path, label_column=-1, has_header=True): """Reads a comma separated file into a dataset. Everything except the label column becomes the features, and anything that looks like a number is turned into one. This is how a file you uploaded becomes something you can train on.""" try: with open(path) as handle: rows = [line.rstrip("\n").rstrip("\r") for line in handle] except OSError: print("Could not find " + path + ". Check the name in the Files panel.") return dataset rows = [row for row in rows if row.strip()] # A file with no label column is a perfectly normal thing to load: it is # what a test set looks like before you have predicted anything. labelled = str(label_column).strip().lower() not in ("none", "", "no", "-") header = [] if has_header and rows: header = [cell.strip() for cell in rows[0].split(",")] rows = rows[1:] added = 0 for row in rows: cells = [cell.strip() for cell in row.split(",")] if not cells or (labelled and len(cells) < 2): continue index = -1 if labelled: index = int(label_column) if index < 0: index = len(cells) + index if index < 0 or index >= len(cells): continue label = cells[index] if labelled else "" keep = [i for i in range(len(cells)) if i != index] typed = [] for i in keep: try: typed.append(float(cells[i])) except ValueError: typed.append(cells[i]) # Every kept column gets its header name, so "the km of a row" works. fields = {} for position, i in enumerate(keep): if i < len(header) and header[i]: fields[header[i]] = typed[position] if header and not dataset.columns: dataset.columns = [header[i] for i in keep if i < len(header)] dataset.add(typed[0] if len(typed) == 1 else typed, label, fields) added += 1 print("Loaded " + str(added) + " rows from " + path + ".") return dataset def save_submission(predictions, dataset, path="submission.csv"): """Prints your predictions in the exact two column format a round is scored in: a header, then one id and one label per line. It is printed rather than saved to a file because the console is the one place you can copy it from. Paste it into a new file in the Files panel, or straight into the upload box.""" labels = list(predictions or []) rows = list(dataset) if dataset is not None else [] if len(labels) != len(rows): print( "You have " + str(len(labels)) + " predictions for " + str(len(rows)) + " rows. Those have to match before this means anything." ) return lines = ["id,label"] for position, row in enumerate(rows): # Looked up directly rather than through row_field, which would # complain on every row of a file that simply has no id column. fields = row.get("fields") or {} row_id = "" for key in fields: if str(key).strip().lower() == "id": row_id = fields[key] break if not str(row_id).strip(): row_id = "r" + str(position + 1).zfill(2) if isinstance(row_id, float) and row_id == int(row_id): row_id = int(row_id) lines.append(str(row_id) + "," + str(labels[position]).strip()) print("--- submission.csv, copy from here ---") for line in lines: print(line) print("--- to here, " + str(len(lines) - 1) + " rows ---") def _words_in(value): letters = [] for character in str(value).lower(): letters.append(character if character.isalnum() else " ") return [word for word in "".join(letters).split() if word] # Words that turn up in every kind of sentence carry no signal about the# label, so the word matching model looks past them._EVERYDAY_WORDS = set( "a an and are as at be been but by can could did do for from had has have " "he her his i if in is it its me my not of on or our she so than that the " "their them then there they this to too us was we were what when which " "who will with would you your".split()) def _content_words(value): words = _words_in(value) kept = [word for word in words if word not in _EVERYDAY_WORDS] return kept if kept else words def _as_numbers(value): if isinstance(value, (list, tuple)): return [float(item) for item in value] return [float(value)] def _is_numeric(value): try: _as_numbers(value) return True except (TypeError, ValueError): return False def _distance(left, right): """How far apart two examples are. Numbers use straight line distance, text uses how many words the two do not share.""" if _is_numeric(left) and _is_numeric(right): a = _as_numbers(left) b = _as_numbers(right) while len(a) < len(b): a.append(0.0) while len(b) < len(a): b.append(0.0) total = 0.0 for i in range(len(a)): total += (a[i] - b[i]) ** 2 return math.sqrt(total) a = set(_words_in(left)) b = set(_words_in(right)) if not a and not b: return 0.0 shared = len(a & b) return 1.0 - (shared / float(len(a | b))) class Model: """Three small classifiers behind one name. nearest looks for the closest example it was trained on words scores how many words the input shares with each label common always answers with the most common label, the baseline to beat """ def __init__(self, kind="nearest", name="model"): self.kind = kind self.name = name self.data = None self.word_scores = {} def train(self, dataset): self.data = dataset self.word_scores = {} if self.kind == "words": for row in dataset.rows: bucket = self.word_scores.setdefault(row["label"], {}) for word in _content_words(row["features"]): bucket[word] = bucket.get(word, 0) + 1 print("Trained " + self.name + " on " + str(dataset.size()) + " examples.") def _scores(self, features): if self.data is None or not self.data.rows: return {} if self.kind == "common": counts = {} for row in self.data.rows: counts[row["label"]] = counts.get(row["label"], 0) + 1 return counts if self.kind == "words": # Score each label by how much of its training vocabulary shows up # in the input, divided by how much text that label was trained on # so a label with more examples cannot win on volume alone. labels = self.data.labels() scores = {} for label in labels: bucket = self.word_scores.get(label, {}) seen = sum(bucket.values()) or 1 running = 0.0 for word in _content_words(features): running += bucket.get(word, 0) / float(seen) scores[label] = round(running, 6) if sum(scores.values()) == 0: # Nothing in the input was ever seen in training. Say so by # splitting the vote evenly, which reads as low confidence. return dict((label, 1) for label in labels) return scores ranked = sorted(self.data.rows, key=lambda row: _distance(features, row["features"])) neighbours = ranked[: min(3, len(ranked))] scores = {} for row in neighbours: scores[row["label"]] = scores.get(row["label"], 0) + 1 return scores def predict(self, features): scores = self._scores(features) if not scores: return "" return max(scores, key=lambda label: scores[label]) def confidence(self, features): """How much of the vote the winning label took, out of 100.""" scores = self._scores(features) total = sum(scores.values()) if not scores or total == 0: return 0.0 best = max(scores.values()) return round(100.0 * best / float(total), 1) def accuracy(self, dataset): if not dataset.rows: return 0.0 right = 0 for row in dataset.rows: if self.predict(row["features"]) == row["label"]: right += 1 return round(100.0 * right / float(len(dataset.rows)), 1) def show(self): print(self.name + " is a " + self.kind + " model.") if self.data is None: print(" It has not been trained yet.") return print(" Trained on " + str(self.data.size()) + " examples.") print(" Labels it can answer with: " + ", ".join(self.data.labels())) def new_model(kind="nearest", name="model"): return Model(kind, name) # ---------------------------------------------------------------# Your script# --------------------------------------------------------------- deliveries = new_dataset("deliveries")load_csv(deliveries, "couriers-practice.csv", -1, True)training, reserved = deliveries.split(60)validation, final_test = reserved.split(60)candidate = new_model("nearest", "candidate")candidate.train(training)baseline = new_model("common", "baseline")baseline.train(training)print("Training count", training.size())print("Validation count", validation.size())print("Final-test count", final_test.size())print("Candidate validation", candidate.accuracy(validation))print("Baseline validation", baseline.accuracy(validation))if (candidate.accuracy(validation) >= baseline.accuracy(validation)): print("Chosen by validation: closest-example candidate") print("Final-test accuracy", candidate.accuracy(final_test))else: print("Chosen by validation: constant-label baseline") print("Final-test accuracy", baseline.accuracy(final_test))print("Two final rows are a demonstration, not a dependable estimate")Compare this with your version. Different names and personal choices are fine when the program follows the same logic.
Keep your progress
Sign in and every reading, quiz, and exercise you finish is saved.