The baseline you must beat
Choose a fair reference method, compare on the same data, and avoid judging useful rare-class detection by accuracy alone.
A score needs a comparison
Suppose a spam model scores 88% accuracy on a development set of ninety ordinary messages and ten spam messages. A baseline that always says ordinary scores 90% on the same messages. By overall accuracy, the model is worse.
That does not prove it learned nothing or caught no spam. It may catch some spam while also raising too many false alarms. To judge whether it is useful, look at both kinds of mistake and decide which measure matches the job. Accuracy is one comparison, not a complete verdict.
This lesson builds a baseline and a candidate model in blocks, compares them fairly, and then shows why accuracy alone can hide the one class that matters most. The stages use a tiny invented set of support messages, each labelled routine or urgent.
A baseline chosen without peeking
The majority-class baseline, which you met in Level 2, finds the most common label in the training rows and predicts that label for every input, without looking at the input at all.
Build with me · 1
Inspect what constant guessing could exploit
Three training messages are routine and one is urgent. A rule that always says routine can exploit that imbalance without reading the message at all. It is useful precisely because it shows how much success comes just from one label being more common.
Use training labels to choose the constant. Looking at final evaluation labels first would give your method information it should not have.
Create the four training examples and print the most common label.
What to look for
Training majority is routine.
Make it yours
Change one training label and recalculate the counts. Restore the original before comparing the supplied versions.
Build with me · 2
Train the baseline as a real model
The common-label algorithm stores the majority choice when trained. Prediction ignores the incoming features, so both ordinary and urgent messages receive the same label. That is the intended reference behaviour, not a broken text parser.
A useful baseline can be simple enough to describe in one sentence. Its role is to make improvement measurable, not to be impressive by itself.
Select always the most common label, train baseline on training, then ask it about two different inputs.
What to look for
Both predictions are routine.
Make it yours
Try nonsense text. Explain why an unchanged result is expected for this baseline.
Compare on the same evaluation rows
Next come a separate set of messages to test on, and a word-matching candidate to compare with the baseline. Both are scored on exactly the same rows.
Build with me · 3
Create different examples for evaluation
The four development sentences use different wording from the training rows. They retain the same two label definitions. The split is written out here so you can read every row and see that the evaluation answers are never used for training.
These examples are tiny teaching evidence. Once you use their results to choose the candidate or features, call them development rather than an untouched final test.
Create a second dataset named development and enter its four labelled rows. Keep training separate.
What to look for
Development count is 4, including three routine rows and one urgent row.
Make it yours
Identify the urgent development phrase and explain what correct detection would look like.
Build with me · 4
Calculate the baseline result by hand
The baseline gets the three routine messages right and misses the urgent one. Three divided by four is 0.75, so accuracy is 75%. It found zero of the urgent cases even though its overall number looks respectable.
Accuracy answers how often the label matched across all rows. It does not by itself answer whether the important minority class was found.
Put baseline accuracy on development into a labelled say. Compare its output with your manual three-out-of-four count.
What to look for
Baseline accuracy is 75.
Make it yours
Imagine 95 routine and five urgent cases. Calculate the same baseline's accuracy and urgent detection count.
Build with me · 5
Compare the candidate on identical evidence
The candidate word matcher can use associations such as urgent and broken. Both models trained on the same four training rows and are scored on the same four development rows. That shared evidence makes the displayed metric comparable.
A candidate advantage on four invented sentences is a demonstration, not a general performance guarantee. Inspect which examples changed and collect more representative evidence before relying on it.
Print candidate accuracy directly beside baseline accuracy, selecting development in both score blocks.
What to look for
The candidate reaches 100 on this tiny supplied development set while the constant baseline reaches 75.
Make it yours
Try an urgent message without urgent or broken in a later probe. Predict why the apparently perfect candidate can still fail.
Make the comparison reproducible
Write down the training rows, the development rows, the measure (such as accuracy), the baseline, and the candidate. Keep the same preparation and the same evaluation rows for both. A model tested on easy photos and a baseline tested on hard ones are not a fair pair.
The split block in the last lesson gives the same split every time for the same data. So pressing Run again and again is not a test across different splits. To see how much a result depends on the split, set up other training and development groups on purpose, keep a separate final test untouched, and record which split produced each result.
A small dataset gives shaky conclusions. One extra correct answer on a three-row set is a big jump in percentage but very little new evidence. Gather more independent examples before trusting a narrow win.
Rare classes change the question
Imagine one hundred support requests, only five of which are urgent. A baseline that always says routine gets 95% accuracy and finds none of the urgent ones. Now suppose another method flags twelve requests as urgent, and four of those really are urgent. Two questions describe it far better than accuracy does.
Build with me · 6
Count how many urgent cases were found
Recall asks how many actual urgent cases were found. In this worked example there are five urgent requests and the system finds four. Four divided by five, times one hundred, is 80%. The missed fifth case remains important even if many routine cases were easy.
These counts come from the imagined hundred requests above, not from the four-row model you just ran.
Store the two counts, nest found divided by actual inside multiplication by 100, and print the result.
What to look for
Urgent recall percent is 80.0.
Make it yours
Change found_urgent to three. Explain what happened to recall without changing the number of routine requests.
Build with me · 7
Count how many alerts were useful
Precision asks how many flagged requests were truly urgent. Four correct urgent alerts among twelve total alerts gives one third, or about 33.3%. The other eight alerts are false alarms requiring review.
Recall and precision divide by different counts. Recall divides by all actual urgent cases; precision divides by all alerts. Mixing those counts changes the question rather than merely rounding the answer.
Store found_urgent and all_alerts, then build the precision fraction. Keep the labelled output explicit about which metric it represents.
What to look for
Urgent precision percent is about 33.333.
Make it yours
Keep four true alerts but reduce all_alerts to eight. Explain why precision improves while recall from the previous five-urgent example stays unchanged.
Whether that method is useful depends on the cost of checking eight false alarms compared with the cost of missing one urgent request. A method can have lower overall accuracy and still be better at the job the application cares about. Say what that job is, and report the measures that match it, instead of declaring every method below the accuracy baseline useless.
Write the result up
A useful report reads like this: "The constant-label baseline got 10 of 20 development examples right. The model got 13 of 20. It improved glass recognition but still missed half the glasses. Both were tested on the same held-back objects."
That names the methods, the counts, the comparison, and the remaining weakness.
Build with me · 8
Report the method, metric, and limits
A report should name the baseline, candidate, evaluation set, and metric. It should also say what the important errors are and how much evidence supports the result. The supplied example improves from three correct out of four to four out of four, but it is still only four invented observations.
Keep a final set separate after selecting the method. If you repeatedly alter the model using final-test errors, that set has become development evidence and needs to be described honestly.
Run the complete comparison and write a short report using counts as well as percentages. State that the rare-class arithmetic above was a separate illustrative scenario.
What to look for
The reproducible report shows the two scores, counts, and an urgent probe, followed by its limitation.
Make it yours
Propose one new development probe and one genuinely independent final-test collection plan. Explain why they have different jobs.
Before the rehearsal, check that your program can create and train both models, print their scores on the same labelled rows, and save one prediction per unlabelled row. Those skills carry over even when the next dataset has different features and labels.
Full reference solution
This is the complete worked program. Try building it yourself first, then use this reference to find the first place your version behaves differently. The Python below is generated from these exact blocks; helper functions are included so its behaviour can be inspected.
# ---------------------------------------------------------------# Building blocks, written out in plain Python.# This part is generated for you. Your script starts further down.# --------------------------------------------------------------- import mathimport random class Dataset: """A pile of labelled examples. Features can be a number, some text, or a list of numbers. The label is whatever answer you want back.""" def __init__(self, name="dataset"): self.name = name self.rows = [] # Column names from the file's header row, when it came from one. # Without these "the km of a row" has nothing to look the name up in. self.columns = [] def add(self, features, label, fields=None): self.rows.append( { "features": features, "label": str(label), "fields": dict(fields) if fields else {}, } ) def size(self): return len(self.rows) def __len__(self): return len(self.rows) def __iter__(self): """Walking a dataset gives you its rows, so "for each row in list [testing]" reads the way it sounds.""" return iter(self.rows) def labels(self): seen = [] for row in self.rows: if row["label"] not in seen: seen.append(row["label"]) return seen def most_common_label(self): if not self.rows: return "" counts = {} for row in self.rows: counts[row["label"]] = counts.get(row["label"], 0) + 1 return max(counts, key=lambda label: counts[label]) def split(self, train_percent=80): """Keeps the given percent for training and hands back the rest as a test set. The shuffle is seeded, so you get the same split every run.""" order = list(range(len(self.rows))) random.Random(0).shuffle(order) cut = int(len(order) * train_percent / 100) train = Dataset(self.name + " (train)") test = Dataset(self.name + " (test)") train.columns = list(self.columns) test.columns = list(self.columns) for position, index in enumerate(order): row = self.rows[index] target = train if position < cut else test target.add(row["features"], row["label"], row.get("fields")) return train, test def show(self, limit=10): print(self.name + ": " + str(len(self.rows)) + " examples") for row in self.rows[:limit]: print(" " + str(row["features"]) + " -> " + row["label"]) if len(self.rows) > limit: print(" ... and " + str(len(self.rows) - limit) + " more") def new_dataset(name="dataset"): return Dataset(name) def row_field(row, name): """One named piece of a row: "the km of this row", "the label of it". Names come from the header line of the csv. "label" always works, even on a file with no header, because every row has one.""" wanted = str(name).strip() if not isinstance(row, dict): print("That is not a row. Use this inside a for each over a dataset.") return "" if wanted.lower() == "label": return row.get("label", "") fields = row.get("fields") or {} if wanted in fields: return fields[wanted] # Header names are matched loosely, so "Rain" finds the "rain" column. for key in fields: if str(key).strip().lower() == wanted.lower(): return fields[key] known = ", ".join([str(k) for k in fields]) if fields else "none" print( "No column called " + wanted + " in this row. Columns here: " + known + "." ) return "" def load_csv(dataset, path, label_column=-1, has_header=True): """Reads a comma separated file into a dataset. Everything except the label column becomes the features, and anything that looks like a number is turned into one. This is how a file you uploaded becomes something you can train on.""" try: with open(path) as handle: rows = [line.rstrip("\n").rstrip("\r") for line in handle] except OSError: print("Could not find " + path + ". Check the name in the Files panel.") return dataset rows = [row for row in rows if row.strip()] # A file with no label column is a perfectly normal thing to load: it is # what a test set looks like before you have predicted anything. labelled = str(label_column).strip().lower() not in ("none", "", "no", "-") header = [] if has_header and rows: header = [cell.strip() for cell in rows[0].split(",")] rows = rows[1:] added = 0 for row in rows: cells = [cell.strip() for cell in row.split(",")] if not cells or (labelled and len(cells) < 2): continue index = -1 if labelled: index = int(label_column) if index < 0: index = len(cells) + index if index < 0 or index >= len(cells): continue label = cells[index] if labelled else "" keep = [i for i in range(len(cells)) if i != index] typed = [] for i in keep: try: typed.append(float(cells[i])) except ValueError: typed.append(cells[i]) # Every kept column gets its header name, so "the km of a row" works. fields = {} for position, i in enumerate(keep): if i < len(header) and header[i]: fields[header[i]] = typed[position] if header and not dataset.columns: dataset.columns = [header[i] for i in keep if i < len(header)] dataset.add(typed[0] if len(typed) == 1 else typed, label, fields) added += 1 print("Loaded " + str(added) + " rows from " + path + ".") return dataset def save_submission(predictions, dataset, path="submission.csv"): """Prints your predictions in the exact two column format a round is scored in: a header, then one id and one label per line. It is printed rather than saved to a file because the console is the one place you can copy it from. Paste it into a new file in the Files panel, or straight into the upload box.""" labels = list(predictions or []) rows = list(dataset) if dataset is not None else [] if len(labels) != len(rows): print( "You have " + str(len(labels)) + " predictions for " + str(len(rows)) + " rows. Those have to match before this means anything." ) return lines = ["id,label"] for position, row in enumerate(rows): # Looked up directly rather than through row_field, which would # complain on every row of a file that simply has no id column. fields = row.get("fields") or {} row_id = "" for key in fields: if str(key).strip().lower() == "id": row_id = fields[key] break if not str(row_id).strip(): row_id = "r" + str(position + 1).zfill(2) if isinstance(row_id, float) and row_id == int(row_id): row_id = int(row_id) lines.append(str(row_id) + "," + str(labels[position]).strip()) print("--- submission.csv, copy from here ---") for line in lines: print(line) print("--- to here, " + str(len(lines) - 1) + " rows ---") def _words_in(value): letters = [] for character in str(value).lower(): letters.append(character if character.isalnum() else " ") return [word for word in "".join(letters).split() if word] # Words that turn up in every kind of sentence carry no signal about the# label, so the word matching model looks past them._EVERYDAY_WORDS = set( "a an and are as at be been but by can could did do for from had has have " "he her his i if in is it its me my not of on or our she so than that the " "their them then there they this to too us was we were what when which " "who will with would you your".split()) def _content_words(value): words = _words_in(value) kept = [word for word in words if word not in _EVERYDAY_WORDS] return kept if kept else words def _as_numbers(value): if isinstance(value, (list, tuple)): return [float(item) for item in value] return [float(value)] def _is_numeric(value): try: _as_numbers(value) return True except (TypeError, ValueError): return False def _distance(left, right): """How far apart two examples are. Numbers use straight line distance, text uses how many words the two do not share.""" if _is_numeric(left) and _is_numeric(right): a = _as_numbers(left) b = _as_numbers(right) while len(a) < len(b): a.append(0.0) while len(b) < len(a): b.append(0.0) total = 0.0 for i in range(len(a)): total += (a[i] - b[i]) ** 2 return math.sqrt(total) a = set(_words_in(left)) b = set(_words_in(right)) if not a and not b: return 0.0 shared = len(a & b) return 1.0 - (shared / float(len(a | b))) class Model: """Three small classifiers behind one name. nearest looks for the closest example it was trained on words scores how many words the input shares with each label common always answers with the most common label, the baseline to beat """ def __init__(self, kind="nearest", name="model"): self.kind = kind self.name = name self.data = None self.word_scores = {} def train(self, dataset): self.data = dataset self.word_scores = {} if self.kind == "words": for row in dataset.rows: bucket = self.word_scores.setdefault(row["label"], {}) for word in _content_words(row["features"]): bucket[word] = bucket.get(word, 0) + 1 print("Trained " + self.name + " on " + str(dataset.size()) + " examples.") def _scores(self, features): if self.data is None or not self.data.rows: return {} if self.kind == "common": counts = {} for row in self.data.rows: counts[row["label"]] = counts.get(row["label"], 0) + 1 return counts if self.kind == "words": # Score each label by how much of its training vocabulary shows up # in the input, divided by how much text that label was trained on # so a label with more examples cannot win on volume alone. labels = self.data.labels() scores = {} for label in labels: bucket = self.word_scores.get(label, {}) seen = sum(bucket.values()) or 1 running = 0.0 for word in _content_words(features): running += bucket.get(word, 0) / float(seen) scores[label] = round(running, 6) if sum(scores.values()) == 0: # Nothing in the input was ever seen in training. Say so by # splitting the vote evenly, which reads as low confidence. return dict((label, 1) for label in labels) return scores ranked = sorted(self.data.rows, key=lambda row: _distance(features, row["features"])) neighbours = ranked[: min(3, len(ranked))] scores = {} for row in neighbours: scores[row["label"]] = scores.get(row["label"], 0) + 1 return scores def predict(self, features): scores = self._scores(features) if not scores: return "" return max(scores, key=lambda label: scores[label]) def confidence(self, features): """How much of the vote the winning label took, out of 100.""" scores = self._scores(features) total = sum(scores.values()) if not scores or total == 0: return 0.0 best = max(scores.values()) return round(100.0 * best / float(total), 1) def accuracy(self, dataset): if not dataset.rows: return 0.0 right = 0 for row in dataset.rows: if self.predict(row["features"]) == row["label"]: right += 1 return round(100.0 * right / float(len(dataset.rows)), 1) def show(self): print(self.name + " is a " + self.kind + " model.") if self.data is None: print(" It has not been trained yet.") return print(" Trained on " + str(self.data.size()) + " examples.") print(" Labels it can answer with: " + ", ".join(self.data.labels())) def new_model(kind="nearest", name="model"): return Model(kind, name) # ---------------------------------------------------------------# Your script# --------------------------------------------------------------- training = new_dataset("training")training.add("calm parcel", "routine")training.add("ordinary parcel", "routine")training.add("usual delivery", "routine")training.add("urgent broken package", "urgent")development = new_dataset("development")development.add("ordinary delivery", "routine")development.add("calm delivery", "routine")development.add("usual parcel", "routine")development.add("urgent broken delivery", "urgent")baseline = new_model("common", "baseline")baseline.train(training)candidate = new_model("words", "candidate")candidate.train(training)print("Training rows", training.size())print("Development rows", development.size())print("Baseline accuracy", baseline.accuracy(development))print("Candidate accuracy", candidate.accuracy(development))print("Urgent probe", candidate.predict("urgent broken delivery"))print("This is a four-row development comparison, not a final deployment claim")Compare this with your version. Different names and personal choices are fine when the program follows the same logic.
Keep your progress
Sign in and every reading, quiz, and exercise you finish is saved.