0%
BuildRound 3 rehearsal: predict from a tableabout 28 min, 8 steps

The baseline you must beat

Choose a fair reference method, compare on the same data, and avoid judging useful rare-class detection by accuracy alone.

A score needs a comparison

Suppose a spam model scores 88% accuracy on a development set of ninety ordinary messages and ten spam messages. A baseline that always says ordinary scores 90% on the same messages. By overall accuracy, the model is worse.

That does not prove it learned nothing or caught no spam. It may catch some spam while also raising too many false alarms. To judge whether it is useful, look at both kinds of mistake and decide which measure matches the job. Accuracy is one comparison, not a complete verdict.

This lesson builds a baseline and a candidate model in blocks, compares them fairly, and then shows why accuracy alone can hide the one class that matters most. The stages use a tiny invented set of support messages, each labelled routine or urgent.

A baseline chosen without peeking

The majority-class baseline, which you met in Level 2, finds the most common label in the training rows and predicts that label for every input, without looking at the input at all.

Build with me · 1

Inspect what constant guessing could exploit

Three training messages are routine and one is urgent. A rule that always says routine can exploit that imbalance without reading the message at all. It is useful precisely because it shows how much success comes just from one label being more common.

Use training labels to choose the constant. Looking at final evaluation labels first would give your method information it should not have.

Blocks at this stageWorked example
make a dataset calledtraining
add totrainingthe examplecalm parcellabeledroutine
add totrainingthe exampleordinary parcellabeledroutine
add totrainingthe exampleusual deliverylabeledroutine
add totrainingthe exampleurgent broken packagelabeledurgent
show datasettraining
sayTraining majoritythen
the most common label intraining

Create the four training examples and print the most common label.

What to look for

Training majority is routine.

Make it yours

Change one training label and recalculate the counts. Restore the original before comparing the supplied versions.

Build with me · 2

Train the baseline as a real model

The common-label algorithm stores the majority choice when trained. Prediction ignores the incoming features, so both ordinary and urgent messages receive the same label. That is the intended reference behaviour, not a broken text parser.

A useful baseline can be simple enough to describe in one sentence. Its role is to make improvement measurable, not to be impressive by itself.

Blocks at this stageWorked example
make a dataset calledtraining
add totrainingthe examplecalm parcellabeledroutine
add totrainingthe exampleordinary parcellabeledroutine
add totrainingthe exampleusual deliverylabeledroutine
add totrainingthe exampleurgent broken packagelabeledurgent
make aalways the most common labelmodel calledbaseline
trainbaselineontraining
sayOrdinary inputthen
whatbaselinesays aboutordinary delivery
sayUrgent inputthen
whatbaselinesays abouturgent broken package

Select always the most common label, train baseline on training, then ask it about two different inputs.

What to look for

Both predictions are routine.

Make it yours

Try nonsense text. Explain why an unchanged result is expected for this baseline.

Compare on the same evaluation rows

Next come a separate set of messages to test on, and a word-matching candidate to compare with the baseline. Both are scored on exactly the same rows.

Build with me · 3

Create different examples for evaluation

The four development sentences use different wording from the training rows. They retain the same two label definitions. The split is written out here so you can read every row and see that the evaluation answers are never used for training.

These examples are tiny teaching evidence. Once you use their results to choose the candidate or features, call them development rather than an untouched final test.

Blocks at this stageWorked example
make a dataset calledtraining
add totrainingthe examplecalm parcellabeledroutine
add totrainingthe exampleordinary parcellabeledroutine
add totrainingthe exampleusual deliverylabeledroutine
add totrainingthe exampleurgent broken packagelabeledurgent
make a dataset calleddevelopment
add todevelopmentthe exampleordinary deliverylabeledroutine
add todevelopmentthe examplecalm deliverylabeledroutine
add todevelopmentthe exampleusual parcellabeledroutine
add todevelopmentthe exampleurgent broken deliverylabeledurgent
show datasetdevelopment
sayDevelopment countthen
how many examples indevelopment

Create a second dataset named development and enter its four labelled rows. Keep training separate.

What to look for

Development count is 4, including three routine rows and one urgent row.

Make it yours

Identify the urgent development phrase and explain what correct detection would look like.

Build with me · 4

Calculate the baseline result by hand

The baseline gets the three routine messages right and misses the urgent one. Three divided by four is 0.75, so accuracy is 75%. It found zero of the urgent cases even though its overall number looks respectable.

Accuracy answers how often the label matched across all rows. It does not by itself answer whether the important minority class was found.

Blocks at this stageWorked example
make a dataset calledtraining
add totrainingthe examplecalm parcellabeledroutine
add totrainingthe exampleordinary parcellabeledroutine
add totrainingthe exampleusual deliverylabeledroutine
add totrainingthe exampleurgent broken packagelabeledurgent
make a dataset calleddevelopment
add todevelopmentthe exampleordinary deliverylabeledroutine
add todevelopmentthe examplecalm deliverylabeledroutine
add todevelopmentthe exampleusual parcellabeledroutine
add todevelopmentthe exampleurgent broken deliverylabeledurgent
make aalways the most common labelmodel calledbaseline
trainbaselineontraining
make aword matchingmodel calledcandidate
traincandidateontraining
sayBaseline accuracythen
out of 100, the accuracy ofbaselineondevelopment

Put baseline accuracy on development into a labelled say. Compare its output with your manual three-out-of-four count.

What to look for

Baseline accuracy is 75.

Make it yours

Imagine 95 routine and five urgent cases. Calculate the same baseline's accuracy and urgent detection count.

Build with me · 5

Compare the candidate on identical evidence

The candidate word matcher can use associations such as urgent and broken. Both models trained on the same four training rows and are scored on the same four development rows. That shared evidence makes the displayed metric comparable.

A candidate advantage on four invented sentences is a demonstration, not a general performance guarantee. Inspect which examples changed and collect more representative evidence before relying on it.

Blocks at this stageWorked example
make a dataset calledtraining
add totrainingthe examplecalm parcellabeledroutine
add totrainingthe exampleordinary parcellabeledroutine
add totrainingthe exampleusual deliverylabeledroutine
add totrainingthe exampleurgent broken packagelabeledurgent
make a dataset calleddevelopment
add todevelopmentthe exampleordinary deliverylabeledroutine
add todevelopmentthe examplecalm deliverylabeledroutine
add todevelopmentthe exampleusual parcellabeledroutine
add todevelopmentthe exampleurgent broken deliverylabeledurgent
make aalways the most common labelmodel calledbaseline
trainbaselineontraining
make aword matchingmodel calledcandidate
traincandidateontraining
sayBaseline accuracythen
out of 100, the accuracy ofbaselineondevelopment
sayCandidate accuracythen
out of 100, the accuracy ofcandidateondevelopment

Print candidate accuracy directly beside baseline accuracy, selecting development in both score blocks.

What to look for

The candidate reaches 100 on this tiny supplied development set while the constant baseline reaches 75.

Make it yours

Try an urgent message without urgent or broken in a later probe. Predict why the apparently perfect candidate can still fail.

Make the comparison reproducible

Write down the training rows, the development rows, the measure (such as accuracy), the baseline, and the candidate. Keep the same preparation and the same evaluation rows for both. A model tested on easy photos and a baseline tested on hard ones are not a fair pair.

The split block in the last lesson gives the same split every time for the same data. So pressing Run again and again is not a test across different splits. To see how much a result depends on the split, set up other training and development groups on purpose, keep a separate final test untouched, and record which split produced each result.

A small dataset gives shaky conclusions. One extra correct answer on a three-row set is a big jump in percentage but very little new evidence. Gather more independent examples before trusting a narrow win.

Rare classes change the question

Imagine one hundred support requests, only five of which are urgent. A baseline that always says routine gets 95% accuracy and finds none of the urgent ones. Now suppose another method flags twelve requests as urgent, and four of those really are urgent. Two questions describe it far better than accuracy does.

Build with me · 6

Count how many urgent cases were found

Recall asks how many actual urgent cases were found. In this worked example there are five urgent requests and the system finds four. Four divided by five, times one hundred, is 80%. The missed fifth case remains important even if many routine cases were easy.

These counts come from the imagined hundred requests above, not from the four-row model you just ran.

Blocks at this stageWorked example
settrue_urgentto5
setfound_urgentto4
sayUrgent recall percentthen
found_urgent
/
true_urgent
x100

Store the two counts, nest found divided by actual inside multiplication by 100, and print the result.

What to look for

Urgent recall percent is 80.0.

Make it yours

Change found_urgent to three. Explain what happened to recall without changing the number of routine requests.

Build with me · 7

Count how many alerts were useful

Precision asks how many flagged requests were truly urgent. Four correct urgent alerts among twelve total alerts gives one third, or about 33.3%. The other eight alerts are false alarms requiring review.

Recall and precision divide by different counts. Recall divides by all actual urgent cases; precision divides by all alerts. Mixing those counts changes the question rather than merely rounding the answer.

Blocks at this stageWorked example
setfound_urgentto4
setall_alertsto12
sayUrgent precision percentthen
round
found_urgent
/
all_alerts
x100
to3places

Store found_urgent and all_alerts, then build the precision fraction. Keep the labelled output explicit about which metric it represents.

What to look for

Urgent precision percent is about 33.333.

Make it yours

Keep four true alerts but reduce all_alerts to eight. Explain why precision improves while recall from the previous five-urgent example stays unchanged.

Whether that method is useful depends on the cost of checking eight false alarms compared with the cost of missing one urgent request. A method can have lower overall accuracy and still be better at the job the application cares about. Say what that job is, and report the measures that match it, instead of declaring every method below the accuracy baseline useless.

Write the result up

A useful report reads like this: "The constant-label baseline got 10 of 20 development examples right. The model got 13 of 20. It improved glass recognition but still missed half the glasses. Both were tested on the same held-back objects."

That names the methods, the counts, the comparison, and the remaining weakness.

Build with me · 8

Report the method, metric, and limits

A report should name the baseline, candidate, evaluation set, and metric. It should also say what the important errors are and how much evidence supports the result. The supplied example improves from three correct out of four to four out of four, but it is still only four invented observations.

Keep a final set separate after selecting the method. If you repeatedly alter the model using final-test errors, that set has become development evidence and needs to be described honestly.

Blocks at this stageWorked example
make a dataset calledtraining
add totrainingthe examplecalm parcellabeledroutine
add totrainingthe exampleordinary parcellabeledroutine
add totrainingthe exampleusual deliverylabeledroutine
add totrainingthe exampleurgent broken packagelabeledurgent
make a dataset calleddevelopment
add todevelopmentthe exampleordinary deliverylabeledroutine
add todevelopmentthe examplecalm deliverylabeledroutine
add todevelopmentthe exampleusual parcellabeledroutine
add todevelopmentthe exampleurgent broken deliverylabeledurgent
make aalways the most common labelmodel calledbaseline
trainbaselineontraining
make aword matchingmodel calledcandidate
traincandidateontraining
sayTraining rowsthen
how many examples intraining
sayDevelopment rowsthen
how many examples indevelopment
sayBaseline accuracythen
out of 100, the accuracy ofbaselineondevelopment
sayCandidate accuracythen
out of 100, the accuracy ofcandidateondevelopment
sayUrgent probethen
whatcandidatesays abouturgent broken delivery
sayThis is a four-row development comparison, not a final deployment claim

Run the complete comparison and write a short report using counts as well as percentages. State that the rare-class arithmetic above was a separate illustrative scenario.

What to look for

The reproducible report shows the two scores, counts, and an urgent probe, followed by its limitation.

Make it yours

Propose one new development probe and one genuinely independent final-test collection plan. Explain why they have different jobs.

Before the rehearsal, check that your program can create and train both models, print their scores on the same labelled rows, and save one prediction per unlabelled row. Those skills carry over even when the next dataset has different features and labels.

Full reference solution

This is the complete worked program. Try building it yourself first, then use this reference to find the first place your version behaves differently. The Python below is generated from these exact blocks; helper functions are included so its behaviour can be inspected.

make a dataset calledtraining
add totrainingthe examplecalm parcellabeledroutine
add totrainingthe exampleordinary parcellabeledroutine
add totrainingthe exampleusual deliverylabeledroutine
add totrainingthe exampleurgent broken packagelabeledurgent
make a dataset calleddevelopment
add todevelopmentthe exampleordinary deliverylabeledroutine
add todevelopmentthe examplecalm deliverylabeledroutine
add todevelopmentthe exampleusual parcellabeledroutine
add todevelopmentthe exampleurgent broken deliverylabeledurgent
make aalways the most common labelmodel calledbaseline
trainbaselineontraining
make aword matchingmodel calledcandidate
traincandidateontraining
sayTraining rowsthen
how many examples intraining
sayDevelopment rowsthen
how many examples indevelopment
sayBaseline accuracythen
out of 100, the accuracy ofbaselineondevelopment
sayCandidate accuracythen
out of 100, the accuracy ofcandidateondevelopment
sayUrgent probethen
whatcandidatesays abouturgent broken delivery
sayThis is a four-row development comparison, not a final deployment claim
PythonHover over a line to see an explanation
# ---------------------------------------------------------------# Building blocks, written out in plain Python.# This part is generated for you. Your script starts further down.# ---------------------------------------------------------------  import mathimport random  class Dataset:    """A pile of labelled examples. Features can be a number, some text, or a    list of numbers. The label is whatever answer you want back."""     def __init__(self, name="dataset"):        self.name = name        self.rows = []        # Column names from the file's header row, when it came from one.        # Without these "the km of a row" has nothing to look the name up in.        self.columns = []     def add(self, features, label, fields=None):        self.rows.append(            {                "features": features,                "label": str(label),                "fields": dict(fields) if fields else {},            }        )     def size(self):        return len(self.rows)     def __len__(self):        return len(self.rows)     def __iter__(self):        """Walking a dataset gives you its rows, so "for each row in list        [testing]" reads the way it sounds."""        return iter(self.rows)     def labels(self):        seen = []        for row in self.rows:            if row["label"] not in seen:                seen.append(row["label"])        return seen     def most_common_label(self):        if not self.rows:            return ""        counts = {}        for row in self.rows:            counts[row["label"]] = counts.get(row["label"], 0) + 1        return max(counts, key=lambda label: counts[label])     def split(self, train_percent=80):        """Keeps the given percent for training and hands back the rest as a        test set. The shuffle is seeded, so you get the same split every run."""        order = list(range(len(self.rows)))        random.Random(0).shuffle(order)        cut = int(len(order) * train_percent / 100)        train = Dataset(self.name + " (train)")        test = Dataset(self.name + " (test)")        train.columns = list(self.columns)        test.columns = list(self.columns)        for position, index in enumerate(order):            row = self.rows[index]            target = train if position < cut else test            target.add(row["features"], row["label"], row.get("fields"))        return train, test     def show(self, limit=10):        print(self.name + ": " + str(len(self.rows)) + " examples")        for row in self.rows[:limit]:            print("  " + str(row["features"]) + "  ->  " + row["label"])        if len(self.rows) > limit:            print("  ... and " + str(len(self.rows) - limit) + " more")  def new_dataset(name="dataset"):    return Dataset(name)  def row_field(row, name):    """One named piece of a row: "the km of this row", "the label of it".     Names come from the header line of the csv. "label" always works, even on    a file with no header, because every row has one."""    wanted = str(name).strip()    if not isinstance(row, dict):        print("That is not a row. Use this inside a for each over a dataset.")        return ""    if wanted.lower() == "label":        return row.get("label", "")    fields = row.get("fields") or {}    if wanted in fields:        return fields[wanted]    # Header names are matched loosely, so "Rain" finds the "rain" column.    for key in fields:        if str(key).strip().lower() == wanted.lower():            return fields[key]    known = ", ".join([str(k) for k in fields]) if fields else "none"    print(        "No column called "        + wanted        + " in this row. Columns here: "        + known        + "."    )    return ""  def load_csv(dataset, path, label_column=-1, has_header=True):    """Reads a comma separated file into a dataset.     Everything except the label column becomes the features, and anything that    looks like a number is turned into one. This is how a file you uploaded    becomes something you can train on."""    try:        with open(path) as handle:            rows = [line.rstrip("\n").rstrip("\r") for line in handle]    except OSError:        print("Could not find " + path + ". Check the name in the Files panel.")        return dataset     rows = [row for row in rows if row.strip()]     # A file with no label column is a perfectly normal thing to load: it is    # what a test set looks like before you have predicted anything.    labelled = str(label_column).strip().lower() not in ("none", "", "no", "-")     header = []    if has_header and rows:        header = [cell.strip() for cell in rows[0].split(",")]        rows = rows[1:]     added = 0    for row in rows:        cells = [cell.strip() for cell in row.split(",")]        if not cells or (labelled and len(cells) < 2):            continue         index = -1        if labelled:            index = int(label_column)            if index < 0:                index = len(cells) + index            if index < 0 or index >= len(cells):                continue         label = cells[index] if labelled else ""        keep = [i for i in range(len(cells)) if i != index]         typed = []        for i in keep:            try:                typed.append(float(cells[i]))            except ValueError:                typed.append(cells[i])         # Every kept column gets its header name, so "the km of a row" works.        fields = {}        for position, i in enumerate(keep):            if i < len(header) and header[i]:                fields[header[i]] = typed[position]        if header and not dataset.columns:            dataset.columns = [header[i] for i in keep if i < len(header)]         dataset.add(typed[0] if len(typed) == 1 else typed, label, fields)        added += 1     print("Loaded " + str(added) + " rows from " + path + ".")    return dataset  def save_submission(predictions, dataset, path="submission.csv"):    """Prints your predictions in the exact two column format a round is    scored in: a header, then one id and one label per line.     It is printed rather than saved to a file because the console is the one    place you can copy it from. Paste it into a new file in the Files panel,    or straight into the upload box."""    labels = list(predictions or [])    rows = list(dataset) if dataset is not None else []     if len(labels) != len(rows):        print(            "You have "            + str(len(labels))            + " predictions for "            + str(len(rows))            + " rows. Those have to match before this means anything."        )        return     lines = ["id,label"]    for position, row in enumerate(rows):        # Looked up directly rather than through row_field, which would        # complain on every row of a file that simply has no id column.        fields = row.get("fields") or {}        row_id = ""        for key in fields:            if str(key).strip().lower() == "id":                row_id = fields[key]                break        if not str(row_id).strip():            row_id = "r" + str(position + 1).zfill(2)        if isinstance(row_id, float) and row_id == int(row_id):            row_id = int(row_id)        lines.append(str(row_id) + "," + str(labels[position]).strip())     print("--- submission.csv, copy from here ---")    for line in lines:        print(line)    print("--- to here, " + str(len(lines) - 1) + " rows ---")  def _words_in(value):    letters = []    for character in str(value).lower():        letters.append(character if character.isalnum() else " ")    return [word for word in "".join(letters).split() if word]  # Words that turn up in every kind of sentence carry no signal about the# label, so the word matching model looks past them._EVERYDAY_WORDS = set(    "a an and are as at be been but by can could did do for from had has have "    "he her his i if in is it its me my not of on or our she so than that the "    "their them then there they this to too us was we were what when which "    "who will with would you your".split())  def _content_words(value):    words = _words_in(value)    kept = [word for word in words if word not in _EVERYDAY_WORDS]    return kept if kept else words  def _as_numbers(value):    if isinstance(value, (list, tuple)):        return [float(item) for item in value]    return [float(value)]  def _is_numeric(value):    try:        _as_numbers(value)        return True    except (TypeError, ValueError):        return False  def _distance(left, right):    """How far apart two examples are. Numbers use straight line distance,    text uses how many words the two do not share."""    if _is_numeric(left) and _is_numeric(right):        a = _as_numbers(left)        b = _as_numbers(right)        while len(a) < len(b):            a.append(0.0)        while len(b) < len(a):            b.append(0.0)        total = 0.0        for i in range(len(a)):            total += (a[i] - b[i]) ** 2        return math.sqrt(total)    a = set(_words_in(left))    b = set(_words_in(right))    if not a and not b:        return 0.0    shared = len(a & b)    return 1.0 - (shared / float(len(a | b)))  class Model:    """Three small classifiers behind one name.     nearest  looks for the closest example it was trained on    words    scores how many words the input shares with each label    common   always answers with the most common label, the baseline to beat    """     def __init__(self, kind="nearest", name="model"):        self.kind = kind        self.name = name        self.data = None        self.word_scores = {}     def train(self, dataset):        self.data = dataset        self.word_scores = {}        if self.kind == "words":            for row in dataset.rows:                bucket = self.word_scores.setdefault(row["label"], {})                for word in _content_words(row["features"]):                    bucket[word] = bucket.get(word, 0) + 1        print("Trained " + self.name + " on " + str(dataset.size()) + " examples.")     def _scores(self, features):        if self.data is None or not self.data.rows:            return {}        if self.kind == "common":            counts = {}            for row in self.data.rows:                counts[row["label"]] = counts.get(row["label"], 0) + 1            return counts        if self.kind == "words":            # Score each label by how much of its training vocabulary shows up            # in the input, divided by how much text that label was trained on            # so a label with more examples cannot win on volume alone.            labels = self.data.labels()            scores = {}            for label in labels:                bucket = self.word_scores.get(label, {})                seen = sum(bucket.values()) or 1                running = 0.0                for word in _content_words(features):                    running += bucket.get(word, 0) / float(seen)                scores[label] = round(running, 6)            if sum(scores.values()) == 0:                # Nothing in the input was ever seen in training. Say so by                # splitting the vote evenly, which reads as low confidence.                return dict((label, 1) for label in labels)            return scores        ranked = sorted(self.data.rows, key=lambda row: _distance(features, row["features"]))        neighbours = ranked[: min(3, len(ranked))]        scores = {}        for row in neighbours:            scores[row["label"]] = scores.get(row["label"], 0) + 1        return scores     def predict(self, features):        scores = self._scores(features)        if not scores:            return ""        return max(scores, key=lambda label: scores[label])     def confidence(self, features):        """How much of the vote the winning label took, out of 100."""        scores = self._scores(features)        total = sum(scores.values())        if not scores or total == 0:            return 0.0        best = max(scores.values())        return round(100.0 * best / float(total), 1)     def accuracy(self, dataset):        if not dataset.rows:            return 0.0        right = 0        for row in dataset.rows:            if self.predict(row["features"]) == row["label"]:                right += 1        return round(100.0 * right / float(len(dataset.rows)), 1)     def show(self):        print(self.name + " is a " + self.kind + " model.")        if self.data is None:            print("  It has not been trained yet.")            return        print("  Trained on " + str(self.data.size()) + " examples.")        print("  Labels it can answer with: " + ", ".join(self.data.labels()))  def new_model(kind="nearest", name="model"):    return Model(kind, name)  # ---------------------------------------------------------------# Your script# ---------------------------------------------------------------  training = new_dataset("training")training.add("calm parcel", "routine")training.add("ordinary parcel", "routine")training.add("usual delivery", "routine")training.add("urgent broken package", "urgent")development = new_dataset("development")development.add("ordinary delivery", "routine")development.add("calm delivery", "routine")development.add("usual parcel", "routine")development.add("urgent broken delivery", "urgent")baseline = new_model("common", "baseline")baseline.train(training)candidate = new_model("words", "candidate")candidate.train(training)print("Training rows", training.size())print("Development rows", development.size())print("Baseline accuracy", baseline.accuracy(development))print("Candidate accuracy", candidate.accuracy(development))print("Urgent probe", candidate.predict("urgent broken delivery"))print("This is a four-row development comparison, not a final deployment claim")

Compare this with your version. Different names and personal choices are fine when the program follows the same logic.

Keep your progress

Sign in and every reading, quiz, and exercise you finish is saved.

Sign in