0%
BuildYour first AI appabout 29 min, 8 steps

What the model did with your words

Inspect the exact limits of the course word matcher and understand why fluent-looking labels can still be wrong.

Open the black box a little

The word-matching model you trained in the last lesson is small enough to explain completely. It does three things.

First, it tidies each message. It makes every letter lower case, splits the text into words at spaces and punctuation, and drops very common words such as the, is, and can. If dropping them would leave nothing, it keeps all the words.

Second, during training, it counts how often each remaining word appears under each label.

Third, during prediction, it looks up the words of the new message. Every word that appeared under a label adds support for that label. Each label's support is divided by the total number of words that label was trained on, so a label cannot win just because its examples were longer. The label with the most support wins, and the score says what share of all the support it received.

That is the whole model. It does not look at word order. It does not treat a question mark as a clue. It does not check which word comes first. If a lesson claimed this model had learned "questions start with is", it would be describing a different model.

Work through a tiny collection

Sixteen examples are too many to count by hand, so the stages below use a separate four-row dataset. Keep this table in view while you work through them:

TextLabel
red applefruit
green applefruit
blue trainvehicle
fast trainvehicle

Before each stage, work out the answer on paper, then run the program to check.

Build with me · 1

Make every training association visible

Read every word under fruit: red, apple, green, apple. There are four occurrences in total. Read vehicle: blue, train, fast, train. It also has four. The repeated words apple and train each contribute two occurrences in their class.

The model lower-cases the text and counts words. It is not storing a dictionary definition of fruit or a photograph of a train. This tiny collection lets us predict which associations the implementation can use.

Blocks at this stageWorked example
make a dataset calledexamples
add toexamplesthe examplered applelabeledfruit
add toexamplesthe examplegreen applelabeledfruit
add toexamplesthe exampleblue trainlabeledvehicle
add toexamplesthe examplefast trainlabeledvehicle
make aword matchingmodel calledmatcher
trainmatcheronexamples
show datasetexamples

Build the four add-example blocks, then create and train matcher. Keep the dataset inspection at the bottom.

What to look for

Four labelled rows are printed, and training uses four examples.

Make it yours

Replace red with yellow in exactly one row. Which prediction input would directly test that edit?

Build with me · 2

Calculate the support for one word

For apple, fruit receives two matching occurrences divided by its four total words: 2/4. Vehicle receives zero. The displayed confidence divides the winning support by all support, so 0.5 divided by 0.5 is 100%.

That number describes this counting rule on this input. It is not a promise that the real meaning is fruit.

Blocks at this stageWorked example
make a dataset calledexamples
add toexamplesthe examplered applelabeledfruit
add toexamplesthe examplegreen applelabeledfruit
add toexamplesthe exampleblue trainlabeledvehicle
add toexamplesthe examplefast trainlabeledvehicle
make aword matchingmodel calledmatcher
trainmatcheronexamples
setmessagetoapple
sayInputthen
message
sayLabelthen
whatmatchersays about
message
sayScorethen
out of 100, how surematcheris about
message

Store apple under message. Read that same variable in the label and confidence blocks.

What to look for

Label is fruit and Score is 100.

Make it yours

Try train and work out the analogous calculation before running.

Build with me · 3

Combine evidence for both labels

Red occurs once under fruit, giving 1/4 support. Train occurs twice under vehicle, giving 2/4 support. Vehicle receives two thirds of the total support, about 66.7%.

Nothing in these counts checks whether red modifies train. Words contribute individually. The winning label therefore follows the associations, not a grammatical understanding of the phrase.

Blocks at this stageWorked example
make a dataset calledexamples
add toexamplesthe examplered applelabeledfruit
add toexamplesthe examplegreen applelabeledfruit
add toexamplesthe exampleblue trainlabeledvehicle
add toexamplesthe examplefast trainlabeledvehicle
make aword matchingmodel calledmatcher
trainmatcheronexamples
setmessagetored train
sayInputthen
message
sayLabelthen
whatmatchersays about
message
sayScorethen
out of 100, how surematcheris about
message

Change only the stored input to red train. Keep the four-row training stack unchanged.

What to look for

Vehicle wins, with about 66.7 out of 100.

Make it yours

Try apple train. Each side now has two matching occurrences; explain why the score becomes 50.

Build with me · 4

Show what word order cannot change

Swapping the order preserves the words and their counts. Because the model only counts words and ignores their order, both inputs receive the same label and score.

A model cannot use information it throws away. Adding a thousand copies of these same rows would not teach this particular matcher to distinguish two inputs represented by the same counts.

Blocks at this stageWorked example
make a dataset calledexamples
add toexamplesthe examplered applelabeledfruit
add toexamplesthe examplegreen applelabeledfruit
add toexamplesthe exampleblue trainlabeledvehicle
add toexamplesthe examplefast trainlabeledvehicle
make aword matchingmodel calledmatcher
trainmatcheronexamples
sayred trainthen
whatmatchersays aboutred train
saytrain redthen
whatmatchersays abouttrain red
sayFirst scorethen
out of 100, how surematcheris aboutred train
sayReversed scorethen
out of 100, how surematcheris abouttrain red

Add two labelled prediction outputs and two score outputs after a single training step. Enter the phrases in opposite order.

What to look for

Both phrases produce vehicle and the same score.

Make it yours

Try uppercase letters and a question mark. The tidying step removes both, so the word associations stay the same.

Build with me · 5

Inspect the no-overlap fallback

Neither class contains marble. Both raw supports are zero, so ordinary division by total support would be undefined. This implementation explicitly replaces the empty evidence with equal shares across labels.

The returned label still needs to be some string, but the tie does not mean the model recognised the object. Distinguish a documented fallback from a useful prediction.

Blocks at this stageWorked example
make a dataset calledexamples
add toexamplesthe examplered applelabeledfruit
add toexamplesthe examplegreen applelabeledfruit
add toexamplesthe exampleblue trainlabeledvehicle
add toexamplesthe examplefast trainlabeledvehicle
make aword matchingmodel calledmatcher
trainmatcheronexamples
setmessagetomarble
sayInputthen
message
sayLabelthen
whatmatchersays about
message
sayScorethen
out of 100, how surematcheris about
message

Set the stored input to marble and leave the model unchanged.

What to look for

Score is 50 for two classes; the returned tie label is not evidence about marble.

Make it yours

Try several other absent words. Explain why a repeating label here does not show successful generalisation.

With three labels, the same rule would give each one an equal share of about 33. This even split is a rule written into this particular program. It is not a general method AI systems use to detect uncertainty.

Build with me · 6

Check the input that reaches the matcher

Lower-casing changes APPLE to apple, and splitting at punctuation drops the exclamation mark. Common words such as the and is are dropped when other words remain. This tidying happens before the model counts anything; people call it preprocessing.

The last sentence has extra words, but apple supplies its known association. Tidying makes some harmless differences disappear, which helps, while also throwing away differences that another task might need.

Blocks at this stageWorked example
make a dataset calledexamples
add toexamplesthe examplered applelabeledfruit
add toexamplesthe examplegreen applelabeledfruit
add toexamplesthe exampleblue trainlabeledvehicle
add toexamplesthe examplefast trainlabeledvehicle
make aword matchingmodel calledmatcher
trainmatcheronexamples
sayLower casethen
whatmatchersays aboutapple
sayUpper casethen
whatmatchersays aboutAPPLE!
sayMixed sentencethen
whatmatchersays aboutThe apple is here

Place the three prediction values in separately labelled output blocks. Read each literal phrase so you know what changed between probes.

What to look for

All three supplied inputs choose fruit.

Make it yours

Choose a pair of sentences whose meaning changes with order. Explain why this representation cannot reliably distinguish that pair.

Build with me · 7

Construct an input with the wrong meaning

A poster is neither a fruit nor a vehicle, but those are the only available labels. Apple and train each contribute two occurrences, creating a tie. The model has no third class for poster and no rule saying that a mention is different from the thing itself.

This is a problem with how the task was defined, as well as with what the model can see. If your application needs neither, you must plan its behaviour and evaluate it. Merely displaying a confidence number does not add the missing capability.

Blocks at this stageWorked example
make a dataset calledexamples
add toexamplesthe examplered applelabeledfruit
add toexamplesthe examplegreen applelabeledfruit
add toexamplesthe exampleblue trainlabeledvehicle
add toexamplesthe examplefast trainlabeledvehicle
make aword matchingmodel calledmatcher
trainmatcheronexamples
setmessagetoApple train poster
sayInputthen
message
sayLabelthen
whatmatchersays about
message
sayScorethen
out of 100, how surematcheris about
message

Run the poster phrase and compare its actual intended meaning with the available labels.

What to look for

The model chooses one of its two labels with score 50; neither is a correct description of the poster task.

Make it yours

Try red poster. It can score 100 for fruit because red is its only known word. Explain why a 60-point gate would accept a wrong answer.

Apply that explanation to your question model

Now the confident mistake from the last lesson makes sense. After the tidying step drops is, Gravity is interesting leaves two words: gravity and interesting. Gravity appeared once under question and never under statement. Interesting never appeared at all. So question receives all the support, and the score is 100, even though the sentence is a statement.

You could add paired examples about the same subject, such as "Does the parcel arrive today" and "The parcel arrives today". That weakens the link between a topic and a label. But the model still cannot see word order or question marks, which are often the clues that separate a question from a statement. "Is the parcel here" and "The parcel is here" look identical to it: once is and the are dropped, both are just parcel and here. More examples cannot bring back information the model never looks at. Telling those apart needs a model that reads word order, and choosing a model that can see the right clues is part of the job.

Model output and application behaviour

The model returns a label and a score. Everything else is the program you wrote: it reads the input, checks the score against a threshold, and prints a message.

That gives two different kinds of failure. If the program prints the label before checking the score, it can show a confident answer and then say it is unsure; that is a mistake in the program, fixed by moving blocks. If the gate is in the right place but the model gives a wrong answer with a high score, the program is following its rule correctly and the rule cannot catch that case; that is a limit of the model, fixed by better examples or a better model.

The last stage turns these one-off checks into a list you can run again after every change.

Build with me · 8

Keep a small repeatable probe list

A probe list makes it harder to remember only the successful example. Each turn prints the input, label, and score together. Write the intended label, or neither, beside each result.

If you change training, rerun this same list to compare versions. Do not call it an untouched final test after you have used its failures to guide improvements. It has become development evidence.

Blocks at this stageWorked example
make a dataset calledexamples
add toexamplesthe examplered applelabeledfruit
add toexamplesthe examplegreen applelabeledfruit
add toexamplesthe exampleblue trainlabeledvehicle
add toexamplesthe examplefast trainlabeledvehicle
make aword matchingmodel calledmatcher
trainmatcheronexamples
make an empty list calledprobes
addappletoprobes
addred traintoprobes
addtrain redtoprobes
addmarbletoprobes
addred postertoprobes
for eachmessagein listprobes
sayInputthen
message
sayLabelthen
whatmatchersays about
message
sayScorethen
out of 100, how surematcheris about
message

Build the list in the shown order and place all three output statements inside for each. Train once above the list and loop.

What to look for

Five grouped results appear. They include both successful familiar matches and limitations.

Make it yours

Add one paired-topic question and statement to your own question-model probe list, applying the same record-then-change method.

A useful list for your question model contains familiar words in new sentences, words it has never seen, questions without question marks, statements that use question words, and messages that are neither. For each failure, ask three things: did the model see the useful clue, did the training examples cover it, and did the program respond sensibly?

"The confidence was high" describes a result. "The matcher counted gravity under question and ignored word order" explains the mechanism, and a mechanism is something you can test and fix.

Full reference solution

This is the complete worked program. Try building it yourself first, then use this reference to find the first place your version behaves differently. The Python below is generated from these exact blocks; helper functions are included so its behaviour can be inspected.

make a dataset calledexamples
add toexamplesthe examplered applelabeledfruit
add toexamplesthe examplegreen applelabeledfruit
add toexamplesthe exampleblue trainlabeledvehicle
add toexamplesthe examplefast trainlabeledvehicle
make aword matchingmodel calledmatcher
trainmatcheronexamples
make an empty list calledprobes
addappletoprobes
addred traintoprobes
addtrain redtoprobes
addmarbletoprobes
addred postertoprobes
for eachmessagein listprobes
sayInputthen
message
sayLabelthen
whatmatchersays about
message
sayScorethen
out of 100, how surematcheris about
message
PythonHover over a line to see an explanation
# ---------------------------------------------------------------# Building blocks, written out in plain Python.# This part is generated for you. Your script starts further down.# ---------------------------------------------------------------  import mathimport random  class Dataset:    """A pile of labelled examples. Features can be a number, some text, or a    list of numbers. The label is whatever answer you want back."""     def __init__(self, name="dataset"):        self.name = name        self.rows = []        # Column names from the file's header row, when it came from one.        # Without these "the km of a row" has nothing to look the name up in.        self.columns = []     def add(self, features, label, fields=None):        self.rows.append(            {                "features": features,                "label": str(label),                "fields": dict(fields) if fields else {},            }        )     def size(self):        return len(self.rows)     def __len__(self):        return len(self.rows)     def __iter__(self):        """Walking a dataset gives you its rows, so "for each row in list        [testing]" reads the way it sounds."""        return iter(self.rows)     def labels(self):        seen = []        for row in self.rows:            if row["label"] not in seen:                seen.append(row["label"])        return seen     def most_common_label(self):        if not self.rows:            return ""        counts = {}        for row in self.rows:            counts[row["label"]] = counts.get(row["label"], 0) + 1        return max(counts, key=lambda label: counts[label])     def split(self, train_percent=80):        """Keeps the given percent for training and hands back the rest as a        test set. The shuffle is seeded, so you get the same split every run."""        order = list(range(len(self.rows)))        random.Random(0).shuffle(order)        cut = int(len(order) * train_percent / 100)        train = Dataset(self.name + " (train)")        test = Dataset(self.name + " (test)")        train.columns = list(self.columns)        test.columns = list(self.columns)        for position, index in enumerate(order):            row = self.rows[index]            target = train if position < cut else test            target.add(row["features"], row["label"], row.get("fields"))        return train, test     def show(self, limit=10):        print(self.name + ": " + str(len(self.rows)) + " examples")        for row in self.rows[:limit]:            print("  " + str(row["features"]) + "  ->  " + row["label"])        if len(self.rows) > limit:            print("  ... and " + str(len(self.rows) - limit) + " more")  def new_dataset(name="dataset"):    return Dataset(name)  def row_field(row, name):    """One named piece of a row: "the km of this row", "the label of it".     Names come from the header line of the csv. "label" always works, even on    a file with no header, because every row has one."""    wanted = str(name).strip()    if not isinstance(row, dict):        print("That is not a row. Use this inside a for each over a dataset.")        return ""    if wanted.lower() == "label":        return row.get("label", "")    fields = row.get("fields") or {}    if wanted in fields:        return fields[wanted]    # Header names are matched loosely, so "Rain" finds the "rain" column.    for key in fields:        if str(key).strip().lower() == wanted.lower():            return fields[key]    known = ", ".join([str(k) for k in fields]) if fields else "none"    print(        "No column called "        + wanted        + " in this row. Columns here: "        + known        + "."    )    return ""  def load_csv(dataset, path, label_column=-1, has_header=True):    """Reads a comma separated file into a dataset.     Everything except the label column becomes the features, and anything that    looks like a number is turned into one. This is how a file you uploaded    becomes something you can train on."""    try:        with open(path) as handle:            rows = [line.rstrip("\n").rstrip("\r") for line in handle]    except OSError:        print("Could not find " + path + ". Check the name in the Files panel.")        return dataset     rows = [row for row in rows if row.strip()]     # A file with no label column is a perfectly normal thing to load: it is    # what a test set looks like before you have predicted anything.    labelled = str(label_column).strip().lower() not in ("none", "", "no", "-")     header = []    if has_header and rows:        header = [cell.strip() for cell in rows[0].split(",")]        rows = rows[1:]     added = 0    for row in rows:        cells = [cell.strip() for cell in row.split(",")]        if not cells or (labelled and len(cells) < 2):            continue         index = -1        if labelled:            index = int(label_column)            if index < 0:                index = len(cells) + index            if index < 0 or index >= len(cells):                continue         label = cells[index] if labelled else ""        keep = [i for i in range(len(cells)) if i != index]         typed = []        for i in keep:            try:                typed.append(float(cells[i]))            except ValueError:                typed.append(cells[i])         # Every kept column gets its header name, so "the km of a row" works.        fields = {}        for position, i in enumerate(keep):            if i < len(header) and header[i]:                fields[header[i]] = typed[position]        if header and not dataset.columns:            dataset.columns = [header[i] for i in keep if i < len(header)]         dataset.add(typed[0] if len(typed) == 1 else typed, label, fields)        added += 1     print("Loaded " + str(added) + " rows from " + path + ".")    return dataset  def save_submission(predictions, dataset, path="submission.csv"):    """Prints your predictions in the exact two column format a round is    scored in: a header, then one id and one label per line.     It is printed rather than saved to a file because the console is the one    place you can copy it from. Paste it into a new file in the Files panel,    or straight into the upload box."""    labels = list(predictions or [])    rows = list(dataset) if dataset is not None else []     if len(labels) != len(rows):        print(            "You have "            + str(len(labels))            + " predictions for "            + str(len(rows))            + " rows. Those have to match before this means anything."        )        return     lines = ["id,label"]    for position, row in enumerate(rows):        # Looked up directly rather than through row_field, which would        # complain on every row of a file that simply has no id column.        fields = row.get("fields") or {}        row_id = ""        for key in fields:            if str(key).strip().lower() == "id":                row_id = fields[key]                break        if not str(row_id).strip():            row_id = "r" + str(position + 1).zfill(2)        if isinstance(row_id, float) and row_id == int(row_id):            row_id = int(row_id)        lines.append(str(row_id) + "," + str(labels[position]).strip())     print("--- submission.csv, copy from here ---")    for line in lines:        print(line)    print("--- to here, " + str(len(lines) - 1) + " rows ---")  def _words_in(value):    letters = []    for character in str(value).lower():        letters.append(character if character.isalnum() else " ")    return [word for word in "".join(letters).split() if word]  # Words that turn up in every kind of sentence carry no signal about the# label, so the word matching model looks past them._EVERYDAY_WORDS = set(    "a an and are as at be been but by can could did do for from had has have "    "he her his i if in is it its me my not of on or our she so than that the "    "their them then there they this to too us was we were what when which "    "who will with would you your".split())  def _content_words(value):    words = _words_in(value)    kept = [word for word in words if word not in _EVERYDAY_WORDS]    return kept if kept else words  def _as_numbers(value):    if isinstance(value, (list, tuple)):        return [float(item) for item in value]    return [float(value)]  def _is_numeric(value):    try:        _as_numbers(value)        return True    except (TypeError, ValueError):        return False  def _distance(left, right):    """How far apart two examples are. Numbers use straight line distance,    text uses how many words the two do not share."""    if _is_numeric(left) and _is_numeric(right):        a = _as_numbers(left)        b = _as_numbers(right)        while len(a) < len(b):            a.append(0.0)        while len(b) < len(a):            b.append(0.0)        total = 0.0        for i in range(len(a)):            total += (a[i] - b[i]) ** 2        return math.sqrt(total)    a = set(_words_in(left))    b = set(_words_in(right))    if not a and not b:        return 0.0    shared = len(a & b)    return 1.0 - (shared / float(len(a | b)))  class Model:    """Three small classifiers behind one name.     nearest  looks for the closest example it was trained on    words    scores how many words the input shares with each label    common   always answers with the most common label, the baseline to beat    """     def __init__(self, kind="nearest", name="model"):        self.kind = kind        self.name = name        self.data = None        self.word_scores = {}     def train(self, dataset):        self.data = dataset        self.word_scores = {}        if self.kind == "words":            for row in dataset.rows:                bucket = self.word_scores.setdefault(row["label"], {})                for word in _content_words(row["features"]):                    bucket[word] = bucket.get(word, 0) + 1        print("Trained " + self.name + " on " + str(dataset.size()) + " examples.")     def _scores(self, features):        if self.data is None or not self.data.rows:            return {}        if self.kind == "common":            counts = {}            for row in self.data.rows:                counts[row["label"]] = counts.get(row["label"], 0) + 1            return counts        if self.kind == "words":            # Score each label by how much of its training vocabulary shows up            # in the input, divided by how much text that label was trained on            # so a label with more examples cannot win on volume alone.            labels = self.data.labels()            scores = {}            for label in labels:                bucket = self.word_scores.get(label, {})                seen = sum(bucket.values()) or 1                running = 0.0                for word in _content_words(features):                    running += bucket.get(word, 0) / float(seen)                scores[label] = round(running, 6)            if sum(scores.values()) == 0:                # Nothing in the input was ever seen in training. Say so by                # splitting the vote evenly, which reads as low confidence.                return dict((label, 1) for label in labels)            return scores        ranked = sorted(self.data.rows, key=lambda row: _distance(features, row["features"]))        neighbours = ranked[: min(3, len(ranked))]        scores = {}        for row in neighbours:            scores[row["label"]] = scores.get(row["label"], 0) + 1        return scores     def predict(self, features):        scores = self._scores(features)        if not scores:            return ""        return max(scores, key=lambda label: scores[label])     def confidence(self, features):        """How much of the vote the winning label took, out of 100."""        scores = self._scores(features)        total = sum(scores.values())        if not scores or total == 0:            return 0.0        best = max(scores.values())        return round(100.0 * best / float(total), 1)     def accuracy(self, dataset):        if not dataset.rows:            return 0.0        right = 0        for row in dataset.rows:            if self.predict(row["features"]) == row["label"]:                right += 1        return round(100.0 * right / float(len(dataset.rows)), 1)     def show(self):        print(self.name + " is a " + self.kind + " model.")        if self.data is None:            print("  It has not been trained yet.")            return        print("  Trained on " + str(self.data.size()) + " examples.")        print("  Labels it can answer with: " + ", ".join(self.data.labels()))  def new_model(kind="nearest", name="model"):    return Model(kind, name)  # ---------------------------------------------------------------# Your script# ---------------------------------------------------------------  examples = new_dataset("examples")examples.add("red apple", "fruit")examples.add("green apple", "fruit")examples.add("blue train", "vehicle")examples.add("fast train", "vehicle")matcher = new_model("words", "matcher")matcher.train(examples)probes = []probes.append("apple")probes.append("red train")probes.append("train red")probes.append("marble")probes.append("red poster")for message in probes:    print("Input", message)    print("Label", matcher.predict(message))    print("Score", matcher.confidence(message))

Compare this with your version. Different names and personal choices are fine when the program follows the same logic.

Keep your progress

Sign in and every reading, quiz, and exercise you finish is saved.

Sign in