What the model did with your words
Inspect the exact limits of the course word matcher and understand why fluent-looking labels can still be wrong.
Open the black box a little
The word-matching model you trained in the last lesson is small enough to explain completely. It does three things.
First, it tidies each message. It makes every letter lower case, splits the text into words at spaces and punctuation, and drops very common words such as the, is, and can. If dropping them would leave nothing, it keeps all the words.
Second, during training, it counts how often each remaining word appears under each label.
Third, during prediction, it looks up the words of the new message. Every word that appeared under a label adds support for that label. Each label's support is divided by the total number of words that label was trained on, so a label cannot win just because its examples were longer. The label with the most support wins, and the score says what share of all the support it received.
That is the whole model. It does not look at word order. It does not treat a question mark as a clue. It does not check which word comes first. If a lesson claimed this model had learned "questions start with is", it would be describing a different model.
Work through a tiny collection
Sixteen examples are too many to count by hand, so the stages below use a separate four-row dataset. Keep this table in view while you work through them:
| Text | Label |
|---|---|
| red apple | fruit |
| green apple | fruit |
| blue train | vehicle |
| fast train | vehicle |
Before each stage, work out the answer on paper, then run the program to check.
Build with me · 1
Make every training association visible
Read every word under fruit: red, apple, green, apple. There are four occurrences in total. Read vehicle: blue, train, fast, train. It also has four. The repeated words apple and train each contribute two occurrences in their class.
The model lower-cases the text and counts words. It is not storing a dictionary definition of fruit or a photograph of a train. This tiny collection lets us predict which associations the implementation can use.
Build the four add-example blocks, then create and train matcher. Keep the dataset inspection at the bottom.
What to look for
Four labelled rows are printed, and training uses four examples.
Make it yours
Replace red with yellow in exactly one row. Which prediction input would directly test that edit?
Build with me · 2
Calculate the support for one word
For apple, fruit receives two matching occurrences divided by its four total words: 2/4. Vehicle receives zero. The displayed confidence divides the winning support by all support, so 0.5 divided by 0.5 is 100%.
That number describes this counting rule on this input. It is not a promise that the real meaning is fruit.
Store apple under message. Read that same variable in the label and confidence blocks.
What to look for
Label is fruit and Score is 100.
Make it yours
Try train and work out the analogous calculation before running.
Build with me · 3
Combine evidence for both labels
Red occurs once under fruit, giving 1/4 support. Train occurs twice under vehicle, giving 2/4 support. Vehicle receives two thirds of the total support, about 66.7%.
Nothing in these counts checks whether red modifies train. Words contribute individually. The winning label therefore follows the associations, not a grammatical understanding of the phrase.
Change only the stored input to red train. Keep the four-row training stack unchanged.
What to look for
Vehicle wins, with about 66.7 out of 100.
Make it yours
Try apple train. Each side now has two matching occurrences; explain why the score becomes 50.
Build with me · 4
Show what word order cannot change
Swapping the order preserves the words and their counts. Because the model only counts words and ignores their order, both inputs receive the same label and score.
A model cannot use information it throws away. Adding a thousand copies of these same rows would not teach this particular matcher to distinguish two inputs represented by the same counts.
Add two labelled prediction outputs and two score outputs after a single training step. Enter the phrases in opposite order.
What to look for
Both phrases produce vehicle and the same score.
Make it yours
Try uppercase letters and a question mark. The tidying step removes both, so the word associations stay the same.
Build with me · 5
Inspect the no-overlap fallback
Neither class contains marble. Both raw supports are zero, so ordinary division by total support would be undefined. This implementation explicitly replaces the empty evidence with equal shares across labels.
The returned label still needs to be some string, but the tie does not mean the model recognised the object. Distinguish a documented fallback from a useful prediction.
Set the stored input to marble and leave the model unchanged.
What to look for
Score is 50 for two classes; the returned tie label is not evidence about marble.
Make it yours
Try several other absent words. Explain why a repeating label here does not show successful generalisation.
With three labels, the same rule would give each one an equal share of about 33. This even split is a rule written into this particular program. It is not a general method AI systems use to detect uncertainty.
Build with me · 6
Check the input that reaches the matcher
Lower-casing changes APPLE to apple, and splitting at punctuation drops the exclamation mark. Common words such as the and is are dropped when other words remain. This tidying happens before the model counts anything; people call it preprocessing.
The last sentence has extra words, but apple supplies its known association. Tidying makes some harmless differences disappear, which helps, while also throwing away differences that another task might need.
Place the three prediction values in separately labelled output blocks. Read each literal phrase so you know what changed between probes.
What to look for
All three supplied inputs choose fruit.
Make it yours
Choose a pair of sentences whose meaning changes with order. Explain why this representation cannot reliably distinguish that pair.
Build with me · 7
Construct an input with the wrong meaning
A poster is neither a fruit nor a vehicle, but those are the only available labels. Apple and train each contribute two occurrences, creating a tie. The model has no third class for poster and no rule saying that a mention is different from the thing itself.
This is a problem with how the task was defined, as well as with what the model can see. If your application needs neither, you must plan its behaviour and evaluate it. Merely displaying a confidence number does not add the missing capability.
Run the poster phrase and compare its actual intended meaning with the available labels.
What to look for
The model chooses one of its two labels with score 50; neither is a correct description of the poster task.
Make it yours
Try red poster. It can score 100 for fruit because red is its only known word. Explain why a 60-point gate would accept a wrong answer.
Apply that explanation to your question model
Now the confident mistake from the last lesson makes sense. After the tidying step drops is, Gravity is interesting leaves two words: gravity and interesting. Gravity appeared once under question and never under statement. Interesting never appeared at all. So question receives all the support, and the score is 100, even though the sentence is a statement.
You could add paired examples about the same subject, such as "Does the parcel arrive today" and "The parcel arrives today". That weakens the link between a topic and a label. But the model still cannot see word order or question marks, which are often the clues that separate a question from a statement. "Is the parcel here" and "The parcel is here" look identical to it: once is and the are dropped, both are just parcel and here. More examples cannot bring back information the model never looks at. Telling those apart needs a model that reads word order, and choosing a model that can see the right clues is part of the job.
Model output and application behaviour
The model returns a label and a score. Everything else is the program you wrote: it reads the input, checks the score against a threshold, and prints a message.
That gives two different kinds of failure. If the program prints the label before checking the score, it can show a confident answer and then say it is unsure; that is a mistake in the program, fixed by moving blocks. If the gate is in the right place but the model gives a wrong answer with a high score, the program is following its rule correctly and the rule cannot catch that case; that is a limit of the model, fixed by better examples or a better model.
The last stage turns these one-off checks into a list you can run again after every change.
Build with me · 8
Keep a small repeatable probe list
A probe list makes it harder to remember only the successful example. Each turn prints the input, label, and score together. Write the intended label, or neither, beside each result.
If you change training, rerun this same list to compare versions. Do not call it an untouched final test after you have used its failures to guide improvements. It has become development evidence.
Build the list in the shown order and place all three output statements inside for each. Train once above the list and loop.
What to look for
Five grouped results appear. They include both successful familiar matches and limitations.
Make it yours
Add one paired-topic question and statement to your own question-model probe list, applying the same record-then-change method.
A useful list for your question model contains familiar words in new sentences, words it has never seen, questions without question marks, statements that use question words, and messages that are neither. For each failure, ask three things: did the model see the useful clue, did the training examples cover it, and did the program respond sensibly?
"The confidence was high" describes a result. "The matcher counted gravity under question and ignored word order" explains the mechanism, and a mechanism is something you can test and fix.
Full reference solution
This is the complete worked program. Try building it yourself first, then use this reference to find the first place your version behaves differently. The Python below is generated from these exact blocks; helper functions are included so its behaviour can be inspected.
# ---------------------------------------------------------------# Building blocks, written out in plain Python.# This part is generated for you. Your script starts further down.# --------------------------------------------------------------- import mathimport random class Dataset: """A pile of labelled examples. Features can be a number, some text, or a list of numbers. The label is whatever answer you want back.""" def __init__(self, name="dataset"): self.name = name self.rows = [] # Column names from the file's header row, when it came from one. # Without these "the km of a row" has nothing to look the name up in. self.columns = [] def add(self, features, label, fields=None): self.rows.append( { "features": features, "label": str(label), "fields": dict(fields) if fields else {}, } ) def size(self): return len(self.rows) def __len__(self): return len(self.rows) def __iter__(self): """Walking a dataset gives you its rows, so "for each row in list [testing]" reads the way it sounds.""" return iter(self.rows) def labels(self): seen = [] for row in self.rows: if row["label"] not in seen: seen.append(row["label"]) return seen def most_common_label(self): if not self.rows: return "" counts = {} for row in self.rows: counts[row["label"]] = counts.get(row["label"], 0) + 1 return max(counts, key=lambda label: counts[label]) def split(self, train_percent=80): """Keeps the given percent for training and hands back the rest as a test set. The shuffle is seeded, so you get the same split every run.""" order = list(range(len(self.rows))) random.Random(0).shuffle(order) cut = int(len(order) * train_percent / 100) train = Dataset(self.name + " (train)") test = Dataset(self.name + " (test)") train.columns = list(self.columns) test.columns = list(self.columns) for position, index in enumerate(order): row = self.rows[index] target = train if position < cut else test target.add(row["features"], row["label"], row.get("fields")) return train, test def show(self, limit=10): print(self.name + ": " + str(len(self.rows)) + " examples") for row in self.rows[:limit]: print(" " + str(row["features"]) + " -> " + row["label"]) if len(self.rows) > limit: print(" ... and " + str(len(self.rows) - limit) + " more") def new_dataset(name="dataset"): return Dataset(name) def row_field(row, name): """One named piece of a row: "the km of this row", "the label of it". Names come from the header line of the csv. "label" always works, even on a file with no header, because every row has one.""" wanted = str(name).strip() if not isinstance(row, dict): print("That is not a row. Use this inside a for each over a dataset.") return "" if wanted.lower() == "label": return row.get("label", "") fields = row.get("fields") or {} if wanted in fields: return fields[wanted] # Header names are matched loosely, so "Rain" finds the "rain" column. for key in fields: if str(key).strip().lower() == wanted.lower(): return fields[key] known = ", ".join([str(k) for k in fields]) if fields else "none" print( "No column called " + wanted + " in this row. Columns here: " + known + "." ) return "" def load_csv(dataset, path, label_column=-1, has_header=True): """Reads a comma separated file into a dataset. Everything except the label column becomes the features, and anything that looks like a number is turned into one. This is how a file you uploaded becomes something you can train on.""" try: with open(path) as handle: rows = [line.rstrip("\n").rstrip("\r") for line in handle] except OSError: print("Could not find " + path + ". Check the name in the Files panel.") return dataset rows = [row for row in rows if row.strip()] # A file with no label column is a perfectly normal thing to load: it is # what a test set looks like before you have predicted anything. labelled = str(label_column).strip().lower() not in ("none", "", "no", "-") header = [] if has_header and rows: header = [cell.strip() for cell in rows[0].split(",")] rows = rows[1:] added = 0 for row in rows: cells = [cell.strip() for cell in row.split(",")] if not cells or (labelled and len(cells) < 2): continue index = -1 if labelled: index = int(label_column) if index < 0: index = len(cells) + index if index < 0 or index >= len(cells): continue label = cells[index] if labelled else "" keep = [i for i in range(len(cells)) if i != index] typed = [] for i in keep: try: typed.append(float(cells[i])) except ValueError: typed.append(cells[i]) # Every kept column gets its header name, so "the km of a row" works. fields = {} for position, i in enumerate(keep): if i < len(header) and header[i]: fields[header[i]] = typed[position] if header and not dataset.columns: dataset.columns = [header[i] for i in keep if i < len(header)] dataset.add(typed[0] if len(typed) == 1 else typed, label, fields) added += 1 print("Loaded " + str(added) + " rows from " + path + ".") return dataset def save_submission(predictions, dataset, path="submission.csv"): """Prints your predictions in the exact two column format a round is scored in: a header, then one id and one label per line. It is printed rather than saved to a file because the console is the one place you can copy it from. Paste it into a new file in the Files panel, or straight into the upload box.""" labels = list(predictions or []) rows = list(dataset) if dataset is not None else [] if len(labels) != len(rows): print( "You have " + str(len(labels)) + " predictions for " + str(len(rows)) + " rows. Those have to match before this means anything." ) return lines = ["id,label"] for position, row in enumerate(rows): # Looked up directly rather than through row_field, which would # complain on every row of a file that simply has no id column. fields = row.get("fields") or {} row_id = "" for key in fields: if str(key).strip().lower() == "id": row_id = fields[key] break if not str(row_id).strip(): row_id = "r" + str(position + 1).zfill(2) if isinstance(row_id, float) and row_id == int(row_id): row_id = int(row_id) lines.append(str(row_id) + "," + str(labels[position]).strip()) print("--- submission.csv, copy from here ---") for line in lines: print(line) print("--- to here, " + str(len(lines) - 1) + " rows ---") def _words_in(value): letters = [] for character in str(value).lower(): letters.append(character if character.isalnum() else " ") return [word for word in "".join(letters).split() if word] # Words that turn up in every kind of sentence carry no signal about the# label, so the word matching model looks past them._EVERYDAY_WORDS = set( "a an and are as at be been but by can could did do for from had has have " "he her his i if in is it its me my not of on or our she so than that the " "their them then there they this to too us was we were what when which " "who will with would you your".split()) def _content_words(value): words = _words_in(value) kept = [word for word in words if word not in _EVERYDAY_WORDS] return kept if kept else words def _as_numbers(value): if isinstance(value, (list, tuple)): return [float(item) for item in value] return [float(value)] def _is_numeric(value): try: _as_numbers(value) return True except (TypeError, ValueError): return False def _distance(left, right): """How far apart two examples are. Numbers use straight line distance, text uses how many words the two do not share.""" if _is_numeric(left) and _is_numeric(right): a = _as_numbers(left) b = _as_numbers(right) while len(a) < len(b): a.append(0.0) while len(b) < len(a): b.append(0.0) total = 0.0 for i in range(len(a)): total += (a[i] - b[i]) ** 2 return math.sqrt(total) a = set(_words_in(left)) b = set(_words_in(right)) if not a and not b: return 0.0 shared = len(a & b) return 1.0 - (shared / float(len(a | b))) class Model: """Three small classifiers behind one name. nearest looks for the closest example it was trained on words scores how many words the input shares with each label common always answers with the most common label, the baseline to beat """ def __init__(self, kind="nearest", name="model"): self.kind = kind self.name = name self.data = None self.word_scores = {} def train(self, dataset): self.data = dataset self.word_scores = {} if self.kind == "words": for row in dataset.rows: bucket = self.word_scores.setdefault(row["label"], {}) for word in _content_words(row["features"]): bucket[word] = bucket.get(word, 0) + 1 print("Trained " + self.name + " on " + str(dataset.size()) + " examples.") def _scores(self, features): if self.data is None or not self.data.rows: return {} if self.kind == "common": counts = {} for row in self.data.rows: counts[row["label"]] = counts.get(row["label"], 0) + 1 return counts if self.kind == "words": # Score each label by how much of its training vocabulary shows up # in the input, divided by how much text that label was trained on # so a label with more examples cannot win on volume alone. labels = self.data.labels() scores = {} for label in labels: bucket = self.word_scores.get(label, {}) seen = sum(bucket.values()) or 1 running = 0.0 for word in _content_words(features): running += bucket.get(word, 0) / float(seen) scores[label] = round(running, 6) if sum(scores.values()) == 0: # Nothing in the input was ever seen in training. Say so by # splitting the vote evenly, which reads as low confidence. return dict((label, 1) for label in labels) return scores ranked = sorted(self.data.rows, key=lambda row: _distance(features, row["features"])) neighbours = ranked[: min(3, len(ranked))] scores = {} for row in neighbours: scores[row["label"]] = scores.get(row["label"], 0) + 1 return scores def predict(self, features): scores = self._scores(features) if not scores: return "" return max(scores, key=lambda label: scores[label]) def confidence(self, features): """How much of the vote the winning label took, out of 100.""" scores = self._scores(features) total = sum(scores.values()) if not scores or total == 0: return 0.0 best = max(scores.values()) return round(100.0 * best / float(total), 1) def accuracy(self, dataset): if not dataset.rows: return 0.0 right = 0 for row in dataset.rows: if self.predict(row["features"]) == row["label"]: right += 1 return round(100.0 * right / float(len(dataset.rows)), 1) def show(self): print(self.name + " is a " + self.kind + " model.") if self.data is None: print(" It has not been trained yet.") return print(" Trained on " + str(self.data.size()) + " examples.") print(" Labels it can answer with: " + ", ".join(self.data.labels())) def new_model(kind="nearest", name="model"): return Model(kind, name) # ---------------------------------------------------------------# Your script# --------------------------------------------------------------- examples = new_dataset("examples")examples.add("red apple", "fruit")examples.add("green apple", "fruit")examples.add("blue train", "vehicle")examples.add("fast train", "vehicle")matcher = new_model("words", "matcher")matcher.train(examples)probes = []probes.append("apple")probes.append("red train")probes.append("train red")probes.append("marble")probes.append("red poster")for message in probes: print("Input", message) print("Label", matcher.predict(message)) print("Score", matcher.confidence(message))Compare this with your version. Different names and personal choices are fine when the program follows the same logic.
Keep your progress
Sign in and every reading, quiz, and exercise you finish is saved.