Is it a question? Train a text model and gate it
Create a dataset, create a text model, train it, predict a label, and add a branch for when it is unsure.
The program you will build
The program reads a message, predicts whether it is a question or a statement, and then either replies or says it is unsure. You will build it in ten small stages beside this article. Each stage is a complete program: Try this stage loads everything it needs, and training happens in your browser when you press Run.
The steps are the ones you followed in Level 2: collect labelled examples, create a model, train it, ask it about a new input, then decide what to do with its answer. First we get the program working. Then we find where it fails. This model matches words; it does not understand English grammar.
Collect labelled examples
A model learns from examples, so the program starts with somewhere to keep them.
Build with me · 1
Create the collection before filling it
The dataset named examples begins empty. It will hold pairs: some input text and the label we want that text to have. Creating a dataset is bookkeeping; it does not create a model or teach one anything.
Decide the meaning of your labels before collecting rows. For this teaching task, question includes a request for an answer, even without a question mark. Statement means it states something. The model will count word associations, not apply this definition as a grammar rule.
From Data, add make a dataset called examples. Put its size value inside a labelled output below it.
What to look for
Examples is 0.
Make it yours
Rename the dataset in its make-dataset block and in every menu that uses it. The name can change without changing what a question means.
Build with me · 2
Add one input with its answer
An example has two parts. Please explain gravity is the input. question is the target answer supplied by a person. The label is not another input word that the model is allowed to read during prediction.
One example is enough to inspect the mechanics, but not enough to build a useful recogniser. At this stage we only want proof that the add block ran on the intended dataset.
Connect add to examples below the make-dataset block. Type the phrase in features and question in label. Connect the size and show blocks underneath.
What to look for
Examples is 1, and the shown row pairs the phrase with question.
Make it yours
Change the phrase to another request, keeping the label spelling exactly question. Explain why Question would be a different label.
Build with me · 3
Add a contrasting label
A classifier must have alternatives to compare. A second label gives it two possible answers, but the topics are also different: gravity appears only in a request and parcel only in a statement. That accidental connection will matter later.
Notice that both examples are added to the same dataset. Creating a new dataset before every add would erase the previous collection, leaving only the last row when training begins.
Add the second example underneath the first, selecting the existing examples dataset. Keep exactly one make-dataset statement at the top.
What to look for
Examples is 2 and both labels appear in the printed rows.
Make it yours
Write a question about a parcel and a statement about gravity on paper. Keep them for later probes rather than adding them immediately.
Here is the full teaching collection. It follows the rule from the first stage, so a request such as "Please explain gravity" counts as a question even without a question mark. Spell each label exactly the same way every time.
| Example text | Label |
|---|---|
| Please explain gravity | question |
| Could you explain magnets | question |
| Why does ice melt | question |
| How does a battery work | question |
| Where can I find the library | question |
| Can you describe a volcano | question |
| What causes thunder | question |
| How do seeds grow | question |
| The parcel arrived safely | statement |
| My parcel arrived yesterday | statement |
| Dinner tastes delicious | statement |
| Our team won today | statement |
| The blue door is closed | statement |
| Music played softly | statement |
| The dog slept peacefully | statement |
| We painted the fence | statement |
Build with me · 4
Grow the collection deliberately
Here are sixteen labelled teaching sentences. Eight request answers and eight state something. Repetition across words such as explain or parcel makes an association the toy model can count. Variety helps reveal how it uses words, but this is still a tiny invented dataset.
The size checkpoint comes after all add statements. The show operation may display only a preview, so count and preview answer different questions: how many rows exist, and what some rows actually contain.
Build or load the sixteen-row stack. Read the label on each add block. Run and check the final count before inserting any model block.
What to look for
Examples is 16. The preview shows input/label pairs, not a learned model.
Make it yours
Replace one sentence with your own under the same definition. Keep the class counts balanced during this first controlled comparison.
Create a model, then train it
The examples are ready. Now the program needs a model to learn from them.
Build with me · 5
Choose a model and give it a name
The make-model block creates judge and selects word matching as its algorithm. An algorithm is the procedure the model uses to learn and predict. This one compares word associations. Choosing it does not automatically connect it to examples.
The dataset and model have different names because they have different jobs. examples stores labelled rows; judge will store the learned information. You can inspect judge before training to see that creating a model is not training it.
From Model, choose make a word matching model called judge. Put it below the examples. Add show what judge learned after it.
What to look for
The model report identifies a words model with zero training examples. No prediction is made yet.
Make it yours
Locate every occurrence of examples and judge in the stack. Explain what each menu is asking you to choose.
Build with me · 6
Connect the model to the collection
Train judge on examples is the instruction that reads the labelled collection and records its word patterns in the model. It must run after collection and after model creation.
Putting another make judge block below training would replace the trained model with an empty one. The same thing happens in Python later: creating a new model under the same name throws away the trained one.
Insert train judge on examples between the make-model block and its show block. Keep the model and dataset menus in the correct positions.
What to look for
The training message reports 16 examples, and the model report now says it was trained on 16 examples.
Make it yours
Temporarily move a fresh make-model block after training and inspect the reset model, then Undo. Do not diagnose its missing learning by adding more examples.
The order matters: make the dataset, add the examples, make the model, train the model. A second make-dataset block after the examples would empty the collection again, in the same way that a second make-model block after training empties the model.
Use the trained model
A trained model can now be given a message it has never seen. Asking it for a label is prediction, which Level 2 also called inference. The model does not learn anything from the new message; it only answers.
Build with me · 7
Give the trained model a new message
Prediction uses the trained model. It does not add the new message to training, and it does not require that you provide the answer. Store the input under message so the printed input and predicted label refer to the same text.
This familiar wording is a wiring check: if the program reads the wrong value, fix that first. A successful familiar case is not evidence that the model understands new English sentences.
Nest the rounded message variable in what judge says about. Put that prediction value inside a labelled say, after training.
What to look for
The supplied message is printed and the model returns question.
Make it yours
Replace the message with The parcel arrived safely and predict the output. Then try new wording on the same topic.
Build with me · 8
Keep the label and its score together
The model supplies two things: a winning label, and a score out of 100 saying how much of its evidence pointed to that label. We store each because the next decision will need them. Both blocks read the same message and the same model.
These unfamiliar words have no useful overlap with the training vocabulary. This implementation shares support evenly when no words match, giving 50 for two classes. A tie still returns a label; the application must decide whether to present it as a useful answer.
Add assignments for label and sure. Place prediction and confidence value blocks in their sockets, both reading message. Print the stored values underneath.
What to look for
Score is 50 for this no-overlap input. Treat the tie's label as an arbitrary result of the implementation, not a recognition success.
Make it yours
Change message to a known phrase. Compare the score, then explain why a larger number alone cannot prove correctness.
A confidence gate
The model always returns a label, even for input unlike anything it trained on, because question and statement are the only answers it has. The program around the model decides whether that label deserves to be shown. A rule that checks the score before showing the label is called a confidence gate.
Build with me · 9
Decide how the application handles uncertainty
The gate is an application rule you write. If sure is below 60, it prints an unsure message. Otherwise, a second if chooses the question or statement response. The verdict messages must live inside the appropriate branches.
The threshold is a trial choice, not a proven safe boundary. A score of exactly 60 passes because the condition is less than 60. Raising the threshold means more unsure answers and fewer accepted ones; it cannot catch every confidently wrong answer.
Build the outer gate and inner label check. Press Run, answer the waiting Terminal with an unfamiliar message, and keep the score output outside the gate for diagnosis.
What to look for
After answering marble lantern violin, the program prints the unsure response and a score of 50.
Make it yours
Try thresholds 50, 60, and 80 with the same Terminal answer. For each, count the wrong answers it accepted and the times it said unsure.
The gate catches messages the model has no evidence about. It cannot catch a message the model is sure about and gets wrong.
Build with me · 10
Expose a confident mistake
Gravity is interesting is a statement. Yet gravity appeared in a question example, and this matcher ignores the sentence structure that would help distinguish the two. It can confidently choose question.
Keep the raw label and score visible during development so a reassuring response cannot hide how the decision arose. A program can run every block correctly and still make a poor prediction, because the model only counts words and learned from sixteen examples.
Press Run and enter Gravity is interesting when Terminal waits for an answer. Do not change the training collection or gate. Compare the intended label with the actual result before making improvements.
What to look for
The example demonstrates a confident question prediction for a statement. The 60-point gate does not prevent this error.
Make it yours
Try your saved parcel question too. Add paired topics only after recording this first version, retrain, and check whether the change helps the held-out wording.
This is a limit of the model and its examples, not a broken if block. The next lesson opens the model up and shows exactly why it happens.
Test it fairly
The messages you just tried were picked to show how the program works, after you had seen the training examples. They are not a fair test. To measure the model, write a separate development set, as you did in Level 2: new wording, with the same topic written both as a question and as a statement. For each message, record the true label, the predicted label, the score, and whether the gate said unsure. Count the unsure answers too; leaving them out hides how often the program refuses to answer.
Changing the training examples needs another Run, because every Run starts at the top of the program and rebuilds the dataset and the model.
Full reference solution
This is the complete worked program. Try building it yourself first, then use this reference to find the first place your version behaves differently. The Python below is generated from these exact blocks; helper functions are included so its behaviour can be inspected.
# ---------------------------------------------------------------# Building blocks, written out in plain Python.# This part is generated for you. Your script starts further down.# --------------------------------------------------------------- import mathimport random class Dataset: """A pile of labelled examples. Features can be a number, some text, or a list of numbers. The label is whatever answer you want back.""" def __init__(self, name="dataset"): self.name = name self.rows = [] # Column names from the file's header row, when it came from one. # Without these "the km of a row" has nothing to look the name up in. self.columns = [] def add(self, features, label, fields=None): self.rows.append( { "features": features, "label": str(label), "fields": dict(fields) if fields else {}, } ) def size(self): return len(self.rows) def __len__(self): return len(self.rows) def __iter__(self): """Walking a dataset gives you its rows, so "for each row in list [testing]" reads the way it sounds.""" return iter(self.rows) def labels(self): seen = [] for row in self.rows: if row["label"] not in seen: seen.append(row["label"]) return seen def most_common_label(self): if not self.rows: return "" counts = {} for row in self.rows: counts[row["label"]] = counts.get(row["label"], 0) + 1 return max(counts, key=lambda label: counts[label]) def split(self, train_percent=80): """Keeps the given percent for training and hands back the rest as a test set. The shuffle is seeded, so you get the same split every run.""" order = list(range(len(self.rows))) random.Random(0).shuffle(order) cut = int(len(order) * train_percent / 100) train = Dataset(self.name + " (train)") test = Dataset(self.name + " (test)") train.columns = list(self.columns) test.columns = list(self.columns) for position, index in enumerate(order): row = self.rows[index] target = train if position < cut else test target.add(row["features"], row["label"], row.get("fields")) return train, test def show(self, limit=10): print(self.name + ": " + str(len(self.rows)) + " examples") for row in self.rows[:limit]: print(" " + str(row["features"]) + " -> " + row["label"]) if len(self.rows) > limit: print(" ... and " + str(len(self.rows) - limit) + " more") def new_dataset(name="dataset"): return Dataset(name) def row_field(row, name): """One named piece of a row: "the km of this row", "the label of it". Names come from the header line of the csv. "label" always works, even on a file with no header, because every row has one.""" wanted = str(name).strip() if not isinstance(row, dict): print("That is not a row. Use this inside a for each over a dataset.") return "" if wanted.lower() == "label": return row.get("label", "") fields = row.get("fields") or {} if wanted in fields: return fields[wanted] # Header names are matched loosely, so "Rain" finds the "rain" column. for key in fields: if str(key).strip().lower() == wanted.lower(): return fields[key] known = ", ".join([str(k) for k in fields]) if fields else "none" print( "No column called " + wanted + " in this row. Columns here: " + known + "." ) return "" def load_csv(dataset, path, label_column=-1, has_header=True): """Reads a comma separated file into a dataset. Everything except the label column becomes the features, and anything that looks like a number is turned into one. This is how a file you uploaded becomes something you can train on.""" try: with open(path) as handle: rows = [line.rstrip("\n").rstrip("\r") for line in handle] except OSError: print("Could not find " + path + ". Check the name in the Files panel.") return dataset rows = [row for row in rows if row.strip()] # A file with no label column is a perfectly normal thing to load: it is # what a test set looks like before you have predicted anything. labelled = str(label_column).strip().lower() not in ("none", "", "no", "-") header = [] if has_header and rows: header = [cell.strip() for cell in rows[0].split(",")] rows = rows[1:] added = 0 for row in rows: cells = [cell.strip() for cell in row.split(",")] if not cells or (labelled and len(cells) < 2): continue index = -1 if labelled: index = int(label_column) if index < 0: index = len(cells) + index if index < 0 or index >= len(cells): continue label = cells[index] if labelled else "" keep = [i for i in range(len(cells)) if i != index] typed = [] for i in keep: try: typed.append(float(cells[i])) except ValueError: typed.append(cells[i]) # Every kept column gets its header name, so "the km of a row" works. fields = {} for position, i in enumerate(keep): if i < len(header) and header[i]: fields[header[i]] = typed[position] if header and not dataset.columns: dataset.columns = [header[i] for i in keep if i < len(header)] dataset.add(typed[0] if len(typed) == 1 else typed, label, fields) added += 1 print("Loaded " + str(added) + " rows from " + path + ".") return dataset def save_submission(predictions, dataset, path="submission.csv"): """Prints your predictions in the exact two column format a round is scored in: a header, then one id and one label per line. It is printed rather than saved to a file because the console is the one place you can copy it from. Paste it into a new file in the Files panel, or straight into the upload box.""" labels = list(predictions or []) rows = list(dataset) if dataset is not None else [] if len(labels) != len(rows): print( "You have " + str(len(labels)) + " predictions for " + str(len(rows)) + " rows. Those have to match before this means anything." ) return lines = ["id,label"] for position, row in enumerate(rows): # Looked up directly rather than through row_field, which would # complain on every row of a file that simply has no id column. fields = row.get("fields") or {} row_id = "" for key in fields: if str(key).strip().lower() == "id": row_id = fields[key] break if not str(row_id).strip(): row_id = "r" + str(position + 1).zfill(2) if isinstance(row_id, float) and row_id == int(row_id): row_id = int(row_id) lines.append(str(row_id) + "," + str(labels[position]).strip()) print("--- submission.csv, copy from here ---") for line in lines: print(line) print("--- to here, " + str(len(lines) - 1) + " rows ---") def _words_in(value): letters = [] for character in str(value).lower(): letters.append(character if character.isalnum() else " ") return [word for word in "".join(letters).split() if word] # Words that turn up in every kind of sentence carry no signal about the# label, so the word matching model looks past them._EVERYDAY_WORDS = set( "a an and are as at be been but by can could did do for from had has have " "he her his i if in is it its me my not of on or our she so than that the " "their them then there they this to too us was we were what when which " "who will with would you your".split()) def _content_words(value): words = _words_in(value) kept = [word for word in words if word not in _EVERYDAY_WORDS] return kept if kept else words def _as_numbers(value): if isinstance(value, (list, tuple)): return [float(item) for item in value] return [float(value)] def _is_numeric(value): try: _as_numbers(value) return True except (TypeError, ValueError): return False def _distance(left, right): """How far apart two examples are. Numbers use straight line distance, text uses how many words the two do not share.""" if _is_numeric(left) and _is_numeric(right): a = _as_numbers(left) b = _as_numbers(right) while len(a) < len(b): a.append(0.0) while len(b) < len(a): b.append(0.0) total = 0.0 for i in range(len(a)): total += (a[i] - b[i]) ** 2 return math.sqrt(total) a = set(_words_in(left)) b = set(_words_in(right)) if not a and not b: return 0.0 shared = len(a & b) return 1.0 - (shared / float(len(a | b))) class Model: """Three small classifiers behind one name. nearest looks for the closest example it was trained on words scores how many words the input shares with each label common always answers with the most common label, the baseline to beat """ def __init__(self, kind="nearest", name="model"): self.kind = kind self.name = name self.data = None self.word_scores = {} def train(self, dataset): self.data = dataset self.word_scores = {} if self.kind == "words": for row in dataset.rows: bucket = self.word_scores.setdefault(row["label"], {}) for word in _content_words(row["features"]): bucket[word] = bucket.get(word, 0) + 1 print("Trained " + self.name + " on " + str(dataset.size()) + " examples.") def _scores(self, features): if self.data is None or not self.data.rows: return {} if self.kind == "common": counts = {} for row in self.data.rows: counts[row["label"]] = counts.get(row["label"], 0) + 1 return counts if self.kind == "words": # Score each label by how much of its training vocabulary shows up # in the input, divided by how much text that label was trained on # so a label with more examples cannot win on volume alone. labels = self.data.labels() scores = {} for label in labels: bucket = self.word_scores.get(label, {}) seen = sum(bucket.values()) or 1 running = 0.0 for word in _content_words(features): running += bucket.get(word, 0) / float(seen) scores[label] = round(running, 6) if sum(scores.values()) == 0: # Nothing in the input was ever seen in training. Say so by # splitting the vote evenly, which reads as low confidence. return dict((label, 1) for label in labels) return scores ranked = sorted(self.data.rows, key=lambda row: _distance(features, row["features"])) neighbours = ranked[: min(3, len(ranked))] scores = {} for row in neighbours: scores[row["label"]] = scores.get(row["label"], 0) + 1 return scores def predict(self, features): scores = self._scores(features) if not scores: return "" return max(scores, key=lambda label: scores[label]) def confidence(self, features): """How much of the vote the winning label took, out of 100.""" scores = self._scores(features) total = sum(scores.values()) if not scores or total == 0: return 0.0 best = max(scores.values()) return round(100.0 * best / float(total), 1) def accuracy(self, dataset): if not dataset.rows: return 0.0 right = 0 for row in dataset.rows: if self.predict(row["features"]) == row["label"]: right += 1 return round(100.0 * right / float(len(dataset.rows)), 1) def show(self): print(self.name + " is a " + self.kind + " model.") if self.data is None: print(" It has not been trained yet.") return print(" Trained on " + str(self.data.size()) + " examples.") print(" Labels it can answer with: " + ", ".join(self.data.labels())) def new_model(kind="nearest", name="model"): return Model(kind, name) # ---------------------------------------------------------------# Your script# --------------------------------------------------------------- examples = new_dataset("examples")examples.add("Please explain gravity", "question")examples.add("Could you explain magnets", "question")examples.add("Why does ice melt", "question")examples.add("How does a battery work", "question")examples.add("Where can I find the library", "question")examples.add("Can you describe a volcano", "question")examples.add("What causes thunder", "question")examples.add("How do seeds grow", "question")examples.add("The parcel arrived safely", "statement")examples.add("My parcel arrived yesterday", "statement")examples.add("Dinner tastes delicious", "statement")examples.add("Our team won today", "statement")examples.add("The blue door is closed", "statement")examples.add("Music played softly", "statement")examples.add("The dog slept peacefully", "statement")examples.add("We painted the fence", "statement")judge = new_model("words", "judge")judge.train(examples)message = input()label = judge.predict(message)sure = judge.confidence(message)print("Input", message)if (sure < 60): print("I am not sure. Please rephrase.")else: if (label == "question"): print("That looks like a question.") else: print("That looks like a statement.")print("Raw label", label)print("Raw score", sure)Compare this with your version. Different names and personal choices are fine when the program follows the same logic.
Keep your progress
Sign in and every reading, quiz, and exercise you finish is saved.