A Classifier in Machine Learning for Kids and Scratch
Transfer the classifier-and-application pattern to an optional Scratch-based tool without making it a prerequisite for the course.
Learn the operations rather than a second interface
This article's complete greeting/farewell classifier runs in the practice editor beside the reading. The optional Machine Learning for Kids and Scratch comparison explains how the same model-and-application pattern can transfer to another environment. No external account, installation, or second editor is needed for the course's required work.
A model learns associations between inputs and labels. An application collects a new input, asks the model for a result, checks a score if appropriate, and chooses a response. Training does not automatically build the application's branching behaviour. Another tool might display a model as an extension block, while code.guide creates and trains it explicitly in the current program.
Define the labels before gathering phrases
A greeting opens a conversation; a farewell closes it. Our teaching collection uses hello, good morning, nice to see you, welcome back, goodbye, see you later, take care, and until tomorrow. These eight examples show the mechanics of fitting, not a sufficient universal collection for recognising all language.
Build with me · 1
Define greeting and farewell consistently
Greeting opens a conversation and farewell closes it. These eight phrases provide a small teaching collection. A mixed phrase such as hello, I need to go falls outside the simple first definition and should be treated deliberately rather than silently given whichever label is convenient.
Collection and label rules are conceptually the same whether a tool displays rows in a project or blocks in a stack.
Create phrases and inspect all eight labelled examples.
What to look for
The dataset contains four greetings and four farewells.
Make it yours
Write one new phrase for each class and reserve them for a later probe.
Mixed messages such as hello, I need to go require an explicit scope decision. For the first version, treat them as outside the simple two-class task and record that limitation. Keep different phrases for development rather than evaluating only the exact training rows. Once you use a probe to improve the collection or settings, it is development evidence rather than an untouched final test.
Build with me · 2
Create and fit the recogniser
The model uses the labelled collection to learn word associations. Your application still needs to read a message, call prediction, and decide how to respond. Model training does not write that surrounding interaction automatically.
Another visual programming tool may expose a trained model through an extension, while this editor's make and train blocks establish it directly in the current run.
Create a word-matching greeter and train it on phrases below the collection.
What to look for
The model report shows eight training examples.
Make it yours
Point out the creation operation and the training operation separately.
Map the data flow across tools
An ask block and an answer value in one environment correspond to reading a line and storing message here. A trained-model classification block corresponds to what greeter says about message. A score block corresponds to the confidence value. A conditional response is still an if/else tree in either environment.
Build with me · 3
Store the user's input before predicting
Storing the input gives later prediction and score calls one shared value. In other block environments an ask-and-answer pair may perform the same role. Here the ask-text block pauses Terminal until you send an answer.
The vocabulary changes between interfaces, but the data flow remains input to stored value to model.
Press Run. When Terminal waits, type hello and press Enter or Send; the stack then stores that answer in message.
What to look for
After the answer hello, Read message is hello.
Make it yours
Run again with a reserved phrase and confirm the exact text was read.
Store the input once and read the same message in both label and score operations. A label for one input paired with a score for another is not a meaningful result. The model version must also stay the same while comparing those two outputs.
Scores can use different scales across tools. This teaching model reports zero to one hundred. A threshold of seventy corresponds numerically to 0.70 on a zero-to-one scale, but only after checking that the score's meaning is comparable. A confidence-shaped number is not automatically a calibrated probability or a guarantee of correctness.
Build with me · 4
Keep label and confidence on the same input
Both outputs refer to message and greeter. Changing the input between calls could mismatch them. Different tools also use different score scales, so a threshold of seventy here corresponds to 0.70 only in a tool whose score ranges from zero to one.
Inspect the actual scale and meaning before transferring a numerical threshold. A familiar block label does not prove identical model semantics.
Read the same message variable in both model values and print the stored outputs.
What to look for
For hello, the toy returns greeting with strong support from its training vocabulary.
Make it yours
Try an unfamiliar phrase and record the actual label and score without assuming a high score proves the phrase was understood.
Trace the response before connecting it to a character
The outer branch asks whether score is below seventy. If so, it prints the unsure response. Otherwise, the inner branch checks whether the label is greeting and prints Hello or Goodbye. The outer gate must surround the verdict. Printing Hello before checking score would expose a verdict the application intended to defer.
Build with me · 5
Gate the application response before choosing a greeting
The low-confidence branch comes first. Only the other branch chooses Hello or Goodbye. That structure ensures an unsure case does not greet first and apologise later.
A threshold can defer some ambiguous inputs, but a high score on unrelated content remains possible. The condition is an application rule to test, not a guarantee created by the word confidence.
Build the outer below-70 test, then place the greeting/farewell decision inside otherwise. Press Run and enter a no-overlap phrase when Terminal waits.
What to look for
After the answer marble lantern, the 70-point gate defers the message.
Make it yours
Test a phrase that shares a known word but has another meaning. Observe whether the gate actually catches it.
For an illustrative score of eighty-four and label greeting, the first condition is false and the inner greeting check is true, so Hello appears. For farewell at eighty-one, the outer condition is also false, but the greeting check is false, so Goodbye appears. Greeting at fifty-five should print only the unsure response. These numbers explain control flow; they are not promised scores from the real model.
Build with me · 6
Trace an accepted greeting
A known greeting should pass the gate if its score reaches the threshold, then satisfy the greeting comparison. Trace the exact order rather than treating the response as one indivisible model action.
The model selected a label and score. Your if/else selected the human-facing message. These are different responsibilities and can fail separately.
Press Run, enter good morning when Terminal waits, then read the raw outputs and selected response.
What to look for
After that answer, the phrase follows the greeting branch and prints Hello! after its diagnostic values.
Make it yours
Swap the Hello and Goodbye messages deliberately, observe a correct model with wrong application behaviour, then Undo.
Build with me · 7
Trace the other accepted label
The farewell case passes the same outer threshold but fails the inner greeting equality test. It therefore enters the Goodbye branch. A nested otherwise belongs to its own if, which is why visual placement matters.
If a new class is added later, this two-way assumption must be revised. Otherwise every accepted label other than greeting would be treated as farewell.
Press Run, enter goodbye when Terminal waits, and trace both tests in order.
What to look for
After that answer, the phrase follows the farewell response and prints Goodbye!.
Make it yours
Explain how adding a third label would change the response tree rather than only the training data.
The nested block shapes make this ownership visible. A statement below the whole conditional runs regardless of its decision. A diagnostic Raw label output is useful during development, but distinguish it from a user-facing verdict if your application intends to withhold uncertain predictions.
Keep the model and application errors separate
If a new greeting is correctly labelled but the character says Goodbye, inspect the application's branches. If the model predicts farewell with a high score, a correctly wired greeting branch will not fix the representation or collection. Both components can fail, and the remedy depends on which intermediate result first became wrong.
Test a reserved greeting, a reserved farewell, a mixed message, and an unrelated sentence. Record intended interpretation, raw label, score, application response, and whether the gate deferred. Do not report only accepted easy cases while hiding how often the application refuses to answer.
Transfer after understanding the local version
Optional Scratch-based learning services may organise editable training projects, fitted models, and applications as separate artifacts or extensions. Access and exact block names depend on the tool and classroom setup. Those interface details are not prerequisites here. You can explain the transferable roles from the working local program without creating an account elsewhere.
Keep source code separate from model parameters and from an editable training collection when the tool stores them separately. Improving examples requires another training operation and a known model version. Merely supplying a new prediction input is not continued learning. The reference at the bottom records the complete local application so you can compare your own construction one step at a time.
Build with me · 8
Name the operations that transfer across tools
The portable idea is a sequence of roles: create labelled data, fit a model, obtain input, predict and score that input, apply a gate, and respond. Another tool can express those roles through different blocks or typed code.
Save the source application and the training collection or model artifacts according to what each tool actually supports. A project used for editing is not automatically a fitted model file accepted by a different runtime.
Run the complete program here, answer when Terminal waits, and explain each operation by its role. The optional external tools discussed in the reading are comparisons, not prerequisites.
What to look for
After the answer hello, the program produces a labelled diagnostic report and an application response.
Make it yours
Use your reserved greeting, farewell, mixed, and unrelated probes. Record wrong accepted answers and deferrals separately.
Full reference solution
This is the complete worked program. Try building it yourself first, then use this reference to find the first place your version behaves differently. The Python below is generated from these exact blocks; helper functions are included so its behaviour can be inspected.
# ---------------------------------------------------------------# Building blocks, written out in plain Python.# This part is generated for you. Your script starts further down.# --------------------------------------------------------------- import mathimport random class Dataset: """A pile of labelled examples. Features can be a number, some text, or a list of numbers. The label is whatever answer you want back.""" def __init__(self, name="dataset"): self.name = name self.rows = [] # Column names from the file's header row, when it came from one. # Without these "the km of a row" has nothing to look the name up in. self.columns = [] def add(self, features, label, fields=None): self.rows.append( { "features": features, "label": str(label), "fields": dict(fields) if fields else {}, } ) def size(self): return len(self.rows) def __len__(self): return len(self.rows) def __iter__(self): """Walking a dataset gives you its rows, so "for each row in list [testing]" reads the way it sounds.""" return iter(self.rows) def labels(self): seen = [] for row in self.rows: if row["label"] not in seen: seen.append(row["label"]) return seen def most_common_label(self): if not self.rows: return "" counts = {} for row in self.rows: counts[row["label"]] = counts.get(row["label"], 0) + 1 return max(counts, key=lambda label: counts[label]) def split(self, train_percent=80): """Keeps the given percent for training and hands back the rest as a test set. The shuffle is seeded, so you get the same split every run.""" order = list(range(len(self.rows))) random.Random(0).shuffle(order) cut = int(len(order) * train_percent / 100) train = Dataset(self.name + " (train)") test = Dataset(self.name + " (test)") train.columns = list(self.columns) test.columns = list(self.columns) for position, index in enumerate(order): row = self.rows[index] target = train if position < cut else test target.add(row["features"], row["label"], row.get("fields")) return train, test def show(self, limit=10): print(self.name + ": " + str(len(self.rows)) + " examples") for row in self.rows[:limit]: print(" " + str(row["features"]) + " -> " + row["label"]) if len(self.rows) > limit: print(" ... and " + str(len(self.rows) - limit) + " more") def new_dataset(name="dataset"): return Dataset(name) def row_field(row, name): """One named piece of a row: "the km of this row", "the label of it". Names come from the header line of the csv. "label" always works, even on a file with no header, because every row has one.""" wanted = str(name).strip() if not isinstance(row, dict): print("That is not a row. Use this inside a for each over a dataset.") return "" if wanted.lower() == "label": return row.get("label", "") fields = row.get("fields") or {} if wanted in fields: return fields[wanted] # Header names are matched loosely, so "Rain" finds the "rain" column. for key in fields: if str(key).strip().lower() == wanted.lower(): return fields[key] known = ", ".join([str(k) for k in fields]) if fields else "none" print( "No column called " + wanted + " in this row. Columns here: " + known + "." ) return "" def load_csv(dataset, path, label_column=-1, has_header=True): """Reads a comma separated file into a dataset. Everything except the label column becomes the features, and anything that looks like a number is turned into one. This is how a file you uploaded becomes something you can train on.""" try: with open(path) as handle: rows = [line.rstrip("\n").rstrip("\r") for line in handle] except OSError: print("Could not find " + path + ". Check the name in the Files panel.") return dataset rows = [row for row in rows if row.strip()] # A file with no label column is a perfectly normal thing to load: it is # what a test set looks like before you have predicted anything. labelled = str(label_column).strip().lower() not in ("none", "", "no", "-") header = [] if has_header and rows: header = [cell.strip() for cell in rows[0].split(",")] rows = rows[1:] added = 0 for row in rows: cells = [cell.strip() for cell in row.split(",")] if not cells or (labelled and len(cells) < 2): continue index = -1 if labelled: index = int(label_column) if index < 0: index = len(cells) + index if index < 0 or index >= len(cells): continue label = cells[index] if labelled else "" keep = [i for i in range(len(cells)) if i != index] typed = [] for i in keep: try: typed.append(float(cells[i])) except ValueError: typed.append(cells[i]) # Every kept column gets its header name, so "the km of a row" works. fields = {} for position, i in enumerate(keep): if i < len(header) and header[i]: fields[header[i]] = typed[position] if header and not dataset.columns: dataset.columns = [header[i] for i in keep if i < len(header)] dataset.add(typed[0] if len(typed) == 1 else typed, label, fields) added += 1 print("Loaded " + str(added) + " rows from " + path + ".") return dataset def save_submission(predictions, dataset, path="submission.csv"): """Prints your predictions in the exact two column format a round is scored in: a header, then one id and one label per line. It is printed rather than saved to a file because the console is the one place you can copy it from. Paste it into a new file in the Files panel, or straight into the upload box.""" labels = list(predictions or []) rows = list(dataset) if dataset is not None else [] if len(labels) != len(rows): print( "You have " + str(len(labels)) + " predictions for " + str(len(rows)) + " rows. Those have to match before this means anything." ) return lines = ["id,label"] for position, row in enumerate(rows): # Looked up directly rather than through row_field, which would # complain on every row of a file that simply has no id column. fields = row.get("fields") or {} row_id = "" for key in fields: if str(key).strip().lower() == "id": row_id = fields[key] break if not str(row_id).strip(): row_id = "r" + str(position + 1).zfill(2) if isinstance(row_id, float) and row_id == int(row_id): row_id = int(row_id) lines.append(str(row_id) + "," + str(labels[position]).strip()) print("--- submission.csv, copy from here ---") for line in lines: print(line) print("--- to here, " + str(len(lines) - 1) + " rows ---") def _words_in(value): letters = [] for character in str(value).lower(): letters.append(character if character.isalnum() else " ") return [word for word in "".join(letters).split() if word] # Words that turn up in every kind of sentence carry no signal about the# label, so the word matching model looks past them._EVERYDAY_WORDS = set( "a an and are as at be been but by can could did do for from had has have " "he her his i if in is it its me my not of on or our she so than that the " "their them then there they this to too us was we were what when which " "who will with would you your".split()) def _content_words(value): words = _words_in(value) kept = [word for word in words if word not in _EVERYDAY_WORDS] return kept if kept else words def _as_numbers(value): if isinstance(value, (list, tuple)): return [float(item) for item in value] return [float(value)] def _is_numeric(value): try: _as_numbers(value) return True except (TypeError, ValueError): return False def _distance(left, right): """How far apart two examples are. Numbers use straight line distance, text uses how many words the two do not share.""" if _is_numeric(left) and _is_numeric(right): a = _as_numbers(left) b = _as_numbers(right) while len(a) < len(b): a.append(0.0) while len(b) < len(a): b.append(0.0) total = 0.0 for i in range(len(a)): total += (a[i] - b[i]) ** 2 return math.sqrt(total) a = set(_words_in(left)) b = set(_words_in(right)) if not a and not b: return 0.0 shared = len(a & b) return 1.0 - (shared / float(len(a | b))) class Model: """Three small classifiers behind one name. nearest looks for the closest example it was trained on words scores how many words the input shares with each label common always answers with the most common label, the baseline to beat """ def __init__(self, kind="nearest", name="model"): self.kind = kind self.name = name self.data = None self.word_scores = {} def train(self, dataset): self.data = dataset self.word_scores = {} if self.kind == "words": for row in dataset.rows: bucket = self.word_scores.setdefault(row["label"], {}) for word in _content_words(row["features"]): bucket[word] = bucket.get(word, 0) + 1 print("Trained " + self.name + " on " + str(dataset.size()) + " examples.") def _scores(self, features): if self.data is None or not self.data.rows: return {} if self.kind == "common": counts = {} for row in self.data.rows: counts[row["label"]] = counts.get(row["label"], 0) + 1 return counts if self.kind == "words": # Score each label by how much of its training vocabulary shows up # in the input, divided by how much text that label was trained on # so a label with more examples cannot win on volume alone. labels = self.data.labels() scores = {} for label in labels: bucket = self.word_scores.get(label, {}) seen = sum(bucket.values()) or 1 running = 0.0 for word in _content_words(features): running += bucket.get(word, 0) / float(seen) scores[label] = round(running, 6) if sum(scores.values()) == 0: # Nothing in the input was ever seen in training. Say so by # splitting the vote evenly, which reads as low confidence. return dict((label, 1) for label in labels) return scores ranked = sorted(self.data.rows, key=lambda row: _distance(features, row["features"])) neighbours = ranked[: min(3, len(ranked))] scores = {} for row in neighbours: scores[row["label"]] = scores.get(row["label"], 0) + 1 return scores def predict(self, features): scores = self._scores(features) if not scores: return "" return max(scores, key=lambda label: scores[label]) def confidence(self, features): """How much of the vote the winning label took, out of 100.""" scores = self._scores(features) total = sum(scores.values()) if not scores or total == 0: return 0.0 best = max(scores.values()) return round(100.0 * best / float(total), 1) def accuracy(self, dataset): if not dataset.rows: return 0.0 right = 0 for row in dataset.rows: if self.predict(row["features"]) == row["label"]: right += 1 return round(100.0 * right / float(len(dataset.rows)), 1) def show(self): print(self.name + " is a " + self.kind + " model.") if self.data is None: print(" It has not been trained yet.") return print(" Trained on " + str(self.data.size()) + " examples.") print(" Labels it can answer with: " + ", ".join(self.data.labels())) def new_model(kind="nearest", name="model"): return Model(kind, name) # ---------------------------------------------------------------# Your script# --------------------------------------------------------------- phrases = new_dataset("phrases")phrases.add("hello", "greeting")phrases.add("good morning", "greeting")phrases.add("nice to see you", "greeting")phrases.add("welcome back", "greeting")phrases.add("goodbye", "farewell")phrases.add("see you later", "farewell")phrases.add("take care", "farewell")phrases.add("until tomorrow", "farewell")greeter = new_model("words", "greeter")greeter.train(phrases)message = input()label = greeter.predict(message)sure = greeter.confidence(message)print("Input", message)print("Predicted label", label)print("Score out of 100", sure)if (sure < 70): print("I am unsure what you meant.")else: if (label == "greeting"): print("Hello!") else: print("Goodbye!")Compare this with your version. Different names and personal choices are fine when the program follows the same logic.
Worth a look
Keep your progress
Sign in and every reading, quiz, and exercise you finish is saved.