Save your work and debug a block program
Tell a saved program apart from what exists only while it runs, and use a repeatable method to find broken connections, inputs, and model setup.
What is kept, and where
Each lesson's practice editor saves your blocks and text files in this browser, so they are still there after a refresh on the same device. That is not a backup. Clearing site data, using private browsing, switching browser, or moving to another device can lose them.
Download practice code saves a copy of the Python generated from your blocks. Keep it in a folder with any data files it reads and a short note about what it does. The download keeps the instructions, but it cannot be loaded back in as blocks.
Camera programs also rely on this website's camera helpers, so their downloaded Python will not run unchanged on its own on your computer.
Exercise workspaces are separate from each article's practice editor, and keep their own saved blocks. A completed progress tick records that you finished; it is not a copy of your program.
What a run rebuilds
A saved program is a recipe. Every Run starts again at the first block. If that block makes an empty dataset, the add blocks below it must fill it again. If a block makes a model, a train block below it must train it again. Saving the program does not save a trained model or any camera photos.
This matters most for camera programs: each Run clears the previous run's photos and trained model. Put collection, training, and every prediction you want in one program, and use a loop to make several predictions in the same run. The next module also shows how to save a trained image model as a file and load it in a later run without training again.
Start from a version that works
Debugging means finding the first place where what the program does differs from what you meant. The first three stages practise on tiny programs whose answers you can work out in your head.
Build with me · 1
Keep a small version that works
Before investigating a complicated model, practise on a program whose answer you can calculate. Two plus three is five. The checkpoint comes after both operations, so it observes the final value.
Saving this source preserves instructions. It does not store a running score that will resume at five next time. Every new Run starts with the assignment to two.
Run the three-block stack, then run it again without editing. Read the order of assignments each time.
What to look for
Both runs print Score 5, not five then eight.
Make it yours
Change the increment and record the new expected result before running.
Build with me · 2
Recognise text where a value should be
This program is valid, but its final instruction says the literal characters score. There is no error message, because printing text is allowed. The mistake is between what the programmer intended and what the socket actually contains.
Do not repair it by changing the calculation. The stored value was already correct. Fix the observation by inserting the rounded variable.
Load the deliberately incorrect output. Run once, then replace the typed text socket with the score value block.
What to look for
Before repair it prints score. After repair it prints 5.
Make it yours
Use Undo to recover the wrong version, then Redo. Explain what changed in source and why old output does not automatically change with Undo.
Build with me · 3
Find a reset in the wrong place
The set block is inside the loop, so it resets total to zero on every turn. The following addition always changes zero to two. This is an ordering mistake, not a broken repeat block.
Trace the first two turns explicitly. The second turn does not start from the previous two because the first instruction overwrites it. The trace output tells you where to look.
Run this deliberate reset example. Move only set total to 0 above the repeat block, leaving change and output inside.
What to look for
Before repair the three values are 2, 2, 2. After repair they are 2, 4, 6.
Make it yours
Move the labelled output below repeat and predict how many lines remain. Restore the trace while debugging.
Find the first point where reality differs
When a model program gives a wrong answer, do not replace everything. Read the program in order and predict the value at each step. Add labelled say blocks at a few checkpoints: the dataset count after collection, the input after it is read, the label the model returns, and the score before the gate. The first checkpoint that shows something unexpected is where to look.
If the dataset count is zero, a wrong label later on is not the first problem. Check that the add blocks are connected, name the right dataset, and run before training. If the input is empty, look at how it was read before changing anything about the model. The next four stages each hide one of these problems.
Build with me · 4
Check collection before blaming prediction
This stack trains on only the first two request examples. It has no statement examples at all. A strange statement prediction starts with that collection problem, before any threshold or response branch exists.
The dataset count is the earliest useful checkpoint. Inspect the labels too: a count of sixteen would not prove that all intended classes were represented.
Run and inspect Collected, then read both add-example labels. Compare them with the intended two-class task.
What to look for
Collected is 2 and both rows are question examples. The model cannot learn statement associations from absent rows.
Make it yours
Load the complete sixteen-row dataset stage from the classifier article or add a statement example here. Check the count and labels before retraining.
Build with me · 5
Print what was actually read
When input differs from what you intended, changing training can hide the actual problem. Print the stored message immediately after the read. In this one-question program, the Terminal wait collects exactly the answer used by the prediction.
An empty answer is still an answer. Matching each question with the response you send is part of debugging.
Press Run, type the intended message when Terminal waits, and press Enter or Send. Then inspect Read input in Terminal output.
What to look for
Read input exactly matches the Terminal answer before the prediction appears.
Make it yours
Run once with an empty answer, observe what the checkpoint reveals, then run again with the intended message.
Build with me · 6
Check the model after the last make-model block
The second make judge statement creates a new empty model and stores it under the same name. The original trained model is no longer what this name refers to. Its earlier training message can therefore mislead you if you ignore the later block.
Read all statements after training when a model appears untrained. A correct early operation does not prevent a later operation from replacing its result.
Run the program with its make-model block deliberately repeated, and inspect the final model report. Delete only the second make-model statement, then run again.
What to look for
Before repair the final model has zero training examples. After repair its report shows sixteen training examples.
Make it yours
Explain why adding a second train after the duplicate would run, but removing the accidental duplicate expresses the intended program more clearly.
Build with me · 7
Put the verdict behind its condition
Here the program prints a label before checking whether it should say it is unsure. The gate still works, but the earlier verdict has already been shown. That is an application-flow problem: the message was placed outside the decision that should control it.
Keep diagnostic labels visibly marked as diagnostic during development, and decide which results the person using the program should see. A confidence gate does not automatically suppress earlier prints.
Press Run and answer an unfamiliar message when Terminal waits. Move or remove the premature verdict, leaving the gate to choose the response.
What to look for
Before repair the output shows a label and then an unsure response. The repaired application gives only its intended response.
Make it yours
Keep a separately labelled Raw label while investigating, then explain why a raw diagnostic result is different from the application's verdict.
The last stage puts all of these checkpoints into one working program.
Build with me · 8
Run a complete version with useful checkpoints
This reference keeps a count, the exact Terminal answer, and the raw prediction before the application response. Those checkpoints cover collection, input, the model's answer, and the decision, in that order. If something looks wrong, start at the first incorrect checkpoint.
The page saves your source in this browser, and Download practice code keeps a Python source copy. Neither one stores camera photos or a trained model. A new Run rebuilds everything from the recipe. Saving a trained model as a file is a separate step, shown in the next module.
Run the complete stack, answer when Terminal waits, then run again with a different answer. Download the source if you want a portable copy. Keep related text files with it and use the collapsed reference only after comparing your own structure.
What to look for
After the answer The parcel arrived safely, collection shows sixteen and the response identifies a statement. Raw values remain available for diagnosis.
Make it yours
Choose one deliberate change at a time: a missing example, a Terminal answer, or a threshold. Predict which checkpoint will first reveal the effect.
Read an error message
Sometimes a run stops with an error instead of a wrong answer. The last line of the error usually names the problem. A missing name means something was used before the block that creates it. A file-not-found error means a filename is misspelt or the file is not in Files. A type or value error means a socket or an input holds the wrong kind of value, such as a word where a number should be. The line number points at the instruction that failed.
Common structural mistakes
| Symptom | Check first |
|---|---|
| A block runs only once | Is it below the loop rather than inside it? |
| A message prints on every turn | Is it inside a loop that should finish before printing? |
| Output contains a variable's name | Did you type text instead of inserting its value block? |
| A model reports no training | Did a later make-model block replace the trained model? |
| A name menu is empty | Is the block that creates the name missing, or further down? |
| Old output appears unchanged | Did you press Run again after editing? |
| A run appears stuck | Is it waiting for a Terminal answer, a camera capture, training, or a long loop? |
Use Stop if a loop keeps running when it should not. Then fix the loop's count or condition before running again. If the canvas looks empty, bring everything back on screen before assuming the blocks were deleted. Undo reverses edits to your program; it cannot bring back browser data that has already been cleared.
A good debugging note says what you expected, what you observed, the first step that went wrong, and the change that fixed it. The same method works in Python in Level 4, with the same patience and fewer coloured shapes.
Full reference solution
This is the complete worked program. Try building it yourself first, then use this reference to find the first place your version behaves differently. The Python below is generated from these exact blocks; helper functions are included so its behaviour can be inspected.
# ---------------------------------------------------------------# Building blocks, written out in plain Python.# This part is generated for you. Your script starts further down.# --------------------------------------------------------------- import mathimport random class Dataset: """A pile of labelled examples. Features can be a number, some text, or a list of numbers. The label is whatever answer you want back.""" def __init__(self, name="dataset"): self.name = name self.rows = [] # Column names from the file's header row, when it came from one. # Without these "the km of a row" has nothing to look the name up in. self.columns = [] def add(self, features, label, fields=None): self.rows.append( { "features": features, "label": str(label), "fields": dict(fields) if fields else {}, } ) def size(self): return len(self.rows) def __len__(self): return len(self.rows) def __iter__(self): """Walking a dataset gives you its rows, so "for each row in list [testing]" reads the way it sounds.""" return iter(self.rows) def labels(self): seen = [] for row in self.rows: if row["label"] not in seen: seen.append(row["label"]) return seen def most_common_label(self): if not self.rows: return "" counts = {} for row in self.rows: counts[row["label"]] = counts.get(row["label"], 0) + 1 return max(counts, key=lambda label: counts[label]) def split(self, train_percent=80): """Keeps the given percent for training and hands back the rest as a test set. The shuffle is seeded, so you get the same split every run.""" order = list(range(len(self.rows))) random.Random(0).shuffle(order) cut = int(len(order) * train_percent / 100) train = Dataset(self.name + " (train)") test = Dataset(self.name + " (test)") train.columns = list(self.columns) test.columns = list(self.columns) for position, index in enumerate(order): row = self.rows[index] target = train if position < cut else test target.add(row["features"], row["label"], row.get("fields")) return train, test def show(self, limit=10): print(self.name + ": " + str(len(self.rows)) + " examples") for row in self.rows[:limit]: print(" " + str(row["features"]) + " -> " + row["label"]) if len(self.rows) > limit: print(" ... and " + str(len(self.rows) - limit) + " more") def new_dataset(name="dataset"): return Dataset(name) def row_field(row, name): """One named piece of a row: "the km of this row", "the label of it". Names come from the header line of the csv. "label" always works, even on a file with no header, because every row has one.""" wanted = str(name).strip() if not isinstance(row, dict): print("That is not a row. Use this inside a for each over a dataset.") return "" if wanted.lower() == "label": return row.get("label", "") fields = row.get("fields") or {} if wanted in fields: return fields[wanted] # Header names are matched loosely, so "Rain" finds the "rain" column. for key in fields: if str(key).strip().lower() == wanted.lower(): return fields[key] known = ", ".join([str(k) for k in fields]) if fields else "none" print( "No column called " + wanted + " in this row. Columns here: " + known + "." ) return "" def load_csv(dataset, path, label_column=-1, has_header=True): """Reads a comma separated file into a dataset. Everything except the label column becomes the features, and anything that looks like a number is turned into one. This is how a file you uploaded becomes something you can train on.""" try: with open(path) as handle: rows = [line.rstrip("\n").rstrip("\r") for line in handle] except OSError: print("Could not find " + path + ". Check the name in the Files panel.") return dataset rows = [row for row in rows if row.strip()] # A file with no label column is a perfectly normal thing to load: it is # what a test set looks like before you have predicted anything. labelled = str(label_column).strip().lower() not in ("none", "", "no", "-") header = [] if has_header and rows: header = [cell.strip() for cell in rows[0].split(",")] rows = rows[1:] added = 0 for row in rows: cells = [cell.strip() for cell in row.split(",")] if not cells or (labelled and len(cells) < 2): continue index = -1 if labelled: index = int(label_column) if index < 0: index = len(cells) + index if index < 0 or index >= len(cells): continue label = cells[index] if labelled else "" keep = [i for i in range(len(cells)) if i != index] typed = [] for i in keep: try: typed.append(float(cells[i])) except ValueError: typed.append(cells[i]) # Every kept column gets its header name, so "the km of a row" works. fields = {} for position, i in enumerate(keep): if i < len(header) and header[i]: fields[header[i]] = typed[position] if header and not dataset.columns: dataset.columns = [header[i] for i in keep if i < len(header)] dataset.add(typed[0] if len(typed) == 1 else typed, label, fields) added += 1 print("Loaded " + str(added) + " rows from " + path + ".") return dataset def save_submission(predictions, dataset, path="submission.csv"): """Prints your predictions in the exact two column format a round is scored in: a header, then one id and one label per line. It is printed rather than saved to a file because the console is the one place you can copy it from. Paste it into a new file in the Files panel, or straight into the upload box.""" labels = list(predictions or []) rows = list(dataset) if dataset is not None else [] if len(labels) != len(rows): print( "You have " + str(len(labels)) + " predictions for " + str(len(rows)) + " rows. Those have to match before this means anything." ) return lines = ["id,label"] for position, row in enumerate(rows): # Looked up directly rather than through row_field, which would # complain on every row of a file that simply has no id column. fields = row.get("fields") or {} row_id = "" for key in fields: if str(key).strip().lower() == "id": row_id = fields[key] break if not str(row_id).strip(): row_id = "r" + str(position + 1).zfill(2) if isinstance(row_id, float) and row_id == int(row_id): row_id = int(row_id) lines.append(str(row_id) + "," + str(labels[position]).strip()) print("--- submission.csv, copy from here ---") for line in lines: print(line) print("--- to here, " + str(len(lines) - 1) + " rows ---") def _words_in(value): letters = [] for character in str(value).lower(): letters.append(character if character.isalnum() else " ") return [word for word in "".join(letters).split() if word] # Words that turn up in every kind of sentence carry no signal about the# label, so the word matching model looks past them._EVERYDAY_WORDS = set( "a an and are as at be been but by can could did do for from had has have " "he her his i if in is it its me my not of on or our she so than that the " "their them then there they this to too us was we were what when which " "who will with would you your".split()) def _content_words(value): words = _words_in(value) kept = [word for word in words if word not in _EVERYDAY_WORDS] return kept if kept else words def _as_numbers(value): if isinstance(value, (list, tuple)): return [float(item) for item in value] return [float(value)] def _is_numeric(value): try: _as_numbers(value) return True except (TypeError, ValueError): return False def _distance(left, right): """How far apart two examples are. Numbers use straight line distance, text uses how many words the two do not share.""" if _is_numeric(left) and _is_numeric(right): a = _as_numbers(left) b = _as_numbers(right) while len(a) < len(b): a.append(0.0) while len(b) < len(a): b.append(0.0) total = 0.0 for i in range(len(a)): total += (a[i] - b[i]) ** 2 return math.sqrt(total) a = set(_words_in(left)) b = set(_words_in(right)) if not a and not b: return 0.0 shared = len(a & b) return 1.0 - (shared / float(len(a | b))) class Model: """Three small classifiers behind one name. nearest looks for the closest example it was trained on words scores how many words the input shares with each label common always answers with the most common label, the baseline to beat """ def __init__(self, kind="nearest", name="model"): self.kind = kind self.name = name self.data = None self.word_scores = {} def train(self, dataset): self.data = dataset self.word_scores = {} if self.kind == "words": for row in dataset.rows: bucket = self.word_scores.setdefault(row["label"], {}) for word in _content_words(row["features"]): bucket[word] = bucket.get(word, 0) + 1 print("Trained " + self.name + " on " + str(dataset.size()) + " examples.") def _scores(self, features): if self.data is None or not self.data.rows: return {} if self.kind == "common": counts = {} for row in self.data.rows: counts[row["label"]] = counts.get(row["label"], 0) + 1 return counts if self.kind == "words": # Score each label by how much of its training vocabulary shows up # in the input, divided by how much text that label was trained on # so a label with more examples cannot win on volume alone. labels = self.data.labels() scores = {} for label in labels: bucket = self.word_scores.get(label, {}) seen = sum(bucket.values()) or 1 running = 0.0 for word in _content_words(features): running += bucket.get(word, 0) / float(seen) scores[label] = round(running, 6) if sum(scores.values()) == 0: # Nothing in the input was ever seen in training. Say so by # splitting the vote evenly, which reads as low confidence. return dict((label, 1) for label in labels) return scores ranked = sorted(self.data.rows, key=lambda row: _distance(features, row["features"])) neighbours = ranked[: min(3, len(ranked))] scores = {} for row in neighbours: scores[row["label"]] = scores.get(row["label"], 0) + 1 return scores def predict(self, features): scores = self._scores(features) if not scores: return "" return max(scores, key=lambda label: scores[label]) def confidence(self, features): """How much of the vote the winning label took, out of 100.""" scores = self._scores(features) total = sum(scores.values()) if not scores or total == 0: return 0.0 best = max(scores.values()) return round(100.0 * best / float(total), 1) def accuracy(self, dataset): if not dataset.rows: return 0.0 right = 0 for row in dataset.rows: if self.predict(row["features"]) == row["label"]: right += 1 return round(100.0 * right / float(len(dataset.rows)), 1) def show(self): print(self.name + " is a " + self.kind + " model.") if self.data is None: print(" It has not been trained yet.") return print(" Trained on " + str(self.data.size()) + " examples.") print(" Labels it can answer with: " + ", ".join(self.data.labels())) def new_model(kind="nearest", name="model"): return Model(kind, name) # ---------------------------------------------------------------# Your script# --------------------------------------------------------------- examples = new_dataset("examples")examples.add("Please explain gravity", "question")examples.add("Could you explain magnets", "question")examples.add("Why does ice melt", "question")examples.add("How does a battery work", "question")examples.add("Where can I find the library", "question")examples.add("Can you describe a volcano", "question")examples.add("What causes thunder", "question")examples.add("How do seeds grow", "question")examples.add("The parcel arrived safely", "statement")examples.add("My parcel arrived yesterday", "statement")examples.add("Dinner tastes delicious", "statement")examples.add("Our team won today", "statement")examples.add("The blue door is closed", "statement")examples.add("Music played softly", "statement")examples.add("The dog slept peacefully", "statement")examples.add("We painted the fence", "statement")judge = new_model("words", "judge")judge.train(examples)print("Examples before prediction", examples.size())message = input()label = judge.predict(message)sure = judge.confidence(message)print("Read input", message)print("Raw label", label)print("Raw score", sure)if (sure < 60): print("I am not sure. Please rephrase.")else: if (label == "question"): print("That looks like a question.") else: print("That looks like a statement.")Compare this with your version. Different names and personal choices are fine when the program follows the same logic.
Keep your progress
Sign in and every reading, quiz, and exercise you finish is saved.