How a model sees a picture
Translate photos into pixels, channels, features, and model inputs, and explain what changes between camera models and handwritten digits.
Turn an image into numbers
Zoom into a photograph until you can see its square pixels. Each pixel is stored as numbers. A grayscale pixel needs one number: its brightness. A colour pixel needs three: how much red, how much green, and how much blue. Each of those three is called a channel, and a photo stored this way is called RGB, after the three colours. Mixing the three channels produces the colour you see.
So a 200-by-200 grayscale image is 40,000 numbers, and the same image in colour is 120,000 numbers, three per pixel. Both look like a square photo on screen, but to a model they are different shapes of input. A model needs exactly the shape it was built to accept.
In an ordinary photo, each channel value runs from 0 (none) to 255 (the most). Other datasets use other ranges. The small handwritten digit dataset in Level 4 runs from 0 to 16, so dividing every dataset by 255 would be a mistake.
Preprocessing creates the expected input
A camera photo may be large, rectangular, and in colour, while a model might need a small square with values in a set range. Before a photo reaches the model, it is prepared: resized, cropped, perhaps turned to grayscale, and rescaled. This preparation is called preprocessing.
Training photos and prediction photos must go through the same preprocessing. If training shows a light digit centred on a dark background, a small dark digit on a large white page gives very different numbers. Preprocessing is part of the recogniser, not an optional extra.
The image blocks do all of this for you inside the browser. In Level 4 you will write the same steps yourself for a drawn digit. Even when a library does the work, you should be able to say what shape and range of numbers reaches the model.
Features and layers
In Level 2 you met the pretrained network: a network trained earlier on a very large collection of photos, which turns any photo into numerical features. Your camera models in this module use the same arrangement. The pretrained part turns pixels into features (people sometimes call it a feature extractor), and your photos train only a small classifier on top. Reusing a network trained on another task in this way is the transfer learning you met in Level 2. It is why thirty photos can produce a working demonstration, and also why thirty photos cannot guarantee good recognition across people, cameras, and rooms.
Inside, the network applies layers of calculations one after another. In many image networks, early layers respond to small local patterns such as edges and textures, and later layers combine them into larger patterns. Treat that as a helpful picture, not a promise that any single part of the network stands for a nameable thing like "finger".
A feature is any of these calculated numbers that helps with the task. The classifier combines features to score each label. It is numbers all the way through. Nowhere is there a written rule such as "two fingers means scissors" unless someone separately programmed one.
Probe the recogniser one condition at a time
The stages below train a model on two everyday objects, then change one condition at a time, such as distance or lighting, to see which changes the model copes with.
Build with me · 1
Choose an observable two-class task
An image is a grid of numerical brightness or colour values. Its label is supplied by the training task. Choosing light-object and dark-object here gives us a convenient experiment, but those labels refer to the selected objects rather than a complete universal rule about brightness.
The pretrained network turns pixels into features. Your small training collection then links those features with your class names.
Pick two safe everyday objects and write down the classes. Create objects in the editor.
What to look for
The task reminder prints without training or camera capture.
Make it yours
Predict which changes that should not matter might still affect it: distance, lighting, background, rotation, or part of an object being hidden.
Build with me · 2
Collect both classes in comparable conditions
A model can use the background as a shortcut if the background lines up with a label. Photographing one class only against a keyboard and the other only against a wall makes that shortcut easy. Keeping both classes across the same backgrounds weakens it.
Balance does not mean identical photos. You want useful variation in pose and distance while avoiding an irrelevant condition that uniquely identifies the class.
Capture each labelled group with similar backgrounds and lighting. Inspect the collection before training.
What to look for
The report contains both classes. Its count alone cannot establish that shortcuts were avoided.
Make it yours
Make a two-column list of conditions represented for each class. Identify a gap before collecting a final evaluation.
Why a shortcut survives the layers
Suppose every scissors photo had a bright window behind the hand, and no other class did. The features describe the whole photo, so they can describe that window as well as the hand. The classifier uses whatever helps it get the training photos right, and the window helps.
Move to another room and the window is gone, and so is the model's accuracy. You can investigate this with controlled tests: keep the gesture the same while changing the background, then keep the background the same while changing the gesture. A score alone cannot tell you which feature the model relied on.
The fix starts with collection: vary the things that should not matter across every label, and include the conditions where the model will really be used.
The next three stages test your trained model under changed conditions, starting from a reference photo to compare against.
Build with me · 3
Record a reference prediction
A reference view gives later comparisons a starting point. Use a photo that was not one of the collected training frames, while keeping a familiar distance and background. Record the intended class before looking at the output.
The score compares only the classes the model knows. It says nothing about how well the model has captured the rest of the image.
Build setup and one prediction. Run and capture the reference view after training finishes.
What to look for
The actual label and score print; either can reveal an error.
Make it yours
Write the result with the conditions. A high score with a wrong label is still a failed prediction.
Build with me · 4
Compare two distances with one trained model
These two predictions share one trained model. Moving farther away makes the object occupy fewer pixels, while more background enters the image. That changes the numbers the model receives, even though the object is the same.
The source deliberately prints an instruction before each capture. You can follow the test without leaving the article or remembering which photo was intended first.
Capture the same object first at a familiar distance and then farther away. Keep the background and lighting as stable as you can.
What to look for
Two labelled result groups appear. A change in score or label is an observation, not a guaranteed outcome.
Make it yours
Repeat with a second object. Decide whether the first result was a consistent pattern or a single example.
Build with me · 5
Vary lighting without changing the class
Different illumination changes pixel values and sometimes shadows or reflections. A light object in shadow can resemble a darker object under bright light. The trained model may or may not cope with that change.
Do not change several conditions at once if you want to identify the likely cause. Hold the object, distance, and background approximately fixed while comparing lighting.
Capture the same object in two ordinary lighting conditions. Do not use intense lights or anything uncomfortable; the experiment only needs a visible change.
What to look for
Both measured labels and scores are printed for comparison.
Make it yours
If a failure appears, propose collecting examples of that condition for both classes. Keep a different view reserved for checking the revised model.
A model with two classes has only two answers, whatever it is shown.
Build with me · 6
Try something outside the label set
A classifier asked to choose between two classes often still chooses one when neither is appropriate. Its class scores compare available alternatives; they do not automatically include a verified none-of-the-above answer.
This explains why a high confidence score on an empty desk is possible. The response tells you about the model and label set, not that the desk secretly is one of your objects.
After training, capture an ordinary scene containing neither selected object. Compare the prediction with the actual task definition.
What to look for
One of the trained labels can appear, even with a high score. There is no promise of a low score on unfamiliar content.
Make it yours
Record what the application should do in this case, and distinguish that desired behaviour from what a threshold actually demonstrated.
Build with me · 7
Evaluate an application threshold
The gate makes a visible decision: accept the label, or defer it, which means hold it back instead of using it. Test the gate on ordinary examples and unusual ones, and record whether each label was correct separately from whether it was accepted. A correct label that was held back and a wrong label that was accepted are different outcomes.
Changing the threshold does not change learned features. It changes the application's rule for using scores, which can be evaluated separately from collecting or retraining.
Keep raw label and score visible before the accept/defer message. Capture three planned conditions with the same trained model.
What to look for
Each capture has a label, score, and gate decision. Those three values together support a useful observation.
Make it yours
Try a different threshold after recording the first version. Count the labels held back and the wrong labels accepted.
Build with me · 8
Finish with a repeatable inspection program
A useful explanation connects layers of the process: the scene produced pixels, preprocessing and the pretrained network turned them into features, your trained classifier produced scores, and your program interpreted those scores. You have observed behaviour rather than identified the exact purpose of every internal unit.
Use this complete stack to make four planned comparisons. Once you use the failures to revise data or settings, those images are development evidence. Keep a genuinely separate final check if you want an honest result for the selected version.
Plan four images before running, then record each result by its number. Compare one changed condition at a time where possible.
What to look for
Four numbered results use one trained model. Your written explanation should refer to the actual observed images and outputs.
Make it yours
Explain one failure without saying the model understands or does not understand the object. Name the changed input condition and the evidence you have.
Connect images to the rest of the course
The text matcher turned words into counts. The camera network turns pixels into learned features. A table model, later in this level, starts with columns such as distance and rain. In every case, what the model is given decides what it can possibly learn.
Try two questions. Why can turning a digit image to grayscale be sensible? Its shape carries the answer, and colour does not. Why would keeping only its average brightness lose too much? A 3 and an 8 can have similar average brightness while their strokes are arranged differently. A good way of turning an input into numbers keeps the differences the task needs.
Full reference solution
This is the complete worked program. Try building it yourself first, then use this reference to find the first place your version behaves differently. The Python below is generated from these exact blocks; helper functions are included so its behaviour can be inspected.
# ---------------------------------------------------------------# This file uses image blocks, which need a camera and a browser.# That part runs here rather than anywhere Python runs.# Everything else below is ordinary Python.# --------------------------------------------------------------- # ---------------------------------------------------------------# Building blocks, written out in plain Python.# This part is generated for you. Your script starts further down.# --------------------------------------------------------------- # These blocks need a camera and a browser, so this part runs here rather# than anywhere Python runs. Everything above this line is ordinary Python.from js_images import pagefrom pyodide.ffi import run_sync class ImageModel: """A model that tells photos apart. It does not look at the pixels itself. A pretrained network turns each photo into a list of numbers, and a small model on top learns which lists go with which label.""" def __init__(self, name="model"): self.name = name self.labels = [] self.trained = False def capture(self, count, label): """Opens the camera panel and waits while you take the photos.""" taken = run_sync(page.capture(self.name, int(count), str(label))) if str(label) not in self.labels: self.labels.append(str(label)) print("Took " + str(taken) + " photos of '" + str(label) + "'.") def train(self): if len(self.labels) < 2: raise ValueError( "A model needs at least two labels to tell anything apart. " "Add photos under a second label before training." ) run_sync(page.train(self.name)) self.trained = True print("Trained " + self.name + " on " + str(len(self.labels)) + " labels.") def _ready(self): if not self.trained: raise ValueError( "This model has not been trained yet. Add photos and train it first." ) def classify(self, photo): """What the model thinks is in the photo.""" self._ready() return run_sync(page.classify(self.name, photo)) def confidence(self, photo): """Largest class score out of 100. Unfamiliar photos can still receive a high score; measure correctness separately on labelled examples.""" self._ready() return run_sync(page.confidence(self.name, photo)) def show(self): run_sync(page.show(self.name)) def new_image_model(name="model"): return ImageModel(name) def capture_one(): """Opens the camera panel for one photo and hands it back.""" return run_sync(page.capture_one()) def load_image_file(path, name="model"): """Restore local learned weights without training again.""" with open(str(path), "r", encoding="utf-8") as source: contents = source.read() model = ImageModel(name) model.labels = run_sync(page.load_image_file(name, contents)) model.trained = True print("Loaded " + name + " with " + str(len(model.labels)) + " labels.") return model def save_image_file(model, path="image-model.codeguide.json"): """Save the fitted model, separately from the program and training photos.""" model._ready() contents = run_sync(page.save_image_file(model.name)) with open(str(path), "w", encoding="utf-8") as target: target.write(contents) print("Saved " + str(path)) def load_teachable_machine(url, name="model"): """Loads a model you published from Teachable Machine. The link is the one under Export Model, Upload my model.""" model = ImageModel(name) model.labels = run_sync(page.load_teachable_machine(name, str(url))) model.trained = True print("Loaded " + name + " with " + str(len(model.labels)) + " labels.") return model # ---------------------------------------------------------------# Your script# --------------------------------------------------------------- objects = new_image_model("objects")objects.capture(10, "light-object")objects.capture(10, "dark-object")objects.show()objects.train()attempt = 0for _ in range(int(4)): attempt = attempt + 1 print("Test image", attempt) photo = capture_one() move = objects.classify(photo) sure = objects.confidence(photo) print("Recognised", move) print("Score", sure)print("Compare object, pixels, learned features, score, and application decision separately")Compare this with your version. Different names and personal choices are fine when the program follows the same logic.
Keep your progress
Sign in and every reading, quiz, and exercise you finish is saved.