Rock, paper, scissors, with the webcam
Build a complete camera script that collects three classes, trains once, recognises several moves, and applies game rules.
Plan one complete run
This lesson builds a rock, paper, scissors game that reads your hand through the webcam. You need a working webcam. The browser asks for camera permission only when your program reaches a capture block, so reading the article never turns the camera on. If your device has no camera, you can still read every stage and build the game rules, which do not use one.
Each Run starts a new camera session. The photos and the trained model from the previous run are gone. So the program that matters has collection, training, and prediction all in one stack. The stages below build it one piece at a time. You can run any of them to practise, but collecting ten photos, adding a block, and running again starts the collection from the beginning.
Collect three labelled groups
The first four stages create the model and collect ten photos for each move. Decide what each move looks like before taking a single photo, and keep to it for every photo.
Build with me · 1
Create an image model before collecting
An image model starts with a name and no examples from your task. It works like the Level 2 image lab: a pretrained network, trained earlier on many other photos, turns each photo into numerical features, and your labelled photos teach a small classifier on top of them. Creating hands is not yet training it on your gestures.
Choose three definitions: closed fist is rock, open palm is paper, and two extended fingers is scissors. Consistent labels matter more than the particular word chosen for the model variable.
Add make an image model called hands from Images. Put a simple status message below it.
What to look for
The status line prints. There is no trained rock-paper-scissors prediction yet.
Make it yours
Explain the difference between the model name hands and its three class names. Renaming hands does not add a class.
Build with me · 2
Collect one labelled group
The capture statement asks for ten photos under one label. All ten must show rock according to your definition. Vary pose and distance a little while keeping the class correct. Ten adjacent nearly identical frames are less varied evidence than ten meaningfully different views.
This partial stage is for learning collection. It is not yet a two-or-more-class training task. Starting another Run clears this temporary collection, so the final program must include every group before training.
Connect add 10 webcam photos to hands labeled rock. Run only if you want to practise the capture interface; allow the camera when your browser requests it.
What to look for
Capture opens for rock, and the inspection reports the collected group. No classification claim is made.
Make it yours
Try three distances for the same fist. Keep faces and unrelated personal material out of your chosen scene if they are unnecessary for the task.
Build with me · 3
Give the model a contrasting group
The second capture group uses the same camera and label rules, with an open palm for paper. Keep backgrounds and lighting shared across classes. If every rock photo has one background and every paper photo another, background becomes an easy shortcut.
Both groups belong to hands. A second make-image-model statement here would reset the collection instead of extending it.
Add the paper capture below rock without another make-image-model block. Inspect the two capture labels before running.
What to look for
The program asks for rock, then paper, and reports both groups.
Make it yours
Choose one background used for both classes. Predict why changing only the gesture is a cleaner first experiment than changing gesture and room together.
Build with me · 4
Complete and inspect the training collection
Scissors completes the label set the game rules will expect. The collection now describes a three-class recognition task. Inspect it before training; a confidently wrong label on a training photo can teach the opposite of what you intended.
Count is one check and content is another. Equal numbers do not guarantee useful diversity or correct labels. Make a note of which hand, angles, lighting, and backgrounds you used so later tests can explore something new.
Add scissors as the third capture group. Keep show photos after all three groups. Read the label shown by each capture prompt.
What to look for
Three labelled groups exist, with ten photos in each if each requested capture completed.
Make it yours
Plan a new pose or a partner's hand for later testing. Do not quietly include that exact test photo in this training collection.
Ten photos per class is a small demonstration set: enough to see the idea work, not enough for reliable recognition. Between photos, move your hand and change its distance and angle. Use the same backgrounds and lighting for all three classes, so that nothing except the hand shape tells them apart.
Train, then recognise one move
With all three groups collected, the model can be trained and then asked about a new photo.
Build with me · 5
Train after all classes are present
Train image model hands trains the classifier for your three classes, using the collected photos. The work happens in your browser. Initial model resources may need to load first, and training time depends on the device. Wait for status updates instead of repeatedly pressing Run.
Prediction comes next. Training sees the labels and adjusts the classifier. Prediction receives a new photo with no label and asks which class gets the strongest support.
Place train image model hands below the inspection. Run the whole collection-and-training stack when you are ready to capture all three classes.
What to look for
After capture, the browser reports training progress and completion. No future photo has been classified yet.
Make it yours
Explain why placing training before the scissors capture would train on an incomplete version of the task.
Build with me · 6
Store one fresh image before asking about it
Capture one test photo and store it as photo. Both image-label and confidence read that same value. If each socket contained a separate camera block, you could classify one gesture and display the score for another.
The result is measured, so this article cannot promise a particular number. Compare the label with the gesture you actually showed. A high score is evidence about this model's preferences, not a correctness guarantee.
Below training, set photo to a webcam photo. Set move and sure using the rounded photo variable. Put labelled output statements after those assignments.
What to look for
One prediction photo is requested. Output shows one of the trained labels and its score on that same image.
Make it yours
Use a new pose. Record intended label, predicted label, and score together instead of saving only successful results.
When you run this stage, complete each capture group, wait for training to finish, then take the prediction photo. If the camera does not start, check the browser's camera permission and any privacy cover over the lens. If training is still loading, wait for the status to change rather than pressing Run again. Pressing Stop means the next Run starts collecting again.
The question recogniser needed a confidence gate, and so does this one.
Build with me · 7
Write a response for an uncertain result
The outer condition decides whether to use the recognition result in the game. A score below 60 makes the program say it is unsure. The other branch prints the label and score. This placement stops the game from treating an unsure label as a valid move.
An unfamiliar elbow or empty frame may still get a high score for one of the three classes. A threshold is not a reliable neither detector, so evaluate the rule instead of assuming it works on everything.
Put both accepted-result output blocks inside otherwise. Put only the unsure message in the below-60 branch.
What to look for
Exactly one branch runs for the captured image. Which branch appears depends on the measured score.
Make it yours
Try several captures of real moves and several of things that are not moves. Count the wrong answers the gate accepted and the right answers it held back before changing the threshold.
The game rules
The game itself needs rules for who wins. Build and test them with moves you type in, before any camera is involved, so that a recognition mistake cannot hide a wiring mistake.
Build with me · 8
Test game rules without a camera
The game is a separate system from the recogniser. Start with known moves so you can check the if/else logic without collecting photos. First handle equality as a draw, then the three combinations that make the player win. All other valid pairs make the computer win.
This is why each winning test uses AND: both the player's move and the opponent's move must match the specific pair. Checking rock alone would wrongly award a win against paper.
Build the nested decision using fixed move and computer assignments. Load this stage to isolate the rules, then inspect each condition and branch.
What to look for
Rock against scissors prints You win.
Make it yours
Try all nine valid pairs: three draws, three wins, and three losses. Keep a small table of expected and actual outcomes.
Build with me · 9
Pick a valid computer move from a list
The list defines the allowed opponent choices. A random whole number from one to three supplies a valid item position for this block interface. Randomness changes the opponent, while the game rules always give the same result for the same pair of moves.
The list order is rock, paper, scissors. If you add a fourth class later, changing the random bound alone will not make the existing rules valid; the game definition must change too.
Create moves and append the three labels. Nest a random whole number in item-of-moves, then store that item under computer before the verdict.
What to look for
The opponent is one valid label. The verdict matches that label against the fixed rock move; it can differ between runs.
Make it yours
For each observed opponent, predict the verdict before reading it. Repeated identical opponent choices can occur by chance.
Put the game together
The last stage joins the recogniser, the gate, and the rules into one program that trains once and plays several rounds.
Build with me · 10
Train once and play several rounds
The final stack keeps collection and training above the round loop. Each of five turns captures one new photo, reads its prediction and score, applies the gate, and plays a round only if the gate accepts the move. The learned image model is reused throughout that one run.
An unsure answer uses up an attempt in this version; the program does not retry forever until it is confident. That explicit rule prevents an unclear gesture from trapping the program in an endless loop. A fresh Run restarts collection, while another loop turn reuses the same trained model.
Build the final loop below training and the move list. Place photo capture, label, score, and gate inside repeat. Keep Five attempts complete below the loop.
What to look for
Training occurs once, then up to five gesture results are processed. The ending message prints once after five attempts.
Make it yours
Choose three attempts while testing, then five. Evaluate recogniser mistakes separately from rule mistakes so you know which part needs improvement.
Test the recogniser on its own
You checked the rules with fixed moves. Now test the recogniser with new poses and, if you can, a partner's hand. For each attempt, write down the move you made, the move the model recognised, its score, and whether the gate said unsure. A game can apply its rules perfectly to a misrecognised move and still feel wrong to the player.
Use what these tests show to improve the photos you collect, and keep a separate set of test hands for the final check, as you did in Level 2.
Full reference solution
This is the complete worked program. Try building it yourself first, then use this reference to find the first place your version behaves differently. The Python below is generated from these exact blocks; helper functions are included so its behaviour can be inspected.
# ---------------------------------------------------------------# This file uses image blocks, which need a camera and a browser.# That part runs here rather than anywhere Python runs.# Everything else below is ordinary Python.# --------------------------------------------------------------- # ---------------------------------------------------------------# Building blocks, written out in plain Python.# This part is generated for you. Your script starts further down.# --------------------------------------------------------------- import random # These blocks need a camera and a browser, so this part runs here rather# than anywhere Python runs. Everything above this line is ordinary Python.from js_images import pagefrom pyodide.ffi import run_sync class ImageModel: """A model that tells photos apart. It does not look at the pixels itself. A pretrained network turns each photo into a list of numbers, and a small model on top learns which lists go with which label.""" def __init__(self, name="model"): self.name = name self.labels = [] self.trained = False def capture(self, count, label): """Opens the camera panel and waits while you take the photos.""" taken = run_sync(page.capture(self.name, int(count), str(label))) if str(label) not in self.labels: self.labels.append(str(label)) print("Took " + str(taken) + " photos of '" + str(label) + "'.") def train(self): if len(self.labels) < 2: raise ValueError( "A model needs at least two labels to tell anything apart. " "Add photos under a second label before training." ) run_sync(page.train(self.name)) self.trained = True print("Trained " + self.name + " on " + str(len(self.labels)) + " labels.") def _ready(self): if not self.trained: raise ValueError( "This model has not been trained yet. Add photos and train it first." ) def classify(self, photo): """What the model thinks is in the photo.""" self._ready() return run_sync(page.classify(self.name, photo)) def confidence(self, photo): """Largest class score out of 100. Unfamiliar photos can still receive a high score; measure correctness separately on labelled examples.""" self._ready() return run_sync(page.confidence(self.name, photo)) def show(self): run_sync(page.show(self.name)) def new_image_model(name="model"): return ImageModel(name) def capture_one(): """Opens the camera panel for one photo and hands it back.""" return run_sync(page.capture_one()) def load_image_file(path, name="model"): """Restore local learned weights without training again.""" with open(str(path), "r", encoding="utf-8") as source: contents = source.read() model = ImageModel(name) model.labels = run_sync(page.load_image_file(name, contents)) model.trained = True print("Loaded " + name + " with " + str(len(model.labels)) + " labels.") return model def save_image_file(model, path="image-model.codeguide.json"): """Save the fitted model, separately from the program and training photos.""" model._ready() contents = run_sync(page.save_image_file(model.name)) with open(str(path), "w", encoding="utf-8") as target: target.write(contents) print("Saved " + str(path)) def load_teachable_machine(url, name="model"): """Loads a model you published from Teachable Machine. The link is the one under Export Model, Upload my model.""" model = ImageModel(name) model.labels = run_sync(page.load_teachable_machine(name, str(url))) model.trained = True print("Loaded " + name + " with " + str(len(model.labels)) + " labels.") return model # ---------------------------------------------------------------# Your script# --------------------------------------------------------------- hands = new_image_model("hands")hands.capture(10, "rock")hands.capture(10, "paper")hands.capture(10, "scissors")hands.show()hands.train()moves = []moves.append("rock")moves.append("paper")moves.append("scissors")for _ in range(int(5)): photo = capture_one() move = hands.classify(photo) sure = hands.confidence(photo) if (sure < 60): print("I am not sure which move you made.") else: computer = moves[int(random.randint(1, 3)) - 1] print("Recognised", move) print("Score", sure) print("Computer", computer) if (move == computer): print("Draw") else: if ((move == "rock") and (computer == "scissors")): print("You win") else: if ((move == "paper") and (computer == "rock")): print("You win") else: if ((move == "scissors") and (computer == "paper")): print("You win") else: print("Computer wins")print("Five attempts complete")Compare this with your version. Different names and personal choices are fine when the program follows the same logic.
Keep your progress
Sign in and every reading, quiz, and exercise you finish is saved.