A set has 90 mugs and 10 glasses. What does an always-mug baseline score?
Which data should compare two candidate training settings?
You change backgrounds and epochs together, then improve the score. What can you conclude?
Most mistakes are glasses predicted as mugs. What is a targeted next step?
The same image appears in training and final test. Why is this a problem?
What should an experiment log connect?
How should target IDs 18, 3, and 42 appear in a labels export?
What is needed to reuse a trained image model without retraining?