Runs in the browser · no account needed to try
Describe what you want to find.
Get a detector that finds it.
Spotr takes a sentence and a list of classes, generates a synthetic dataset, labels it with bounding boxes, trains a YOLO detector on a free GPU, and hands you a link. Open the link on any phone or laptop and detection runs entirely on-device — your camera frames never leave the browser.
Try a demo
Street
/streetPeople and vehicles, the classic street scene.
- person
- car
- bus
- truck
- bicycle
- motorcycle
Open detector →
Fruit
/fruitBananas, apples and oranges on a kitchen counter.
- banana
- apple
- orange
Open detector →
Pets
/petsDogs, cats and birds around the house.
- dog
- cat
- bird
Open detector →
All three demos run one COCO-pretrained YOLO26n graph, filtered to a focused class list. Models trained by Spotr replace them once the training pipeline lands.
How it works
- 01
Name your classes
Give the model a name and the things it should find — “hard hat, safety vest”.
- 02
Generate or upload
A language model expands your scene into ~40 varied prompts, or bring your own photos.
- 03
Review the labels
Boxes are drawn automatically by an open-vocabulary detector. Delete the bad ones before training.
- 04
Share the link
Training exports to ONNX and you get a six-character URL that runs on any modern browser.
Built to cost nothing to run
Every part of Spotr sits on a free tier, which is a design constraint rather than a footnote. Inference runs on your device instead of a server, so a popular model page costs the same as an unpopular one. Training runs on donated serverless GPU credits with a hard per-job time cap, and falls back to a second free backend when the monthly budget runs out.
- Inference
- On-device, WebGPU with a WASM fallback
- Detector
- Ultralytics YOLO26 nano at 640px
- Training cap
- 10 minutes per model
- Frames uploaded
- None, ever
- Licence
- AGPL-3.0, source on GitHub
- Cost to you
- Free