Preparing your data
What makes labeling fast and accurate: file quality, stable IDs, groups, hints and privacy.
Files
- JPEG or PNG, long edge 1,000 to 4,000 pixels. Larger files only slow experts down; smaller ones hide details.
- One subject per image where possible. If a photo shows two loads, split it before pushing.
- Keep the original orientation. We do not auto-rotate; fix EXIF orientation on your side.
- Send a
checksum. It costs you one hash and protects against corrupted or swapped files.
Stable IDs
Your id is the join key for everything: results, flags, re-labels, invoices. Derive it from something that never changes (a database primary key, a camera event ID), not from a filename that might be renamed. If you have to re-push a corrected file, use a new ID and treat the old one as withdrawn.
Groups
Use group for the dimension you want quality reported on: site, camera, customer, model version. Agreement and flag rates are broken down per group in the dashboard, which is how you spot a camera with bad lighting or a site with unusual material mixes.
Model hints and active learning
Two kinds of hints. Hidden ones: model_label and model_confidence are never shown to experts, so they cannot bias them; we use them to put uncertain items first and to report where humans disagree with the model. Pre-annotations: label (a class name) or regions (boxes {"type":"rect","label":"Can","x":12,"y":30,"w":20,"h":25} or polygons {"type":"polygon","label":"Can","points":[[x,y],...]}, all in percent of the image) are shown to the expert as a suggestion to confirm or correct, which is much faster than drawing from scratch. Once your model is decent, push only items below a confidence threshold; that typically cuts labeling volume by 60 to 80 percent.
Daily batches
Keep one running task per label schema and push each day’s items into it, ideally in a few large requests rather than one request per file. Set captured_at so queue order and the 24-hour turnaround reports are accurate. Use one Idempotency-Key per batch.