Not a marketplace routing your data to strangers. A trained, managed workforce with a published rule set, run by the engineers who also build the models.
A managed network held to a written standard, with per-annotator scoring and a three-violation disqualification threshold.
Segmented into atomic-action clips and described hand by hand — on a live program against a 300,000-hour corpus.
Interiors, engineering drawings, agriculture — polygons, masks, boxes, and keypoints, delivered in your schema.
Not just labels: segmentation and detection models, drone-to-orthomosaic pipelines, deployed inference.
A single minute of washing dishes holds twenty distinct actions, performed by two hands doing different things, under a camera that moves with the operator's head. We cut it to frame-accurate boundaries and describe each hand independently — object-grounded, with camera motion and scene as structured attributes.
How the pipeline works →Two axes — boundaries and language — each reviewed twice. An attributed internal pass that coaches the annotator, then an independent external pass that gates the batch. Work only moves forward on approval.
Internal rejection returns the clip with feedback and coaching. An external rejection sends a large chunk of the batch back for re-work — not just the failing clip.
Annotators work in a purpose-built timeline tool: scrub the source, set boundaries frame by frame, loop the segment to confirm the cut. Boundary accuracy at 0.5-second resolution cannot be eyeballed at 1×.
A bad boundary silently corrupts the description written on top of it. Every boundary is re-checked twice before any description work begins — an attributed internal pass that coaches, then a blind external pass that judges the batch, not the person.
External rejection is deliberately expensive: a large chunk of the batch goes back for re-work. That cost is what keeps internal review honest.
"Washes the bowl" is not sufficient supervision. "Washes the blue and white ceramic bowl with a blue sponge near the faucet" names the target with disambiguating attributes, names the tool, and locates the action — what a VLA policy must resolve at inference time.
One hand stabilises while the other acts — collapsing both into one sentence destroys the structure a bimanual policy needs.
Lets training weight or filter clips where apparent object motion is really head motion.
Sink, counter, shelf — a usable conditioning signal, not an unnormalised string.
Where the reviewer rewrites a description, the record keeps both — the corrected text and the original beneath it. Preserved originals turn the review queue into a labelled corpus of annotator error: per-annotator quality scoring, targeted retraining, guideline refinement. Then the independent external team audits the batch before delivery.
Annotators are held to it directly — three violations disqualifies an annotator from the queue.
Every approved clip is emitted as a structured record — frame-accurate timecodes, hand enum, two independent descriptions, camera motion, scene, and review state with full correction history.
Left hand: The left hand holds the blue and white ceramic bowl in the stainless steel sink.
Right hand: The right hand washes the blue and white ceramic bowl with a blue sponge near the faucet.





Labeled clips, a QA report with error taxonomy, and the questions we'd ask before scaling. Your taxonomy or ours.