A kitchen and a warehouse ask completely different things of a policy. We organise first-person capture by environment, because that is what decides the action vocabulary, the object density and where the annotation actually gets hard.
Continuous video is cut into atomic-action clips against a published rule set, described independently for each hand, then gated by two segmentation reviews and two description reviews. Records from a kitchen and a warehouse share a schema, so they train together.
Kitchens, living spaces and personal care. The densest source of bimanual manipulation we capture.

Cooking, washing up and food preparation in real home kitchens, cut into atomic actions and described hand by hand.

Folding, tidying and object handling on sofas, tables and floors, where the working surface moves with the operator.

Grooming and bathroom routines, where tools are small, actions are short and the mirror doubles every object in frame.
Shelf work in live store aisles, where dozens of near-identical products make object reference the hard part.
Line cooking and service at tempo, where the same action repeats hundreds of times a shift.
Assembly, materials handling and warehouse work, with narrow action sets and tight tolerances.

Outdoor sorting and stacking of bulk materials, where objects are irregular, heavy and handled with both hands.
Assembly, fastening and inspection at a fixed station, where the action set is narrow and tolerances are tight.
Picking, scanning, packing and labelling across a warehouse floor, with long traverses between short bursts of manipulation.
The rule set, review model and record schema are already fixed. Adding an environment is a capture and vocabulary problem, not a methodology one, so a new domain starts against a known standard.