Folding, tidying and object handling on sofas, tables and floors, where the working surface moves with the operator.
Folding a garment in the lap, stack resting on the sofa armFolding laundry on a sofaReaching across a table to retrieve an objectHandling a bowl at floor level
01
What this environment asks for
Living spaces are where deformable-object manipulation actually happens. Garments change shape continuously as they are handled, the working surface is often the operator's own lap, and both hands are almost always in play.
The same rule set applies as in the kitchen, but the segmentation problem is harder. Folding has no crisp start and end the way pouring does, so boundaries are set at the pauses between folds rather than at contact events.
02
How the work runs
Segmented folding sequences at the natural pause between folds rather than mid-motion.
Described the garment by type and colour so the target is resolvable at inference time.
Separated the stabilising hand from the acting hand in every bimanual clip.
Handled lap-level work where the surface moves with the operator.
Recorded large camera motion where the operator turned between the folding surface and the stack.
Reviewed boundaries twice before any description work began.
03
What makes it hard
Garments deform continuously, so the object looks different at each boundary frame.
Folding is repetitive with no inherent action boundary.
Lap-level work means the working surface moves with the head camera.
Both hands frequently grip the same object for different purposes.
Low-contrast fabrics on similar-coloured furniture blur the object edge.
04
Result
Atomic-action coverage of deformable-object handling in real living spaces, with per-hand roles preserved rather than collapsed into a single description.
05
What comes out
Structured records with frame-accurate boundaries, independent left and right hand descriptions, camera motion and scene attributes, locked after sign-off.