A human demonstration can show what a task looks like without telling a robot exactly how to perform it. That makes video useful, but the useful information must be identified. A clip of someone arranging cartons can expose object relationships, action order and an outcome while leaving robot-specific control entirely unspecified.
What the recording can support
Annotations can connect an instruction to a scene, identify the object being manipulated and describe the change after each action. Timed steps make it easier to locate the evidence for a label. They also expose omissions: an object may leave the camera view, a hand may hide contact, or the final state may not establish success.
- Scene descriptions explain where the task takes place.
- Object labels distinguish materials, containers and relevant landmarks.
- Action segments describe an observed sequence.
- Outcome labels describe what should be true afterwards.
- Instructions connect language to the intended task.
These labels can support task-understanding experiments and research into representations learned from human activity. They do not guarantee that a particular model will improve, and their suitability depends on collection quality, rights and the training method.
What a phone clip does not contain
Ordinary human video does not supply synchronised robot joint positions, actuator commands, calibrated robot observations or contact-force measurements. A human hand and a robot gripper also have different geometry and capabilities. Copying an observed movement is therefore not the same as producing an executable control trajectory.
The EgoVLA research project illustrates this distinction through an approach combining human egocentric data with robot demonstrations and retargeting. Human recordings complement robot data within a learning method; they are not presented here as a direct replacement for the measurements a robot needs.
Build a traceable handoff
In the 9jaBots task studio, an annotation export contains instructions, timed actions, objects, outcomes and declared provenance. The original video stays separate. Keep that source recording and its permission documentation alongside the exported JSON.
Review labels against the actual clip and distinguish self-review from independent review. Do not relabel an authored template as a collected demonstration merely because its steps sound plausible. The source type is part of the evidence another team needs to judge the data.
Evaluate the layer you have
Begin by testing whether a system selects the intended action given a task description. The task benchmark supports that narrow comparison. The grid lab separately checks navigation under explicit obstacles. Neither currently measures visual understanding of uploaded video.
When hardware becomes available, add synchronised observations and actions under a documented collection protocol. Preserve links between the task, recording, robot configuration and outcome so that claims about transfer can be tested rather than inferred from the existence of a video.
Bring people into the project
Use 9jaTesters to request collection, annotation, testing or human review. Describe the task, volume, location or languages, and acceptance criteria in your brief. For contributor opportunities, join the 9jaTesters workforce.