L1
Atomic action
~seconds · e.g. Grasp the carafe handle
A bounded primitive: reach, grasp, pour, place. Semantic boundaries define the action — duration is guidance, not a cut.
Methodology
To map human motion at meaningful scale and transform it into motion intelligence that advances robotics and Physical AI for the benefit of people and society. Motion Atlas captures human activity as synchronized CameraStreams, grounds them to Master Episode Time, and attaches structure that a robot-learning stack can consume. Capture and Episode operations live in Studio — a separate application. Nexus and Dataset Explorer are how the public inspects the result. This page is the public method, not internal infrastructure.
Episode model
Episode → CameraStream → MediaAsset → Master Time. Multiple synchronized viewpoints belong to that Episode. Original media remains protected. Derived media remains traceable. Camera count is per Episode. Nexus reads this chain; it does not author it. This site plays PROXY only.
01
Capture
Real-world human activity, as observed
02
Synchronize
One Master Time across views and annotations
03
Structure
Episode, task hierarchy, actions, objects, contacts
04
Understand
Observable execution context on the same clock
05
Validate
Quality, provenance, protected originals
06
Deliver
Motion intelligence for robotics and Physical AI
Capture domains
Current domains: HOU Household, HND Handyman / Maintenance, RET Grocery / Retail, CAR Care & Clinical Support. Additional domains can be added without redesigning the contract.
Task decomposition
Actions compose skills; skills compose activities. A long-horizon Episode can remain intact. Useful components are referenced through Master Time rather than by cutting the source into disconnected clips. Semantic task boundaries matter more than arbitrary clip length.
L1
~seconds · e.g. Grasp the carafe handle
A bounded primitive: reach, grasp, pour, place. Semantic boundaries define the action — duration is guidance, not a cut.
L2
~30–90 seconds · e.g. Pour a beverage
A coordinated sequence of actions that accomplishes a recognizable objective. L1 atomics nest under it.
L3
~3–15 minutes · e.g. Prepare breakfast
A longer task composed of interconnected skills and actions. The parent Episode can remain intact while L2/L1 components are referenced on Master Time. Duration is guidance, not a cut.
Demonstration tree
DemonstrationPrepare and serve a hot beverage
Illustrative hierarchy snapshotted from the demonstration Episode. Not a live production EDL.
L2 Retrieve drinking vessel
L2 Pour beverage
L2 Serve beverage
Taxonomy
Motion Atlas does not merely collect clips. Video is the observational foundation. The Episode organizes physical activity into a consistent, versionable hierarchy a learning stack can use.
ECoT
ECoT is structured, observable task-execution context: observation, goal, action plan, decision rule, and success check. It is not private reasoning or hidden thought.
Inside an episode
DemonstrationVideo is the observational foundation. The same Master Episode Time binds action labels, hand keypoints, contacts, object state, ECoT, failure markers, and QA. Not every Episode includes every optional layer. Inspect one layer at a time.
Quality gate
DemonstrationScale alone is not sufficient. Quality should be measurable. Provenance should be traceable. Originals should remain protected. Capture quality and annotation integrity are checked before a clip is training-eligible. This report is a demonstration of the gate, not a production certificate.
Dataset status
Training eligible
Demonstration slice. Production eligibility is per-episode.