Views
CameraStreams share the Episode clock
Teaching Physical AI to understand how people move, interact and accomplish real-world tasks.
Motion Atlas transforms diverse human activity into structured motion intelligence — capturing how actions unfold across time, viewpoints, objects and environments so intelligent machines can learn from the physical world.
The data problem
Artificial intelligence has learned extensively from the digital world. Physical AI introduces a different challenge. A machine operating in the real world must understand more than what an object is.
The physical world must be experienced, not merely described.
Embodied knowledge
Everyday human activity contains an enormous body of physical knowledge. That knowledge appears in ordinary actions people have learned through experience. Much of it was never formally written down. It exists in how people act.
It exists in motion.
Motion Atlas captures physical experience and transforms it into structured data designed to become useful for machines learning about the physical world.
How Motion Atlas works
Motion Atlas captures real-world human activity as Episodes. An Episode is a coherent representation of a physical activity connected across viewpoints, time, media, task structure, objects, interactions, observable execution, quality, and provenance. Capture, synchronize, structure, understand, validate, and deliver — this is the public capability, not a list of internal applications.
One activity. One Episode. Multiple perspectives. One shared timeline.
01
Capture
Real-world human activity, as observed
02
Synchronize
One Master Time across views and annotations
03
Structure
Episode, task hierarchy, actions, objects, contacts
04
Understand
Observable execution context on the same clock
05
Validate
Quality, provenance, protected originals
06
Deliver
Motion intelligence for robotics and Physical AI
Synchronized observation
One physical event. Multiple synchronized observations of the same Episode. Unavailable views are omitted, not invented. Six is a configuration, not a required maximum.
1-camera
A single CameraStream. The Episode still has Master Time, structure, and ECoT.
3-camera
Three synchronized views of the same Episode.
6-camera
Six synchronized views. Six is a configuration, not a required maximum.
EP000001 · 6-camera configuration
DemonstrationPour water into mug






Switch configurations to inspect 1-camera, 3-camera, and 6-camera Episodes. These are supported layouts — not a claim that cameras 2, 4, or 5 are missing from a six-slot rig.
Master Episode Time
Views, actions, contacts, object-state changes, and ECoT records are aligned to Master Episode Time. Demonstration playback may lockstep videos on that clock — that is playback behavior, not a published sync-precision claim.
Views
CameraStreams share the Episode clock
Structure
Actions and contacts occupy the same range
Context
ECoT is timed to the body, not to a caption file
The Episode
Video records what a camera saw. Physical intelligence requires context. Motion Atlas connects synchronized observations with task structure, time, objects, interactions, observable execution, quality, and provenance. Depending on the collection, an Episode may also incorporate richer motion representations and additional sensor-derived information. Not every Episode contains every optional layer.
The result is not simply a collection of recordings. It is a structured representation of human activity designed to become increasingly useful for Physical AI.
Action
DemonstrationPOUR
Transfer water into mug
1.55–3.55s · BOTH · SUCCESS
Objects · contact
RIGHT HAND → CARAFE HANDLE
ECoT
ECoT is structured, observable task-execution context: observation, goal, action plan, decision rule, and success check. It is not private reasoning or hidden thought.
Origin SYNTHETIC · privacy UNKNOWN
Observable task context (ECoT)
ECoT is structured, observable task-execution context: observation, goal, action plan, decision rule, and success check. It is not private reasoning or hidden thought.
Task hierarchy
A hand reaches. An object is grasped. A tool is positioned. An adjustment is made. Another action follows. A mistake may be corrected. Eventually, an objective is accomplished. Isolated movement clips cannot explain that. L1 atomic action. L2 skill. L3 activity. Semantic task boundaries matter more than arbitrary clip length.
Intelligent machines ultimately need to learn not only how people move, but how movement becomes useful action.
L1
~seconds · e.g. Grasp the carafe handle
A bounded primitive: reach, grasp, pour, place. Semantic boundaries define the action — duration is guidance, not a cut.
L2
~30–90 seconds · e.g. Pour a beverage
A coordinated sequence of actions that accomplishes a recognizable objective. L1 atomics nest under it.
L3
~3–15 minutes · e.g. Prepare breakfast
A longer task composed of interconnected skills and actions. The parent Episode can remain intact while L2/L1 components are referenced on Master Time. Duration is guidance, not a cut.
Demonstration tree
DemonstrationPrepare and serve a hot beverage
Illustrative hierarchy snapshotted from the demonstration Episode. Not a live production EDL.
L2 Retrieve drinking vessel
L2 Pour beverage
L2 Serve beverage
Diversity
The same objective can be accomplished differently depending on the person, the environment, the objects, the tools, the workspace, and the technique. Homes differ. Workplaces differ. Objects differ. Tools differ. Techniques differ. People differ. A machine intended to operate beyond a controlled environment must eventually encounter this variation. Motion Atlas is being designed to capture human activity across different people, environments, objects, tasks and ways of accomplishing them — a collection objective, not a claim that the current catalog already represents the world.
The world is diverse. The data that teaches machines to operate within it should reflect that reality.

HOU · Launch domain
Wash, fold, pour, wipe, make the bed. The chores a home actually asks of a body.
Pour a drink · Fold a towel · Wipe a counter

HND · Launch domain
Drill, hinge, fasten, measure, repair. Tools against the built world — contact-rich and two-handed.
Hang a hinge · Drive a fastener · Measure and mark

RET · Launch domain
Select, bag, carry, pay, put away. Motion that leaves the house and still has to land.
Bag produce · Shelf an item · Carry and set down

CAR · Launch domain
Support activities around care — positioning, bed preparation, equipment handling. Not a claim of regulated clinical-procedure datasets.
Bed preparation · Mobility assistance (simulated) · Care-equipment preparation
Four launch domains in the current program. The taxonomy can expand. They are not four permanent pillars.
Future expansion — not currently captured
Quality
Large amounts of data can be valuable. Volume alone does not make physical data useful. Data becomes harder to trust when its origin is uncertain, its timing is unreliable, its quality is unknown, its transformations cannot be traced, or its relationship to the underlying activity has been lost. Motion Atlas is being designed around synchronization, structured metadata, quality assurance, versioning, traceability, provenance, and preservation of original media.
Quality should be measurable. Provenance should be traceable. Originals should remain protected.
Dataset explorer
DemonstrationThere is a difference between watching human activity and representing it for machine learning. Here: one physical event, multiple synchronized perspectives, one shared timeline, structured task context.
Evidence: playable demonstration Episodes. One physical event. Multiple synchronized perspectives. One shared timeline.
EP000001 · MA-HH-KITCHEN-DEMO-001 · L1 · MA-6 · HOU · SYNTHETIC
VIDEO ✓MULTI-VIEW ✓2D POSE ✓3D POSE —HANDS ✓CONTACT ✓STATE ✓ECoT ✓FAILURE ✓
CONTACT-RICH · BIMANUAL · PRECISION · synced web proxy

CAM01 · FRONT · EXO

CAM02 · RIGHT · EXO

CAM03 · LEFT · EXO

CAM04 · CHEST · EGO

CAM05 · FOREHEAD · EGO

CAM06 · 360° · EXO
Task hierarchy · demonstration tree
DemonstrationL3 · Parent task
Prepare and serve a hot beverage
illustrative L3 · 3–15 min in production
Larger purpose
The practical value of Physical AI will increasingly depend on what that intelligence enables machines to do in the physical world — in manufacturing, logistics, healthcare and care environments, agriculture, construction, maintenance, homes, and other forms of useful physical work. Motion Atlas does not claim to transform those industries. It exists to contribute one part of this larger undertaking: transforming knowledge embedded in human activity into motion intelligence that machines can learn from.
Mission
From human experience, to machine understanding, to useful action in the physical world.
Human Motion. Machine Intelligence.
Motion Atlas Platform
Nexus is this site — the public, research, and customer gateway. Dataset Explorer is where you inspect structured sample Episodes. Studio is a separate application for Episode operations.
nexus
AVAILABLEPublic, research, and customer gateway to Motion Atlas.
You are here
explorer
AVAILABLEInspect structured sample Episodes and Motion Atlas metadata.
studio
PILOTMulti-camera Episode editing and dataset operations.
Studio deployment connection required.
Sample
A representative synchronized multi-view episode, under NDA, with episode metadata, task hierarchy, action segmentation, objects, contacts, state transitions, pose, ECoT, failure/recovery labels, QA report, calibration, and schema documentation — where those layers exist for the slice. The aim is not a highlight reel. It is evidence that human activity can be structured so Physical AI can learn useful work in the world people share.
FAQ
One coherent human activity — one physical event — across cameras, time, and structure. Depending on the collection it may include synchronized multi-view video, task hierarchy, actions, objects, contacts, observable execution, quality, and provenance. Camera count is per Episode. Not every Episode contains every optional layer.
Much of physical knowledge is embodied in how people move rather than written down. Physical AI has to learn how actions unfold in the real world. Motion Atlas captures that activity and structures it as motion intelligence machines can learn from.
A single camera records one observation. Synchronized views of the same Episode show how the same action occupies space and time from more than one vantage — what a learning stack may need when a hand occludes an object, or a tool leaves the frame.
No. Camera count is not a platform constant. MA-6 is full multiview, MA-3 is long-horizon, and single-camera episodes are valid. Unavailable views are marked, not faked.
The demonstration overlay uses 21 image-plane hand keypoints. Metric 3D is only claimed when a release documents frame, units, origin, and method.
The current domain is care and clinical support activities — positioning, bed prep, equipment handling. Not a regulated clinical-procedure dataset.
Task failures and recoveries stay. Capture failures do not ship.
Request a sample under NDA. If the fit is real, we talk license and cut.