Motion Atlas — The Global Dataset for Physical AI. Motion Atlas builds structured, synchronized multi-camera human-motion datasets and motion intelligence for robotics and Physical AI.

Teaching Physical AI to understand how people move, interact and accomplish real-world tasks.

Motion Atlas transforms diverse human activity into structured motion intelligence — capturing how actions unfold across time, viewpoints, objects and environments so intelligent machines can learn from the physical world.

The data problem

Intelligence is moving into the physical world.

Artificial intelligence has learned extensively from the digital world. Physical AI introduces a different challenge. A machine operating in the real world must understand more than what an object is.

The physical world must be experienced, not merely described.

  • how objects are approached
  • how they are grasped
  • how they are manipulated
  • how tools are used
  • how two hands coordinate
  • how environments influence movement
  • how actions unfold through time
  • how individual actions combine into meaningful tasks

Embodied knowledge

Human activity contains knowledge.

Everyday human activity contains an enormous body of physical knowledge. That knowledge appears in ordinary actions people have learned through experience. Much of it was never formally written down. It exists in how people act.

It exists in motion.

  • folding a towel
  • organizing a workspace
  • preparing food
  • using a tool
  • handling unfamiliar objects
  • cleaning
  • assembling
  • sorting, carrying, completing the work

Motion Atlas captures physical experience and transforms it into structured data designed to become useful for machines learning about the physical world.

How Motion Atlas works

From human activity to motion intelligence.

Motion Atlas captures real-world human activity as Episodes. An Episode is a coherent representation of a physical activity connected across viewpoints, time, media, task structure, objects, interactions, observable execution, quality, and provenance. Capture, synchronize, structure, understand, validate, and deliver — this is the public capability, not a list of internal applications.

One activity. One Episode. Multiple perspectives. One shared timeline.

  1. 01

    Capture

    Real-world human activity, as observed

  2. 02

    Synchronize

    One Master Time across views and annotations

  3. 03

    Structure

    Episode, task hierarchy, actions, objects, contacts

  4. 04

    Understand

    Observable execution context on the same clock

  5. 05

    Validate

    Quality, provenance, protected originals

  6. 06

    Deliver

    Motion intelligence for robotics and Physical AI

Synchronized observation

Supported configurations: 1-camera, 3-camera, and 6-camera.

One physical event. Multiple synchronized observations of the same Episode. Unavailable views are omitted, not invented. Six is a configuration, not a required maximum.

  • 1-camera

    A single CameraStream. The Episode still has Master Time, structure, and ECoT.

  • 3-camera

    Three synchronized views of the same Episode.

  • 6-camera

    Six synchronized views. Six is a configuration, not a required maximum.

EP000001 · 6-camera configuration

Demonstration

Pour water into mug

CAM01Front
CAM02Right
CAM03Left
CAM04Chest
CAM05Forehead
CAM06360°

Switch configurations to inspect 1-camera, 3-camera, and 6-camera Episodes. These are supported layouts — not a claim that cameras 2, 4, or 5 are missing from a six-slot rig.

Master Episode Time

Every annotation shares one clock.

Views, actions, contacts, object-state changes, and ECoT records are aligned to Master Episode Time. Demonstration playback may lockstep videos on that clock — that is playback behavior, not a published sync-precision claim.

Views

CameraStreams share the Episode clock

Structure

Actions and contacts occupy the same range

Context

ECoT is timed to the body, not to a caption file

The Episode

More than video.

Video records what a camera saw. Physical intelligence requires context. Motion Atlas connects synchronized observations with task structure, time, objects, interactions, observable execution, quality, and provenance. Depending on the collection, an Episode may also incorporate richer motion representations and additional sensor-derived information. Not every Episode contains every optional layer.

The result is not simply a collection of recordings. It is a structured representation of human activity designed to become increasingly useful for Physical AI.

Action

Demonstration

POUR

Transfer water into mug

1.55–3.55s · BOTH · SUCCESS

Objects · contact

  • carafe
  • mug
  • water
  • counter

RIGHT HAND → CARAFE HANDLE

ECoT

ECoT is structured, observable task-execution context: observation, goal, action plan, decision rule, and success check. It is not private reasoning or hidden thought.

Origin SYNTHETIC · privacy UNKNOWN

Observable task context (ECoT)

ECoT is structured, observable task-execution context: observation, goal, action plan, decision rule, and success check. It is not private reasoning or hidden thought.

Observation
What is visible in the workspace
Goal
What this step is for
Action plan
What the body is doing
Decision rule
What would change the next step
Success check
How completion is judged
Narration
Participant speech captured with the Episode, when present
Annotation
Human-authored task context on Master Time

Task hierarchy

From movements to meaningful tasks.

A hand reaches. An object is grasped. A tool is positioned. An adjustment is made. Another action follows. A mistake may be corrected. Eventually, an objective is accomplished. Isolated movement clips cannot explain that. L1 atomic action. L2 skill. L3 activity. Semantic task boundaries matter more than arbitrary clip length.

Intelligent machines ultimately need to learn not only how people move, but how movement becomes useful action.

L1

Atomic action

~seconds · e.g. Grasp the carafe handle

A bounded primitive: reach, grasp, pour, place. Semantic boundaries define the action — duration is guidance, not a cut.

L2

Skill

~30–90 seconds · e.g. Pour a beverage

A coordinated sequence of actions that accomplishes a recognizable objective. L1 atomics nest under it.

L3

Activity

~3–15 minutes · e.g. Prepare breakfast

A longer task composed of interconnected skills and actions. The parent Episode can remain intact while L2/L1 components are referenced on Master Time. Duration is guidance, not a cut.

Demonstration tree

Demonstration

Prepare and serve a hot beverage

Illustrative hierarchy snapshotted from the demonstration Episode. Not a live production EDL.

  1. L2 Retrieve drinking vessel

    • L1 reach.cup
    • L1 grasp.cup
    • L1 lift.cup
    • L1 place.cup
  2. L2 Pour beverage

    • L1 reach.carafe
    • L1 grasp.carafe
    • L1 pour.liquid
    • L1 place.carafe
    • L1 settle.mug
  3. L2 Serve beverage

    • L1 grasp.mug
    • L1 transport.mug
    • L1 place.mug
    • L1 release.mug

Diversity

The physical world has no single way of moving.

The same objective can be accomplished differently depending on the person, the environment, the objects, the tools, the workspace, and the technique. Homes differ. Workplaces differ. Objects differ. Tools differ. Techniques differ. People differ. A machine intended to operate beyond a controlled environment must eventually encounter this variation. Motion Atlas is being designed to capture human activity across different people, environments, objects, tasks and ways of accomplishing them — a collection objective, not a claim that the current catalog already represents the world.

The world is diverse. The data that teaches machines to operate within it should reflect that reality.

HOU · Launch domain

Household

Wash, fold, pour, wipe, make the bed. The chores a home actually asks of a body.

Pour a drink · Fold a towel · Wipe a counter

HND · Launch domain

Handyman / maintenance

Drill, hinge, fasten, measure, repair. Tools against the built world — contact-rich and two-handed.

Hang a hinge · Drive a fastener · Measure and mark

RET · Launch domain

Grocery / retail

Select, bag, carry, pay, put away. Motion that leaves the house and still has to land.

Bag produce · Shelf an item · Carry and set down

CAR · Launch domain

Care & clinical support

Support activities around care — positioning, bed preparation, equipment handling. Not a claim of regulated clinical-procedure datasets.

Bed preparation · Mobility assistance (simulated) · Care-equipment preparation

Four launch domains in the current program. The taxonomy can expand. They are not four permanent pillars.

Future expansion — not currently captured

  • Manufacturing
  • Warehousing
  • Hospitality
  • Construction
  • Agriculture
  • Office

Quality

Scale matters. Trust matters more.

Large amounts of data can be valuable. Volume alone does not make physical data useful. Data becomes harder to trust when its origin is uncertain, its timing is unreliable, its quality is unknown, its transformations cannot be traced, or its relationship to the underlying activity has been lost. Motion Atlas is being designed around synchronization, structured metadata, quality assurance, versioning, traceability, provenance, and preservation of original media.

Quality should be measurable. Provenance should be traceable. Originals should remain protected.

Dataset explorer

Demonstration

See human activity as a machine may need to understand it.

There is a difference between watching human activity and representing it for machine learning. Here: one physical event, multiple synchronized perspectives, one shared timeline, structured task context.

Evidence: playable demonstration Episodes. One physical event. Multiple synchronized perspectives. One shared timeline.

Workspace

EP000001 · MA-HH-KITCHEN-DEMO-001 · L1 · MA-6 · HOU · SYNTHETIC

Pour water into mug

VIDEO ✓MULTI-VIEW ✓2D POSE ✓3D POSE —HANDS ✓CONTACT ✓STATE ✓ECoT ✓FAILURE ✓

CONTACT-RICH · BIMANUAL · PRECISION · synced web proxy

CAM01 · FRONT · EXO

CAM02 · RIGHT · EXO

CAM03 · LEFT · EXO

CAM04 · CHEST · EGO

CAM05 · FOREHEAD · EGO

CAM06 · 360° · EXO

00:00.000 / 00:06.000
Synchronized multi-view
L2
L1
Contact
State
▲▲▲▲▲▲▲▲▲
Fail

Task hierarchy · demonstration tree

Demonstration
Click a timed row to seek

L3 · Parent task

Prepare and serve a hot beverage

illustrative L3 · 3–15 min in production

Larger purpose

Building useful intelligence for the physical world.

The practical value of Physical AI will increasingly depend on what that intelligence enables machines to do in the physical world — in manufacturing, logistics, healthcare and care environments, agriculture, construction, maintenance, homes, and other forms of useful physical work. Motion Atlas does not claim to transform those industries. It exists to contribute one part of this larger undertaking: transforming knowledge embedded in human activity into motion intelligence that machines can learn from.

Mission

To map human motion at meaningful scale and transform it into motion intelligence that advances robotics and Physical AI for the benefit of people and society.

From human experience, to machine understanding, to useful action in the physical world.

Human Motion. Machine Intelligence.

Motion Atlas Platform

One motion intelligence platform. Specialized workspaces.

Nexus is this site — the public, research, and customer gateway. Dataset Explorer is where you inspect structured sample Episodes. Studio is a separate application for Episode operations.

nexus

AVAILABLE

Nexus

Public, research, and customer gateway to Motion Atlas.

You are here

explorer

AVAILABLE

Dataset Explorer

Inspect structured sample Episodes and Motion Atlas metadata.

studio

PILOT

Motion Atlas Studio

Multi-camera Episode editing and dataset operations.

Studio deployment connection required.

Sample

Inspect the data yourself.

A representative synchronized multi-view episode, under NDA, with episode metadata, task hierarchy, action segmentation, objects, contacts, state transitions, pose, ECoT, failure/recovery labels, QA report, calibration, and schema documentation — where those layers exist for the slice. The aim is not a highlight reel. It is evidence that human activity can be structured so Physical AI can learn useful work in the world people share.

FAQ

Straight answers before a license call.

What is an Episode?

One coherent human activity — one physical event — across cameras, time, and structure. Depending on the collection it may include synchronized multi-view video, task hierarchy, actions, objects, contacts, observable execution, quality, and provenance. Camera count is per Episode. Not every Episode contains every optional layer.

Why capture human activity?

Much of physical knowledge is embodied in how people move rather than written down. Physical AI has to learn how actions unfold in the real world. Motion Atlas captures that activity and structures it as motion intelligence machines can learn from.

Why multiple synchronized views?

A single camera records one observation. Synchronized views of the same Episode show how the same action occupies space and time from more than one vantage — what a learning stack may need when a hand occludes an object, or a tool leaves the frame.

Is every sequence six cameras?

No. Camera count is not a platform constant. MA-6 is full multiview, MA-3 is long-horizon, and single-camera episodes are valid. Unavailable views are marked, not faked.

Is pose 3D?

The demonstration overlay uses 21 image-plane hand keypoints. Metric 3D is only claimed when a release documents frame, units, origin, and method.

Do you capture medical procedures?

The current domain is care and clinical support activities — positioning, bed prep, equipment handling. Not a regulated clinical-procedure dataset.

Are failures kept?

Task failures and recoveries stay. Capture failures do not ship.

How do we start?

Request a sample under NDA. If the fit is real, we talk license and cut.