Methodology

How human motion becomes an Episode.

To map human motion at meaningful scale and transform it into motion intelligence that advances robotics and Physical AI for the benefit of people and society. Motion Atlas captures human activity as synchronized CameraStreams, grounds them to Master Episode Time, and attaches structure that a robot-learning stack can consume. Capture and Episode operations live in Studio — a separate application. Nexus and Dataset Explorer are how the public inspects the result. This page is the public method, not internal infrastructure.

Episode model

One physical event becomes one Episode.

Episode → CameraStream → MediaAsset → Master Time. Multiple synchronized viewpoints belong to that Episode. Original media remains protected. Derived media remains traceable. Camera count is per Episode. Nexus reads this chain; it does not author it. This site plays PROXY only.

  1. 01

    Capture

    Real-world human activity, as observed

  2. 02

    Synchronize

    One Master Time across views and annotations

  3. 03

    Structure

    Episode, task hierarchy, actions, objects, contacts

  4. 04

    Understand

    Observable execution context on the same clock

  5. 05

    Validate

    Quality, provenance, protected originals

  6. 06

    Deliver

    Motion intelligence for robotics and Physical AI

Capture domains

Launch set, not a closed taxonomy.

Current domains: HOU Household, HND Handyman / Maintenance, RET Grocery / Retail, CAR Care & Clinical Support. Additional domains can be added without redesigning the contract.

Task decomposition

L1, L2, and L3 share one parent clock.

Actions compose skills; skills compose activities. A long-horizon Episode can remain intact. Useful components are referenced through Master Time rather than by cutting the source into disconnected clips. Semantic task boundaries matter more than arbitrary clip length.

L1

Atomic action

~seconds · e.g. Grasp the carafe handle

A bounded primitive: reach, grasp, pour, place. Semantic boundaries define the action — duration is guidance, not a cut.

L2

Skill

~30–90 seconds · e.g. Pour a beverage

A coordinated sequence of actions that accomplishes a recognizable objective. L1 atomics nest under it.

L3

Activity

~3–15 minutes · e.g. Prepare breakfast

A longer task composed of interconnected skills and actions. The parent Episode can remain intact while L2/L1 components are referenced on Master Time. Duration is guidance, not a cut.

Demonstration tree

Demonstration

Prepare and serve a hot beverage

Illustrative hierarchy snapshotted from the demonstration Episode. Not a live production EDL.

  1. L2 Retrieve drinking vessel

    • L1 reach.cup
    • L1 grasp.cup
    • L1 lift.cup
    • L1 place.cup
  2. L2 Pour beverage

    • L1 reach.carafe
    • L1 grasp.carafe
    • L1 pour.liquid
    • L1 place.carafe
    • L1 settle.mug
  3. L2 Serve beverage

    • L1 grasp.mug
    • L1 transport.mug
    • L1 place.mug
    • L1 release.mug

Taxonomy

A structured language for physical activity.

Motion Atlas does not merely collect clips. Video is the observational foundation. The Episode organizes physical activity into a consistent, versionable hierarchy a learning stack can use.

  1. DomainHousehold↓
  2. WorkflowKitchen cleanup↓
  3. TaskClear dining table↓
  4. SkillTransport dishes↓
  5. AtomicReach → Grasp → Lift → Transport → Place → Release↓
  6. ContactRight hand → Plate↓
  7. StatePlate.table → Plate.hand → Plate.sink↓
  8. OutcomeSuccess / failure / recovery

ECoT

Observable task-execution context.

ECoT is structured, observable task-execution context: observation, goal, action plan, decision rule, and success check. It is not private reasoning or hidden thought.

Inside an episode

Demonstration

Video is the first layer. Not the product.

Video is the observational foundation. The same Master Episode Time binds action labels, hand keypoints, contacts, object state, ECoT, failure markers, and QA. Not every Episode includes every optional layer. Inspect one layer at a time.

  1. 01VIDEO
  2. 02TIMESTAMPS
  3. 03TASK HIERARCHY
  4. 04ACTION SEGMENTS
  5. 05HAND / BODY POSE
  6. 06OBJECTS
  7. 07CONTACT EVENTS
  8. 08OBJECT STATE
  9. 09FAILURE / RECOVERY
  10. 10ECoT
  11. 11QA
  12. = training data for physical AI

Quality gate

Demonstration

Motion Atlas quality gate.

Scale alone is not sufficient. Quality should be measurable. Provenance should be traceable. Originals should remain protected. Capture quality and annotation integrity are checked before a clip is training-eligible. This report is a demonstration of the gate, not a production certificate.

  • Temporal synchronizationPASS
  • Camera calibrationPASS
  • Frame integrityPASS
  • ExposurePASS
  • Motion blurPASS
  • Hand visibilityPASS
  • Critical object visibilityPASS
  • Pose confidencePASS
  • Annotation completenessPASS
  • Privacy reviewPASS
  • Sequence continuityPASS

Dataset status

Training eligible

Demonstration slice. Production eligibility is per-episode.