Physical Intelligence · Hand Pose

Where every finger was, even when the camera could not see it.

Egocentric capture paired with wrist cameras and marker gloves, so the delivered hand pose is measured rather than estimated — including the keypoints the head camera loses behind an object or the other hand.

What it is

What gets captured, and why it is worth having.

Environment
Workshop bench, window light
Captured with
Marker gloves, two wrist cameras, head-mounted RGB
Task
Thread a nut and tighten a bracket
Workshop bench, window light · Marker gloves, two wrist cameras, head-mounted RGB

Monocular and depth rigs estimate hand pose; this one records it. Thin marker gloves and two wrist-mounted cameras see the fingers from below and behind, so a grasp inside a drawer or a hand wrapped around a jar still has 21 keypoints per frame, with a confidence value where a marker was briefly lost.

The same rooms, tasks and people as the egocentric line, with the dexterous work pulled forward: fastening, peeling, threading, pouring — the tasks where a policy fails because the hand it learned from was half hidden.

What it trains
  • Dexterous manipulation policies
  • Hand-pose estimators and their benchmarks
  • Hand–object contact and affordance models
  • Retargeting from human hand to robot hand
The spec

How it is captured, and what you receive.

How it is captured
Primary rig
Marker gloves, 21 keypoints per hand, both hands
Wrist cameras
Two cameras looking at the palm and fingers
Head camera
Head- or chest-mounted RGB, or a depth rig from the depth line
Inertial gloves
Optional, for capture without line of sight
Calibration
Per-session camera and glove calibration files
Sync
Head, wrist and glove streams on one clock
What you receive
Video
H.264 in MP4, head camera plus two wrist streams
Resolution
1080p head, 720p wrist, 4K on request
Frame rate
30 fps, 60 fps on request
Hand pose
21 keypoints per hand per frame, 3D, with confidence
Calibration
Per-session camera and glove calibration files
Episode
1–8 minutes, one task each
Environments and tasks

Where it is recorded, and what people are asked to do.

Tasks are recorded where they actually happen — not in a studio dressed to look like a kitchen. Hover a room to see the kind of setting the protocol calls for.

Kitchens

Indian, European and Brazilian home kitchens, lived-in and cluttered, plus small commercial kitchens and pantries.

Cooking preparationWashing upCupboards and drawersAppliancesPouring and scooping

Living rooms and bedrooms

Real Western homes in Spain, Germany, France and Italy, and apartments in India, Brazil, Japan and Korea. Sofas, wardrobes, shelves, laundry.

Tidying and foldingMaking bedsShelving and storageLaundryFurniture interaction

Offices

Open-plan and small offices: desks, pantries, meeting rooms, filing and cable clutter.

Desk tasksDocument handlingCables and peripheralsPantry tasksDrawers and storage

Retail and small shops

Kirana stores, groceries, pharmacies and supermarket aisles, with real stock and real customers out of focus.

Shelf restockingPicking to a listBagging and countersInventory checksCrates and boxes

Warehouses and light industrial

Racking aisles, packing benches, courier depots and small production lines with pallets, trolleys and parts bins.

Pallet to shelfParcel sortingPacking and tapingParts movementInspection

Workshops and labs

Mechanical and electronics workshops, repair benches and robotics labs where a follower rig can be set up.

Tool handlingAssembly and fasteningPegboards and racksBench tidyingTeleoperation setups

Agriculture and outdoor

Polytunnels, packing sheds, market stalls and yards, in daylight that changes from one episode to the next.

Sorting produceCrates and traysPruning and pickingWeighing and baggingLoading

Client-defined and confidential

Your site, or a space recreated to your drawings, run by an NDA'd crew with limited-visibility workflows and no clip leaving your region.

Proprietary workflowsRecreated spacesStealth programsOn-site with your rigYour robot, our operators
What ships

Deliverables, and the labels you can add.

Deliverables
  • Head and wrist video per episode
  • 3D hand keypoints per frame, both hands, with confidence
  • Task and environment metadata per episode
  • QC status and reviewer chain per episode
  • Consent reference on every file
  • Quality report per batch: accepted, rejected, and why
  • Weekly drops to your bucket, region pinned at kickoff
Annotation add-ons
  • Object pose alongside hand pose
  • Contact events per finger
  • Task and sub-task segmentation
  • Captions per segment

Labelled on the same bench as our speech work, to a written guide, behind the same calibration test.

How a pilot runs

Ten to thirty episodes, then the volume.

  1. 01day 0
    Brief

    The task family, the environments, the rig, the sensors, the episode count. One page, agreed before anything is bought or anyone is recruited.

  2. 02week 1
    SOP and kit

    A written capture protocol — framing, lighting, where a task starts and ends, what counts as a failed take — and the kit list. Operators are trained and tested on it before a single episode is paid for.

  3. 03week 2
    Pilot episodes

    A small batch, ten to thirty episodes, delivered with QC status and a quality report. This is where the protocol breaks — on purpose, and cheaply.

  4. 04ongoing
    Review and scale

    You review the pilot, we fix the protocol, then the run scales on a weekly delivery cadence with the same report attached to every batch.

Turnaround
48 hrs
Written spec and quote
5 days
Sample episode in your format
2 weeks
Pilot batch with its QA report
Weekly
Delivery cadence once the pilot clears
Before an episode counts

What QC checks, every time.

  • Framing

    The task stays in frame; hands and workspace inside the protocol's bounds for the whole episode.

  • Exposure and focus

    No blown highlights or focus hunting beyond the tolerance set for the room.

  • Motion

    No camera whip, no dropped frames; IMU continuous where the rig has one.

  • Sync

    Every stream aligned within one frame; timestamps monotonic across cameras, state and action logs.

  • Task compliance

    Start and end states match the SOP; no skipped or reordered steps unless the protocol allows it.

  • Duration

    Inside the episode window for the task family; overruns are trimmed to the protocol, not the clip.

  • Metadata

    Environment, task, rig, operator id, session and consent reference present and valid on every episode.

  • Privacy

    No bystander faces, screens or documents without consent; blurred where the protocol flags it.

  • Calibration

    Intrinsics and extrinsics on file per rig and session for depth, stereo and multi-camera capture.

  • Labels

    Where annotation is ordered, agreement against the gold set above the threshold before the batch ships.