Physical Intelligence · Video Annotation

Robot and human video, turned into training data a stranger can check.

Egocentric, teleoperation, robot and policy-rollout video annotated on Foundry: sub-task boundaries, natural-language captions, hand and object masks, success and failure labels — to a written guide, behind a calibration test, with agreement reported per batch.

What it is

What gets captured, and why it is worth having.

Environment
Annotation desk on Foundry, afternoon light
Captured with
Two monitors, timeline tool, written guide open
Task
Mark sub-task boundaries and caption each segment
Annotation desk on Foundry, afternoon light · Two monitors, timeline tool, written guide open

Long-horizon video is where crowd annotation falls apart: a segment boundary is a judgment, and two people make it differently unless the guide says exactly what counts. Ours does. Annotators clear a gold set on your taxonomy before they touch a paid episode, and reviewers sample every batch against the same set.

Model-assisted where it helps — a VLM proposes captions and boundaries, a person accepts, edits or rejects, and the rate of each is reported — so the data is human-verified without paying human time for the easy frames.

What it trains
  • Language-conditioned policies
  • Long-horizon task understanding
  • Policy evaluation and failure analysis
  • Video-language pretraining for robotics
The spec

How it is captured, and what you receive.

How it is captured
Platform
Foundry — timeline, mask and caption tools on one bench
Guide
Your taxonomy and guide, versioned with us
Gold set
Per project, calibration before paid work
Model assist
VLM proposals with accept, edit and reject rates reported
Review
Sampled per batch, adjudicated by a senior reviewer
Inputs
MP4, MOV, image sequences, LeRobot- or RLDS-style episodes
What you receive
Inputs
MP4, MOV, image sequences, LeRobot- or RLDS-style episodes
Segmentation
Task and sub-task boundaries to the frame, nested where the taxonomy is
Captions
Natural language per segment, style guide enforced
Masks
Hand, object and gripper masks per frame or keyframe
Labels
Success and failure per episode, intervention flags, ratings
Output
JSON per episode, COCO-style masks, or your schema
Environments and tasks

Where it is recorded, and what people are asked to do.

Tasks are recorded where they actually happen — not in a studio dressed to look like a kitchen. Hover a room to see the kind of setting the protocol calls for.

Kitchens

Indian, European and Brazilian home kitchens, lived-in and cluttered, plus small commercial kitchens and pantries.

Cooking preparationWashing upCupboards and drawersAppliancesPouring and scooping

Living rooms and bedrooms

Real Western homes in Spain, Germany, France and Italy, and apartments in India, Brazil, Japan and Korea. Sofas, wardrobes, shelves, laundry.

Tidying and foldingMaking bedsShelving and storageLaundryFurniture interaction

Offices

Open-plan and small offices: desks, pantries, meeting rooms, filing and cable clutter.

Desk tasksDocument handlingCables and peripheralsPantry tasksDrawers and storage

Retail and small shops

Kirana stores, groceries, pharmacies and supermarket aisles, with real stock and real customers out of focus.

Shelf restockingPicking to a listBagging and countersInventory checksCrates and boxes

Warehouses and light industrial

Racking aisles, packing benches, courier depots and small production lines with pallets, trolleys and parts bins.

Pallet to shelfParcel sortingPacking and tapingParts movementInspection

Workshops and labs

Mechanical and electronics workshops, repair benches and robotics labs where a follower rig can be set up.

Tool handlingAssembly and fasteningPegboards and racksBench tidyingTeleoperation setups

Agriculture and outdoor

Polytunnels, packing sheds, market stalls and yards, in daylight that changes from one episode to the next.

Sorting produceCrates and traysPruning and pickingWeighing and baggingLoading

Client-defined and confidential

Your site, or a space recreated to your drawings, run by an NDA'd crew with limited-visibility workflows and no clip leaving your region.

Proprietary workflowsRecreated spacesStealth programsOn-site with your rigYour robot, our operators
What ships

Deliverables, and the labels you can add.

Deliverables
  • Labels per episode in your schema
  • Agreement, rework rate and VLM accept rate per batch
  • Task and environment metadata per episode
  • QC status and reviewer chain per episode
  • Consent reference on every file
  • Quality report per batch: accepted, rejected, and why
  • Weekly drops to your bucket, region pinned at kickoff
Annotation add-ons
  • 3D hand pose on annotated frames
  • Object tracking across episodes
  • Language instruction rewriting
  • Adjudicated gold set you keep

Labelled on the same bench as our speech work, to a written guide, behind the same calibration test.

How a pilot runs

Ten to thirty episodes, then the volume.

  1. 01day 0
    Brief

    The task family, the environments, the rig, the sensors, the episode count. One page, agreed before anything is bought or anyone is recruited.

  2. 02week 1
    SOP and kit

    A written capture protocol — framing, lighting, where a task starts and ends, what counts as a failed take — and the kit list. Operators are trained and tested on it before a single episode is paid for.

  3. 03week 2
    Pilot episodes

    A small batch, ten to thirty episodes, delivered with QC status and a quality report. This is where the protocol breaks — on purpose, and cheaply.

  4. 04ongoing
    Review and scale

    You review the pilot, we fix the protocol, then the run scales on a weekly delivery cadence with the same report attached to every batch.

Turnaround
48 hrs
Written spec and quote
5 days
Sample episode in your format
2 weeks
Pilot batch with its QA report
Weekly
Delivery cadence once the pilot clears
Before an episode counts

What QC checks, every time.

  • Framing

    The task stays in frame; hands and workspace inside the protocol's bounds for the whole episode.

  • Exposure and focus

    No blown highlights or focus hunting beyond the tolerance set for the room.

  • Motion

    No camera whip, no dropped frames; IMU continuous where the rig has one.

  • Sync

    Every stream aligned within one frame; timestamps monotonic across cameras, state and action logs.

  • Task compliance

    Start and end states match the SOP; no skipped or reordered steps unless the protocol allows it.

  • Duration

    Inside the episode window for the task family; overruns are trimmed to the protocol, not the clip.

  • Metadata

    Environment, task, rig, operator id, session and consent reference present and valid on every episode.

  • Privacy

    No bystander faces, screens or documents without consent; blurred where the protocol flags it.

  • Calibration

    Intrinsics and extrinsics on file per rig and session for depth, stereo and multi-camera capture.

  • Labels

    Where annotation is ordered, agreement against the gold set above the threshold before the batch ships.