Teleoperation Explained

Teleoperation Explained
A hand closing around a mug in one frame of reference, a gripper closing around it in another, and the split-second of bookkeeping that keeps them the same demonstration

An operator reaches for a mug on a kitchen counter. A two-fingered gripper follows the movement, across the room or over a network connection. The operator controls the action. To train a policy from this episode, the recording must include the joint positions, gripper state and action commands, with timestamps that match them to the camera frames.

Video shows the mug being picked up. A teleoperation episode also records the robot commands that produced the grasp.

Video shows what happened; teleoperation records what was done

Perit’s egocentric video work records people completing tasks from a first-person view. There is no robot state or command stream, but the footage can show what successful execution looks like. UMI gripper capture adds gripper-relative motion without requiring a full robot in the loop.

In teleoperation, a person controls a robot or a robot-equivalent end-effector directly. The recorded actions are commands a policy can learn to predict. Perit delivers these demonstrations episode by episode, with synchronised camera streams.

What one episode actually contains

Each episode contains several streams recorded together:

  • Camera streams: one or more views of the workspace, synchronised to each other and to everything below.
  • State: the robot or end-effector’s own position and configuration at each moment: joint angles, gripper pose, whatever the rig reports about itself.
  • Action: the command sent by the operator’s teleoperation input at that same moment, which is what a policy is ultimately trying to learn to predict.
  • Depth, where the rig supports it: RGB-D, stereo or LiDAR capture, aligned per episode, adding a 3D geometry layer to the 2D camera view.
  • Tactile, where the rig supports it: a contact and pressure stream synchronised alongside the video, for tasks where force matters as much as position.

A camera frame needs a matching action timestamp. A pressure spike needs a matching view of the contact. Without that alignment, the policy cannot reliably connect the command, the state and the result.

The one-frame rule

Perit’s Sync check requires streams to align within one frame, with monotonic timestamps across cameras, state and action logs. Converting that frame into milliseconds makes the tolerance easier to judge.

at 30 fps: 1 frame = 1/30 s ≈ 33.3 ms at 60 fps: 1 frame = 1/60 s ≈ 16.7 ms

At 33 milliseconds, that budget is roughly six times tighter than the approximately 200 milliseconds a person takes to react to a visual stimulus. The comparison is only a reference for scale. The rig needs tight internal timing so a policy can associate a gripper command with the frame in which it happened.

Consider a pick-and-place task with roughly a two-second window for a successful gripper closure. The fingers must reach the object before closing and close soon enough to secure it.

150 ms drift ÷ 2,000 ms window = 7.5% of the action window misattributed

A 150-millisecond drift shifts the action label by 7.5% of that two-second window. An unbuffered USB camera or an unsynchronised logging thread can introduce it. The recording may look normal while repeatedly pairing the command with an earlier or later frame.

Three parallel timelines show camera, state and action ticks at regular intervals. A shaded vertical band marks a one-frame tolerance window around a gripper-close event; one action tick falls inside the band and is marked as passing, while a second, offset example falls outside the band and is marked as failing sync. Camera, state and action streams aligned within one frame A gripper-close action either lands inside the ±1-frame band or fails sync ±1 frame Camera State Action PassFail: outside tolerancetime →
Figure 1. A gripper-close action passes the sync check when it falls within one frame of the matching camera and state ticks. Outside that band, the episode cannot reliably identify which frame belongs to the action.

Checking camera calibration

Depth and multi-camera rigs also need calibration records for each rig and session. Intrinsics describe the camera’s focal length, optical centre and lens distortion, allowing a pixel to be mapped to a ray in 3D space. Extrinsics locate and orient that camera relative to the rig or world, so the ray can be expressed in a shared coordinate frame.

Moving a rig, bumping a camera or adjusting a mount can invalidate the previous extrinsics. The video may still look normal while the resulting 3D coordinates carry a consistent offset. Check and log calibration for each session.

Where the episodes come from

Perit records teleoperation in mechanical and electronics workshops, repair benches and robotics labs, as well as kitchens, living rooms, offices, retail floors, warehouses and agricultural sites. Its embodied-data programme spans 2,500-plus screened contributors, more than 120 homes, shops and worksites, and over 40 cities. Published capacity across all embodied modalities is 200 capture hours a day.

Teams can use their own teleoperation hardware, including ALOHA-class arms. Bring Your Own Rig supplies crews, rooms and QA around the team’s hardware and protocol.

A BYOR scenario: A robotics team has a working leader-follower arm pair and a capture stack tuned to their own action space, but no pipeline for recruiting operators, running sessions across varied real-world sites, or QA-checking sync and calibration at volume. Rather than replicate that infrastructure, they ship the rig and protocol; Perit supplies the crews, the kitchens and workshops to run it in, and the same sync, calibration and privacy checks applied to every other episode on the bench.

The pilot for a rig nobody has run before

A new rig goes through the same four-phase pilot used for Perit’s physical-intelligence work. Before committing volume, the pilot checks whether the SOP produces clean sync, the task suits the arm, and calibration holds across a full session.

Step 01Day 0: BriefTask family, environment, rig, sensors and episode count agreed on one page.
Step 02Week 1: SOP and kitCapture protocol and kit list written; operators trained and tested before payment.
Step 03Week 2: Pilot episodesTen to thirty episodes delivered with QC status and a quality report.
Step 04Ongoing: Review and scaleProtocol refined against pilot findings, then weekly delivery.

A useful brief names the object class, starting and ending states, failure definition and required streams. “Pick and place” alone leaves too much open. Ten to thirty pilot episodes give the team a chance to check timing across a real operator’s session and fix drift before the main order.

Modality What’s synchronised Failure mode if sync breaks
Teleoperation Camera, robot/end-effector state, action commands Policy learns the wrong instant to act; grasp timing drifts systematically
Depth RGB video and depth stream, per episode 3D geometry misaligned with the visual frame it’s supposed to describe
Tactile Video and pressure stream Contact events attributed to the wrong moment in the task
The video is the part you can watch. The synchronisation is the part that decides whether a policy can actually learn anything from it.

Tracing failures with episode metadata

Every episode needs valid environment, task, rig, operator ID, session and consent references. These records become useful when a trained policy behaves unexpectedly and the team needs to trace the examples behind it.

If a policy closes its gripper consistently early, operator and session IDs let you inspect the relevant slice of episodes. You can compare sessions, check for calibration drift and look for differences in grip technique. Without those records, the investigation starts with thousands of episodes and little way to narrow the search.

What ships at the end

The ten checks cover framing, exposure and focus, motion, sync, task compliance, duration, metadata, privacy, calibration, and label agreement where annotation was ordered. Failed episodes are re-shot or rejected without billing. For teleoperation, sync and calibration determine whether the policy can connect the recorded command to the right physical state.

Episode requirements Teleoperation records robot commands alongside camera and state streams. Check that they stay synchronised within one frame, and verify calibration for every session.

Tell us the task, the room and the rig

Perit runs teleoperation, depth and tactile capture in kitchens, workshops, warehouses and client-defined sites across more than 40 cities. Use Perit’s rig or bring your own. See the full physical intelligence lineup or send a one-page brief to get a pilot scoped within 48 hours.

The opening mug-and-gripper scene and the pick-and-place timing example are illustrative, constructed to walk through the synchronisation mechanics rather than describing a specific recorded episode. Human visual reaction time is cited as an order-of-magnitude comparison, not a precise benchmark. Figures on contributor count, site count, city count and daily capture capacity, and the ten-point QC checklist, reflect Perit’s published physical intelligence programme at the time of writing.

Was this article helpful? No ratings yet