From Human Demonstrations to Self-Improving Robots: The Data Behind Physical AI

From Human Demonstrations to Self-Improving Robots: The Data Behind Physical AI

How Robots Learn from Data

Many robot skills start as a set of recordings, and one of the most direct ways to learn from them is by demonstration. A person shows the robot how to do a task, the robot records what it sensed and how it moved, and a model is trained to reproduce that behavior in similar situations. This is imitation learning, and its simplest form, behavior cloning, works like supervised learning: given what the robot currently sees, predict the action the demonstrator took. Demonstrations are usually produced through teleoperation, where an operator controls the robot directly while its sensors record the robot’s observations and actions. (What those recordings contain, from camera views to joint state and touch, is covered in our earlier post on training data for humanoid robots).

The alternative is learning by experience. Here the robot acts on its own, tries something, observes the result, and adjusts. Reinforcement learning formalizes this idea: behavior that leads to good outcomes is reinforced, and behavior that doesn’t is discouraged. The data comes from the robot’s own attempts, collected autonomously, and it includes failures as well as successes, which demonstrations rarely show. The catch is that a robot with no skill yet mostly produces clumsy attempts, so useful signal can be slow to appear, and trial and error in the physical world can be costly and unsafe.

Demonstration gives a robot a strong starting point, and experience lets it go beyond what any demonstrator showed. Each source teaches something different and suits a different stage of learning. The rest of this article looks at how each is collected, what each costs, and how the strongest teams chain them together, starting with humans and gradually letting the robot take over.

Teleoperation Data: Learning from Human Demonstrations

Teleoperation collects demonstrations by putting a human in control of the robot while its sensors record the robot’s observations and actions. The setups vary widely. In leader-follower rigs, an operator moves a lightweight replica arm and the robot mirrors it in real time, the approach behind ALOHA [1] and, with a mobile base, Mobile ALOHA [2]. Other options include VR and controller-based operation, motion-capture gloves for fine hand kinematics, and handheld devices such as the Universal Manipulation Interface (UMI) [3], which lets people collect demonstrations without the robot being present at all.

Capture method How it works Strength Limitation
Leader-follower rig Operator moves a replica arm, robot mirrors it Precise, natural control; data is already in the robot’s own frame Needs the robot on site; operators need training
VR / controller Headset and controllers map hand motion to the robot Easy to set up; supports remote operation Weaker touch feedback; latency affects fine tasks
Handheld device (UMI-style) Human holds a gripper fitted with a camera Collect anywhere, without a robot Must match the robot’s gripper; misses robot-body dynamics
Motion-capture gloves Track finger joints during the task Rich hand detail for dexterous manipulation Needs retargeting to the robot hand and careful calibration

The value of this data is its density. A large share of the data shows deliberate, goal-directed behavior, and operators naturally include the small corrections and recoveries that make a policy robust. A standard way to use it is behavior cloning, where a policy πθ is trained to predict the operator’s action a from what the robot observes, o, by minimizing the gap between the two:

L(θ) = E[ ‖ πθ(o) − a ‖² ]

over (o, a) pairs in the demonstration dataset D

The main limits are practical and statistical. Teleoperation costs more per hour than autonomous collection, throughput depends on skilled operators, and the data carries each operator’s personal style. There is also a well-known technical problem: a policy trained only on demonstrations mainly sees the states visited by the demonstrator. A small mistake can push it into an unfamiliar state, where subsequent errors can compound [4]. The usual fix is to let the robot run and have humans step in and correct it when it drifts [4], which is where the robot’s own experience starts to matter.

Learning from Experience with Autonomous Data

In autonomous collection, the robot acts on its own. It runs a policy, which may be a hand-written script, a partially trained model, or an exploration strategy, and logs every observation, action, and outcome. Early large-scale efforts, such as the grasping work of Levine et al. [5] and QT-Opt [6], used fleets of robots to collect large numbers of grasping attempts. More recently, RoboCat [7] showed a generalist agent generating new training data from its own attempts and feeding it back into training.

The appeal is scale without operators. Once the setup exists, each extra hour costs little, and the data includes failures and near-misses that demonstrations rarely contain. One common framework for learning from this kind of data is reinforcement learning, where the policy π is trained to maximize the expected discounted return:

J(π) = E[ Σ γᵗ r(sₜ, aₜ) ]

where r is the reward the robot receives for taking action aₜ in state sₜ, and γ (between 0 and 1) discounts rewards that arrive later. For reinforcement learning, a reward or success signal provides the feedback that tells the model which outcomes to favor.

The limits mirror those of teleoperation. While a policy is still weak, most attempts fail, so useful signal is sparse. Scenes have to be reset between attempts, by people or extra hardware, and unsupervised trials can damage the robot, the objects, and the surroundings. Attempts still need filtering and success labeling, and because the data reflects what the current policy already does, it can stay narrow. Neither source is complete on its own, and the differences become clearer when the two are compared directly.

Comparing Teleoperation and Autonomous Data

Set side by side, the two sources differ along a handful of dimensions: signal quality, cost, scale, coverage of failures, risk, and the stage of a project each suits. Cost per hour is a misleading measure on its own. Teleoperation costs more per hour, but a large share of what it records is usable. Autonomous hours are cheap, yet while a policy is still weak, most of what it records teaches a model very little. Even demonstration quality isn’t uniform, since it varies with the skill and habits of each operator [8]. The table below lays out the differences.

Dimension Teleoperation data Autonomous data
Signal quality High and intentional, with reliable task completions Sparse while the policy is weak, improves as it gets better
Cost per hour Higher, due to skilled operators and hardware Lower once the setup is running
Scalability Limited by operator availability and collection sites Limited by hardware, scene resets, and safety
Failure and recovery coverage Limited unless collected on purpose Broad, since failures are logged naturally
Diversity Can be narrow with few operators or sites Bounded by what the current policy tries
Risk Low, since a human is in control Higher, with possible damage to robot, objects, and surroundings
Best stage Bootstrapping and dexterous tasks Scaling and refining a capable policy

The graph below shows the core trade-off.

Illustrative only. The curves show the general trade-off between the two data sources, not measured results.

Teleoperation data is valuable from the start and stays fairly steady, while autonomous data begins low and rises as the policy gets more capable. Where the two curves cross depends on the task and the robot, and neither source wins outright. That is why the strongest teams use both in sequence, which is the subject of the flywheel section.

The Data Flywheel That Lets Robots Improve Themselves

In practice, the choice isn’t either-or. Capable teams run both sources as a loop. It starts with teleoperation: a modest set of human demonstrations bootstraps a first policy that is capable enough to attempt the task, even if unreliably. That policy is then deployed to collect data on its own, at a scale and cost that demonstrations can’t match. The loop looks like this.

The data flywheel. Human demonstrations bootstrap a policy, the robot collects data on its own, and human corrections at the edges feed the next round of training.

Human effort shifts rather than disappears. When the policy fails or drifts, an operator steps in and corrects it, and those corrections, recorded at exactly the moments the policy struggled, are among the most informative data in the loop [4] [9]. The policy is then retrained on the growing mix of demonstrations, autonomous attempts, and corrections, and the improved version goes back out to collect the next round. RoboCat [7] follows this pattern: it was adapted to new tasks from a small set of demonstrations, generated additional data from its own attempts, and was trained again on the expanded dataset.

Each turn of the loop raises the share of autonomous data that is actually useful, moving the policy along the curve from the previous section, while human effort concentrates on the edge cases where it matters most. The loop depends on things that are easy to underestimate: reliable success labeling, clean logging, consistent formats, and enough diversity in tasks and environments that the policy doesn’t improve in only one narrow place. That is where data operations, and not model design, start to decide how far the flywheel turns.

Running Data Collection at Scale

Both sources work well inside a single lab. The difficulty appears when collection has to spread across many operators, sites, and robots. Teleoperation at scale becomes a people and logistics problem: operators must be trained to a common protocol, sessions have to be planned around fatigue, hardware must be calibrated the same way at every site, and quality has to be tracked per operator so that one person’s habits don’t quietly dominate the dataset. Distributed efforts such as DROID [10], which gathered demonstrations from many collectors across hundreds of scenes and several institutions, show how much coordination this takes.

Autonomous collection has its own version of the problem. A fleet has to be monitored, scenes have to be reset, failures have to be detected and triaged, and safety limits have to hold when nobody is watching every robot. In both cases the logging layer matters as much as the robots. Consistent formats, reliable success labels, and metadata about the task, environment, robot, and operator are what allow data to be filtered, audited, and traced back to its source later.

Illustrative only. Relative cost per useful hour when the cost per collected hour is held constant, based on the share of collected hours that produce usable data.

This is why scale moves the bottleneck. Early on, the constraint is the algorithm. Later, it is whether data can be produced reliably, consistently, and across enough varied conditions. Teams that treat collection as an operations discipline, with defined protocols, per-operator quality tracking, and automated checks, get more from every hour than teams that only add hours.

How Perit Approaches Human-Led Robot Data

Every turn of the flywheel begins with human-led data, and that is where Perit AI works. Perit’s physical AI data work spans teleoperation, handheld and gripper-based demonstrations, egocentric video, and depth capture, with the setup adapted to each project instead of fixed to one rig. That flexibility matters because the right capture method depends on the task. A dexterous fine-manipulation dataset and a whole-body mobile-manipulation dataset call for different sensors, different operators, and different protocols.

The operational questions raised in this article are the ones Perit works through with partners: how operators are trained to a common protocol, how hardware is calibrated the same way across sites, how quality is tracked per operator, and what metadata the logging layer carries so that data can be filtered and audited later. Those choices decide whether a first set of demonstrations is ready to bootstrap a policy, and whether the corrections that come later are as informative as they should be.

Project type Typical capture setup What matters most
Dexterous fine manipulation Teleoperation or handheld-gripper demonstrations High-frequency hand and gripper data, contact information
Whole-body mobile manipulation Teleoperation with depth and body-state streams Joint and balance signals, wide-area depth
Learning from natural human behavior Egocentric video with depth Realistic environments, variety of people and objects

The table shows how the setup changes with the task.

Perit’s physical AI programs are currently in early access and are being scoped with partners, so each collection setup is built around the team’s task, hardware, and stage in the loop, whether they are gathering a first set of demonstrations or planning where humans fit into a loop that is already running. To discuss a physical AI data project, contact Perit.

Building Robots That Keep Improving

Teleoperation and autonomous collection are not competing answers to the same question. Demonstrations give a robot competence it could not find by itself, autonomous collection extends that competence at a scale no operator team can match, and human corrections keep both honest at the edges. The teams that make progress treat these as stages of a single loop, and choose the mix based on how capable their policy already is.

As robots move from labs into homes, warehouses, and factories, the quality and consistency of the data behind them increasingly decides how far they can improve. Getting the first demonstrations right, and running collection so that every later turn of the loop stays reliable, is what lets a robot go from imitating people to improving on its own. For more on what that data contains, see our earlier post on training data for humanoid robots.

References

[1] T. Z. Zhao et al., “Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware,” RSS, 2023.

[2] Z. Fu et al., “Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation,” CoRL, 2024.

[3] C. Chi et al., “Universal Manipulation Interface: In-The-Wild Robot Teaching Without In-The-Wild Robots,” RSS, 2024.

[4] S. Ross et al., “A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning,” AISTATS, 2011.

[5] S. Levine et al., “Learning Hand-Eye Coordination for Robotic Grasping with Deep Learning and Large-Scale Data Collection,” IJRR, 2018.

[6] D. Kalashnikov et al., “Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation” (QT-Opt), CoRL, 2018.

[7] K. Bousmalis et al., “RoboCat: A Self-Improving Generalist Agent for Robotic Manipulation,” TMLR, 2024.

[8] A. Mandlekar et al., “What Matters in Learning from Offline Human Demonstrations for Robot Manipulation,” CoRL, 2022.

[9] M. Kelly et al., “HG-DAgger: Interactive Imitation Learning with Human Experts,” ICRA, 2019.

[10] A. Khazatsky et al., “DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset,” RSS, 2024.

Was this article helpful? No ratings yet