{"id":781,"date":"2026-09-22T13:41:54","date_gmt":"2026-09-22T13:41:54","guid":{"rendered":"https:\/\/perit.ai\/blogs\/?p=781"},"modified":"2026-09-22T14:20:27","modified_gmt":"2026-09-22T14:20:27","slug":"training-data-for-humanoid-robots-what-physical-ai-needs-to-learn-from-the-real-world","status":"publish","type":"post","link":"https:\/\/perit.ai\/blogs\/training-data-for-humanoid-robots-what-physical-ai-needs-to-learn-from-the-real-world\/","title":{"rendered":"Training Data for Humanoid Robots: What Physical AI Needs to Learn From the Real World"},"content":{"rendered":"<h2>Why Humanoid Robots Need a Different Kind of Training Data<\/h2>\n<p class=\"PDq2pG_selectionAnchorContainer\" dir=\"auto\" data-start=\"227\" data-end=\"864\">Humanoid robots are learning to work in environments designed for people, but teaching them to do so requires more than better hardware or larger AI models. A robot needs to understand what it sees, follow instructions, move its body, interact with objects, and respond when something changes around it. For example, picking up a cup involves much more than recognizing the cup. The robot must understand where it is, how to reach it, how to grasp it, and what to do next. This is why humanoid robots require training data that connects perception, language, movement, and physical interaction rather than isolated images or actions.<\/p>\n<p dir=\"auto\" data-start=\"866\" data-end=\"1141\">This also means that the quality of the training data matters as much as its volume. Data needs to capture different tasks, environments, people, objects, and ways of performing the same activity so that robots can learn to handle situations beyond controlled demonstrations.<\/p>\n<h2 dir=\"auto\" data-start=\"866\" data-end=\"1141\">The Core Components of Humanoid Robot Training Data<\/h2>\n<p class=\"PDq2pG_selectionAnchorContainer\" dir=\"auto\" data-start=\"1428\" data-end=\"1698\">Humanoid robots need more than a collection of images or videos to learn physical tasks. Their training data can combine <strong data-start=\"1549\" data-end=\"1607\">visual, motion, sensor, language, and interaction data<\/strong>, allowing a model to learn not only what is happening but also how an action is performed.<\/p>\n<p dir=\"auto\" data-start=\"1700\" data-end=\"1991\"><strong data-start=\"1700\" data-end=\"1747\">Human demonstrations and robot trajectories<\/strong> are two important sources. Egocentric video can capture how people naturally interact with objects, while robot demonstrations and teleoperation record how those actions translate into robot movements. <a href=\"#reference-1\">[1]<\/a> <a href=\"#reference-2\">[2]<\/a><\/p>\n<p dir=\"auto\" data-start=\"1700\" data-end=\"1991\">A robot trajectory can be represented as:<\/p>\n<p dir=\"auto\" style=\"text-align: center;\" data-start=\"1993\" data-end=\"2045\"><strong data-start=\"1993\" data-end=\"2045\">Trajectory = {(s0, a0), (s1, a1), &#8230;, (sT, aT)}<\/strong><\/p>\n<p dir=\"auto\" data-start=\"2047\" data-end=\"2213\">Here, <strong data-start=\"2053\" data-end=\"2071\">s0, s1, &#8230; sT<\/strong> represent the robot&#8217;s state or observation at different points in time, while <strong data-start=\"2150\" data-end=\"2168\">a0, a1, &#8230; aT<\/strong> represent the actions taken at those points.<\/p>\n<p dir=\"auto\" data-start=\"2215\" data-end=\"2833\">The data can also include <strong data-start=\"2241\" data-end=\"2350\">joint positions, hand and gripper movements, depth, tactile signals, force measurements, and IMU readings<\/strong>. These provide information that ordinary RGB video cannot capture on its own, such as the 3D position of an object or the physical contact involved in a grasp. <strong data-start=\"2511\" data-end=\"2546\">Language and action annotations<\/strong> add another layer of context by connecting instructions such as &#8220;place the bottle on the table&#8221; with actions such as reaching, grasping, moving, and releasing. When these different signals are synchronised through timestamps, they give the model a more complete view of the interaction. <a href=\"#reference-1\">[1]<\/a> <a href=\"#reference-3\">[3]<\/a><\/p>\n<div class=\"\" data-turn-id-container=\"request-WEB:e6331cbb-2172-4882-a726-23f1bda73932-18\" data-is-intersecting=\"true\">\n<section class=\"text-token-text-primary w-full focus:outline-none has-data-writing-block:pointer-events-none [&amp;:has([data-writing-block])&gt;*]:pointer-events-auto R6Vx5W_threadScrollVars scroll-mb-[calc(var(--scroll-root-safe-area-inset-bottom,0px)+var(--thread-response-height))] scroll-mt-[calc(var(--header-height)+min(200px,max(70px,20svh)))]\" dir=\"auto\" data-turn-id=\"request-WEB:e6331cbb-2172-4882-a726-23f1bda73932-18\" data-turn-id-container=\"request-WEB:e6331cbb-2172-4882-a726-23f1bda73932-18\" data-testid=\"conversation-turn-38\" data-turn=\"assistant\">\n<div class=\"text-base my-auto mx-auto pb-8 [--thread-content-margin:var(--thread-content-margin-xs,calc(var(--spacing)*4))] @w-sm\/main:[--thread-content-margin:var(--thread-content-margin-sm,calc(var(--spacing)*6))] @w-lg\/main:[--thread-content-margin:var(--thread-content-margin-lg,calc(var(--spacing)*16))] px-(--thread-content-margin)\">\n<div class=\"[--thread-content-max-width:40rem] @w-lg\/main:[--thread-content-max-width:48rem] mx-auto max-w-(--thread-content-max-width) flex-1 group\/turn-messages focus-visible:outline-hidden relative flex w-full min-w-0 flex-col agent-turn\" data-conversation-screenshot-content=\"\">\n<div class=\"flex max-w-full flex-col gap-4 grow\">\n<div class=\"min-h-8 text-message relative flex w-full flex-col items-end gap-2 text-start break-words whitespace-normal outline-none keyboard-focused:focus-ring [.text-message+&amp;]:mt-1\" dir=\"auto\" tabindex=\"0\" data-message-author-role=\"assistant\" data-message-id=\"5cd1d352-3786-4a37-98ad-bb58c9fedd2f\" data-message-model-slug=\"gpt-5-6\" data-turn-start-message=\"true\">\n<div class=\"flex w-full flex-col gap-1 empty:hidden\">\n<div class=\"LR5Y_W_content markdown prose dark:prose-invert wrap-break-word w-full light markdown-new-styling\">\n<table class=\" aligncenter\" style=\"height: 350px;\" width=\"1103\">\n<thead>\n<tr>\n<th>Data type<\/th>\n<th>What it captures<\/th>\n<th>Why it matters<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Visual data<\/td>\n<td>RGB video, egocentric video, depth<\/td>\n<td>Helps the robot understand objects, people and surroundings<\/td>\n<\/tr>\n<tr>\n<td>Robot trajectories<\/td>\n<td>States, actions and movement sequences<\/td>\n<td>Shows how a physical task is performed over time<\/td>\n<\/tr>\n<tr>\n<td>Proprioception<\/td>\n<td>Joint positions, velocities and robot state<\/td>\n<td>Tells the model how the robot&#8217;s body is positioned and moving<\/td>\n<\/tr>\n<tr>\n<td>Hand and gripper data<\/td>\n<td>Hand pose, finger and gripper movements<\/td>\n<td>Important for grasping and dexterous manipulation<\/td>\n<\/tr>\n<tr>\n<td>Force and tactile data<\/td>\n<td>Contact forces, pressure and touch<\/td>\n<td>Helps the robot understand physical interaction with objects<\/td>\n<\/tr>\n<tr>\n<td>Language data<\/td>\n<td>Instructions and task descriptions<\/td>\n<td>Connects human instructions with physical actions<\/td>\n<\/tr>\n<tr>\n<td>Temporal information<\/td>\n<td>Timestamps and ordered actions<\/td>\n<td>Preserves the relationship between actions and their outcomes<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<h2 dir=\"auto\" data-section-id=\"1ka38oh\" data-start=\"0\" data-end=\"49\">Why Data Diversity Matters for Humanoid Robots<\/h2>\n<p dir=\"ltr\">Task diversity, environmental diversity, and object diversity compound rather than add. A policy trained on 10 tasks in one kitchen with one lighting setup has effectively seen one scenario, not ten. In practice, published robot manipulation datasets now aim for hundreds of distinct scenes and tens of thousands of trajectories specifically to break this narrowness. Open X-Embodiment, for instance, covers 527 distinct skills and roughly 160,000 tasks precisely because narrow coverage is the most common failure mode in deployed policies. <a href=\"#reference-1\">[1]<\/a><\/p>\n<p dir=\"ltr\">Long-horizon tasks add a second axis of difficulty. A single trajectory can be written as a chain:<\/p>\n<p dir=\"ltr\" style=\"text-align: center;\"><strong data-start=\"3947\" data-end=\"4021\">Trajectory = (state, action) \u2192 (state, action) \u2192 &#8230; \u2192 (state, action)<\/strong><\/p>\n<p dir=\"ltr\">and keeping this chain intact during data processing (rather than shuffling frames independently) is what lets a model learn how one sub-action conditions the next. A dataset with a million near-duplicate reach-and-grasp clips can be less useful for training than 50,000 trajectories spanning genuinely different tasks and conditions, because the marginal information per added trajectory drops fast once variation saturates. <a href=\"#reference-5\">[5]<\/a><\/p>\n<h2 class=\"PDq2pG_selectionAnchorContainer\" dir=\"auto\" data-section-id=\"1qmij70\" data-start=\"864\" data-end=\"913\">From Raw Demonstrations to Training-Ready Data<\/h2>\n<p class=\"PDq2pG_selectionAnchorContainer\" dir=\"auto\" data-start=\"4408\" data-end=\"4688\">A single hour of raw demonstration footage is not a training example. It&#8217;s an unstructured bundle of video, depth, robot state, and hand-tracking data that first has to be synchronized to a common clock, segmented into discrete sub-actions, annotated, and filtered before a model can learn from it. Even small timing errors between sensor streams can affect how the model associates an action with the resulting physical interaction. Action segmentation then breaks an episode into meaningful sub-phases such as reach, grasp, lift, transport, and release. Failed or corrupted takes also need to be identified before the dataset is used for training. <a href=\"#reference-5\">[5]<\/a><\/p>\n<p dir=\"auto\" style=\"text-align: center;\" data-start=\"4408\" data-end=\"4688\"><strong data-start=\"399\" data-end=\"430\">Human \/ Robot Demonstration<\/strong><br data-start=\"430\" data-end=\"433\" \/>\u2193<br data-start=\"434\" data-end=\"437\" \/><strong data-start=\"437\" data-end=\"464\">Multimodal Data Capture<\/strong><br data-start=\"464\" data-end=\"467\" \/><em data-start=\"467\" data-end=\"526\">Video \u00b7 Depth \u00b7 Robot State \u00b7 Hand Motion \u00b7 Force\/Tactile<\/em><br data-start=\"526\" data-end=\"529\" \/>\u2193<br data-start=\"530\" data-end=\"533\" \/><strong data-start=\"533\" data-end=\"557\">Time Synchronization<\/strong><br data-start=\"557\" data-end=\"560\" \/><em data-start=\"560\" data-end=\"586\">Align all sensor streams<\/em><br data-start=\"586\" data-end=\"589\" \/>\u2193<br data-start=\"590\" data-end=\"593\" \/><strong data-start=\"593\" data-end=\"616\">Action Segmentation<\/strong><br data-start=\"616\" data-end=\"619\" \/><em data-start=\"619\" data-end=\"663\">Reach \u00b7 Grasp \u00b7 Lift \u00b7 Transport \u00b7 Release<\/em><br data-start=\"663\" data-end=\"666\" \/>\u2193<br data-start=\"667\" data-end=\"670\" \/><strong data-start=\"670\" data-end=\"684\">Annotation<\/strong><br data-start=\"684\" data-end=\"687\" \/><em data-start=\"687\" data-end=\"730\">Actions \u00b7 Objects \u00b7 Instructions \u00b7 Events<\/em><br data-start=\"730\" data-end=\"733\" \/>\u2193<br data-start=\"734\" data-end=\"737\" \/><strong data-start=\"737\" data-end=\"756\">Quality Control<\/strong><br data-start=\"756\" data-end=\"759\" \/><em data-start=\"759\" data-end=\"810\">Remove corrupted, incomplete or failed recordings<\/em><br data-start=\"810\" data-end=\"813\" \/>\u2193<br data-start=\"814\" data-end=\"817\" \/><strong data-start=\"817\" data-end=\"843\">Training-Ready Dataset<\/strong><\/p>\n<h2 dir=\"auto\" data-start=\"3694\" data-end=\"3966\">The Challenges of Collecting Humanoid Robot Training Data<\/h2>\n<p dir=\"ltr\">A single capture rig can combine RGB or stereo cameras, structured-light or ToF depth sensors, IMUs, motion capture markers, hand-tracking cameras, force-torque sensors at the wrist, and tactile arrays on the fingertips, each running on its own clock and sample rate. Camera streams typically run at 30 to 60 fps, while force-torque sensors often sample at 500 Hz to 1 kHz to capture fast contact transients. Reconciling those disparate rates without losing precision is a nontrivial engineering problem before annotation even starts.<\/p>\n<p dir=\"ltr\">Consent, privacy, and secure handling become part of the pipeline as soon as real people are recorded performing the demonstrations, and this typically means de-identifying footage, restricting downstream use, and tracking data provenance per institution when collection spans multiple sites. <a href=\"#reference-2\">[2]<\/a><\/p>\n<div class=\"\" data-turn-id-container=\"request-WEB:e6331cbb-2172-4882-a726-23f1bda73932-33\" data-is-intersecting=\"true\">\n<section class=\"text-token-text-primary w-full focus:outline-none has-data-writing-block:pointer-events-none [&amp;:has([data-writing-block])&gt;*]:pointer-events-auto R6Vx5W_threadScrollVars scroll-mb-[calc(var(--scroll-root-safe-area-inset-bottom,0px)+var(--thread-response-height))] scroll-mt-[calc(var(--header-height)+min(200px,max(70px,20svh)))]\" dir=\"auto\" data-turn-id=\"request-WEB:e6331cbb-2172-4882-a726-23f1bda73932-33\" data-turn-id-container=\"request-WEB:e6331cbb-2172-4882-a726-23f1bda73932-33\" data-testid=\"conversation-turn-68\" data-turn=\"assistant\">\n<div class=\"text-base my-auto mx-auto pb-8 [--thread-content-margin:var(--thread-content-margin-xs,calc(var(--spacing)*4))] @w-sm\/main:[--thread-content-margin:var(--thread-content-margin-sm,calc(var(--spacing)*6))] @w-lg\/main:[--thread-content-margin:var(--thread-content-margin-lg,calc(var(--spacing)*16))] px-(--thread-content-margin)\">\n<div class=\"[--thread-content-max-width:40rem] @w-lg\/main:[--thread-content-max-width:48rem] mx-auto max-w-(--thread-content-max-width) flex-1 group\/turn-messages focus-visible:outline-hidden relative flex w-full min-w-0 flex-col agent-turn\" data-conversation-screenshot-content=\"\">\n<div class=\"flex max-w-full flex-col gap-4 grow\">\n<div class=\"min-h-8 text-message relative flex w-full flex-col items-end gap-2 text-start break-words whitespace-normal outline-none keyboard-focused:focus-ring [.text-message+&amp;]:mt-1\" dir=\"auto\" tabindex=\"0\" data-message-author-role=\"assistant\" data-message-id=\"a58c214c-fdee-4cd8-9329-9717b387fa54\" data-message-model-slug=\"gpt-5-6\" data-turn-start-message=\"true\">\n<div class=\"flex w-full flex-col gap-1 empty:hidden\">\n<div class=\"LR5Y_W_content markdown prose dark:prose-invert wrap-break-word w-full light markdown-new-styling\">\n<h2 dir=\"auto\" data-start=\"0\" data-end=\"40\">From Training Data to Robot Behaviour<\/h2>\n<p class=\"PDq2pG_selectionAnchorContainer\" dir=\"auto\" data-start=\"42\" data-end=\"397\">Imitation learning, and specifically behaviour cloning, treats the problem as supervised learning: given an observation, predict the action. What separates a Vision-Language-Action (VLA) model from earlier behaviour-cloning policies is that it can be trained on vision-language data alongside robot trajectories, connecting semantic knowledge with physical actions. Google DeepMind&#8217;s RT-2 provides a clear example. The model was co-fine-tuned on robotic trajectory data and Internet-scale vision-language tasks, and its evaluation across roughly 6,000 trials showed improved generalisation to novel objects, environments, and instructions. <a href=\"#reference-5\">[5]<\/a> <a href=\"#reference-6\">[6]<\/a><\/p>\n<h2 dir=\"auto\" data-start=\"1058\" data-end=\"1341\">Scaling Physical AI Training Data<\/h2>\n<div class=\"flex max-w-full flex-col gap-4 grow\">\n<div class=\"min-h-8 text-message relative flex w-full flex-col items-end gap-2 text-start break-words whitespace-normal outline-none keyboard-focused:focus-ring [.text-message+&amp;]:mt-1\" dir=\"auto\" tabindex=\"0\" data-message-author-role=\"assistant\" data-message-id=\"a4bad857-e192-4980-a745-d8ad10df767c\" data-message-model-slug=\"gpt-5-6\" data-turn-start-message=\"true\">\n<div class=\"flex w-full flex-col gap-1 empty:hidden\">\n<div class=\"LR5Y_W_content markdown prose dark:prose-invert wrap-break-word w-full light markdown-new-styling\">\n<p class=\"PDq2pG_selectionAnchorContainer\" dir=\"auto\" data-start=\"38\" data-end=\"355\">Task specificity determines which signals actually matter. Dexterous manipulation policies lean heavily on hand and gripper trajectories sampled at high frequency, often 100 Hz or above for finger joint angles, while whole-body locomotion or loco-manipulation tasks depend more on IMU, joint torque, and balance-relevant proprioception. A dataset over-indexed on one modality for a task that needs another (say, thousands of hours of RGB video for a task that is fundamentally about contact force) produces a policy that looks complete on paper but fails at deployment. <a href=\"#reference-1\">[1]<\/a> <a href=\"#reference-6\">[6]<\/a><\/p>\n<\/div>\n<\/div>\n<\/div>\n<\/div>\n<h2 class=\"pointer-events-none -mb-px h-px w-full opacity-0\" aria-hidden=\"true\">How Perit Supports Physical AI Data Collection<\/h2>\n<p class=\"PDq2pG_selectionAnchorContainer\" dir=\"auto\" data-start=\"200\" data-end=\"555\">Perit AI&#8217;s physical AI data work spans egocentric video, depth, teleoperation, and handheld or gripper-based demonstrations, with the capture setup adapted per project rather than fixed to one rig. That matters because a dexterous fine-manipulation dataset and a whole-body mobile-manipulation dataset need genuinely different sensor loadouts, not just different content. Perit AI&#8217;s physical AI data work spans <strong data-start=\"7581\" data-end=\"7660\">egocentric video, depth, teleoperation, and handheld-gripper demonstrations<\/strong>, with capture setups adapted to different collection requirements. Its current physical-AI programs are in early access and being scoped with partners. <a href=\"#reference-8\">[8]<\/a><\/p>\n<h2 dir=\"auto\" data-start=\"1123\" data-end=\"1385\">Where the Data Fits in a Robotics Training Pipeline<\/h2>\n<p dir=\"ltr\">The value of a dataset is set by how cleanly it survives the journey from raw demonstration to trained policy, not by its size alone. A typical pipeline runs:<\/p>\n<p dir=\"ltr\">data collection \u2192 annotation \u2192 quality checks \u2192 dataset formatting \u2192 model training \u2192 evaluation. <a href=\"#reference-5\">[5]<\/a><\/p>\n<p dir=\"ltr\">The failure mode to watch for is silent information loss at the boundaries between these stages. If action segmentation is off by even 100 to 200 milliseconds, or if an annotation pass drops the language instruction tied to a trajectory, the model trains on a mismatched (observation, action) pair and the error never shows up until evaluation, when it looks like a model problem rather than a data problem.<\/p>\n<p dir=\"ltr\">Two published foundation models illustrate why the fit has to be right for the target system, not just technically correct. Octo, trained on 800,000 trajectories drawn from Open X-Embodiment, was built specifically to transfer zero-shot across manipulation tasks <a href=\"#reference-7\">[7]<\/a>. OpenVLA, trained on a larger 970,000-trajectory OXE subset with a 7-billion-parameter Llama 2 backbone, instead prioritized a stronger pretrained vision-language foundation underneath the robot fine-tuning. Same source dataset, two different scale and architecture choices, because the two models were built for different jobs <a href=\"#reference-9\">[9]<\/a>. A dexterous-manipulation model needs dense hand and gripper data. A whole-body locomotion model needs depth and joint-position streams at a different sampling rate. The dataset has to be shaped for the destination, not just cleaned and handed over.<\/p>\n<p dir=\"ltr\" style=\"text-align: center;\"><strong data-start=\"1284\" data-end=\"1303\">Real-World Data<\/strong><br data-start=\"1303\" data-end=\"1306\" \/><em data-start=\"1306\" data-end=\"1360\">Human demonstrations \u00b7 Robot teleoperation \u00b7 Sensors<\/em><br data-start=\"1360\" data-end=\"1363\" \/>\u2193<br data-start=\"1364\" data-end=\"1367\" \/><strong data-start=\"1367\" data-end=\"1386\">Curated Dataset<\/strong><br data-start=\"1386\" data-end=\"1389\" \/><em data-start=\"1389\" data-end=\"1433\">Synchronised \u00b7 Annotated \u00b7 Quality checked<\/em><br data-start=\"1433\" data-end=\"1436\" \/>\u2193<br data-start=\"1437\" data-end=\"1440\" \/><strong data-start=\"1440\" data-end=\"1458\">Model Training<\/strong><br data-start=\"1458\" data-end=\"1461\" \/><em data-start=\"1461\" data-end=\"1507\">Imitation Learning \u00b7 Behaviour Cloning \u00b7 VLA<\/em><br data-start=\"1507\" data-end=\"1510\" \/>\u2193<br data-start=\"1511\" data-end=\"1514\" \/><strong data-start=\"1514\" data-end=\"1530\">Robot Policy<\/strong><br data-start=\"1530\" data-end=\"1533\" \/><em data-start=\"1533\" data-end=\"1586\">Predicts actions from observations and instructions<\/em><br data-start=\"1586\" data-end=\"1589\" \/>\u2193<br data-start=\"1590\" data-end=\"1593\" \/><strong data-start=\"1593\" data-end=\"1618\">Real-World Evaluation<\/strong><br data-start=\"1618\" data-end=\"1621\" \/><em data-start=\"1621\" data-end=\"1651\">Success \u00b7 Failure \u00b7 Recovery<\/em><br data-start=\"1651\" data-end=\"1654\" \/>\u2193<br data-start=\"1655\" data-end=\"1658\" \/><strong data-start=\"1658\" data-end=\"1685\">New Data \/ Improvements<\/strong><\/p>\n<h2 dir=\"auto\" data-start=\"976\" data-end=\"1201\">What Makes Physical AI Data Useful at Scale<\/h2>\n<div class=\"relative basis-auto flex-col -mb-(--composer-overlap-px) pb-(--composer-overlap-px) [--composer-overlap-px:28px] grow flex\">\n<div class=\"flex min-h-0 grow flex-col text-sm\">\n<div class=\"qMYqUG_convSearchResultHighlightRoot\">\n<div class=\"\" data-turn-id-container=\"request-WEB:e6331cbb-2172-4882-a726-23f1bda73932-41\" data-is-intersecting=\"true\">\n<section class=\"text-token-text-primary w-full focus:outline-none has-data-writing-block:pointer-events-none [&amp;:has([data-writing-block])&gt;*]:pointer-events-auto R6Vx5W_threadScrollVars scroll-mb-[calc(var(--scroll-root-safe-area-inset-bottom,0px)+var(--thread-response-height))] scroll-mt-[calc(var(--header-height)+min(200px,max(70px,20svh)))]\" dir=\"auto\" data-turn-id=\"request-WEB:e6331cbb-2172-4882-a726-23f1bda73932-41\" data-turn-id-container=\"request-WEB:e6331cbb-2172-4882-a726-23f1bda73932-41\" data-testid=\"conversation-turn-84\" data-turn=\"assistant\">\n<div class=\"text-base my-auto mx-auto pb-8 [--thread-content-margin:var(--thread-content-margin-xs,calc(var(--spacing)*4))] @w-sm\/main:[--thread-content-margin:var(--thread-content-margin-sm,calc(var(--spacing)*6))] @w-lg\/main:[--thread-content-margin:var(--thread-content-margin-lg,calc(var(--spacing)*16))] px-(--thread-content-margin)\">\n<div class=\"[--thread-content-max-width:40rem] @w-lg\/main:[--thread-content-max-width:48rem] mx-auto max-w-(--thread-content-max-width) flex-1 group\/turn-messages focus-visible:outline-hidden relative flex w-full min-w-0 flex-col agent-turn\" data-conversation-screenshot-content=\"\">\n<div class=\"flex max-w-full flex-col gap-4 grow\">\n<div class=\"min-h-8 text-message relative flex w-full flex-col items-end gap-2 text-start break-words whitespace-normal outline-none keyboard-focused:focus-ring [.text-message+&amp;]:mt-1\" dir=\"auto\" tabindex=\"0\" data-message-author-role=\"assistant\" data-message-id=\"a64ba61f-2603-4078-b519-ccf7714f0061\" data-message-model-slug=\"gpt-5-6\" data-turn-start-message=\"true\">\n<div class=\"flex w-full flex-col gap-1 empty:hidden\">\n<div class=\"LR5Y_W_content markdown prose dark:prose-invert wrap-break-word w-full light markdown-new-styling\">\n<p class=\"PDq2pG_selectionAnchorContainer\" dir=\"auto\" data-start=\"48\" data-end=\"361\">Task specificity determines which signals actually matter. Dexterous manipulation policies lean heavily on hand and gripper trajectories sampled at high frequency, often 100 Hz or above for finger joint angles, while whole-body locomotion or loco-manipulation tasks depend more on IMU, joint torque, and balance-relevant proprioception. A dataset over-indexed on one modality for a task that needs another (say, thousands of hours of RGB video for a task that is fundamentally about contact force) produces a policy that looks complete on paper but fails at deployment. Consistency compounds this: a large dataset with inconsistent annotation schemas across contributors is often less usable than a smaller one annotated to a single standard, because a model trained on inconsistently labeled data learns the label noise as readily as it learns the task.<\/p>\n<h2 dir=\"auto\" data-start=\"982\" data-end=\"1292\">The Role of Real-World Data in the Next Generation of Robotics<\/h2>\n<p class=\"PDq2pG_selectionAnchorContainer\" dir=\"auto\" data-start=\"67\" data-end=\"459\">Simulated environments can generate demonstration volume cheaply, effectively unlimited episodes at near-zero marginal cost, but every simulator carries a sim-to-real gap: mismatches in contact dynamics, friction coefficients, lighting, and sensor noise between the simulated and physical world. That gap is why real-world data continues to carry disproportionate weight per trajectory in published training recipes, even in setups where synthetic data makes up the majority of total training volume by episode count. The emerging pattern is not real-world data being replaced by simulation, but hybrid pipelines where simulation handles breadth (rare edge cases, dangerous scenarios, exhaustive object variation) while real-world demonstrations anchor the physical grounding that policies are ultimately evaluated against. <a href=\"#reference-1\">[1]<\/a> <a href=\"#reference-7\">[7]<\/a><\/p>\n<h2 dir=\"auto\" data-start=\"793\" data-end=\"1142\">The Future of Humanoid Robot Training Data<\/h2>\n<p class=\"PDq2pG_selectionAnchorContainer\" dir=\"auto\" data-start=\"47\" data-end=\"322\">The bottleneck is increasingly shifting from simply collecting more data to structuring it well. As datasets grow, consistent annotation, meaningful task and environment diversity, and clear quality-control criteria become increasingly important for making the data useful for training. Research on imitation learning has shown that the quality and distribution of demonstrations can significantly affect how well a policy performs, meaning that more data does not automatically translate into better results. <a href=\"#reference-5\">[5]<\/a> The practical implication for teams building these pipelines is that annotation schemas and quality-control criteria should be defined before large-scale collection begins, rather than being retrofitted afterward.<\/p>\n<div class=\"flex max-w-full flex-col gap-4 grow\">\n<div class=\"min-h-8 text-message relative flex w-full flex-col items-end gap-2 text-start break-words whitespace-normal outline-none keyboard-focused:focus-ring [.text-message+&amp;]:mt-1\" dir=\"auto\" tabindex=\"0\" data-message-author-role=\"assistant\" data-message-id=\"587b59ae-349a-4169-8b22-ffc9c1914f53\" data-message-model-slug=\"gpt-5-6\" data-turn-start-message=\"true\">\n<div class=\"flex w-full flex-col gap-1 empty:hidden\">\n<div class=\"LR5Y_W_content markdown prose dark:prose-invert wrap-break-word w-full light markdown-new-styling\">\n<h2 class=\"PDq2pG_selectionAnchorContainer\" dir=\"auto\" data-start=\"111\" data-end=\"148\">Where Humanoid Robotics Is Heading<\/h2>\n<p dir=\"auto\" data-start=\"150\" data-end=\"562\">Humanoid robots need training data that binds perception, language, and physical action into one stream, not hardware and model capacity alone. The open question for the field isn&#8217;t whether to collect more data, it&#8217;s whether that data is diverse enough in tasks, environments, and objects, and structured tightly enough (synchronized, segmented, consistently annotated) for a policy to actually learn the behaviours it needs. As real-world data collection matures into its own specialized layer of the robotics stack, the datasets that combine both properties, genuine diversity and training-ready structure, are what will determine how far these systems generalize beyond their training distribution. <a href=\"#reference-1\">[1]<\/a> <a href=\"#reference-5\">[5]<\/a> <a href=\"#reference-6\">[6]<\/a><\/p>\n<h2>References<\/h2>\n<p id=\"reference-1\"><strong>[1]<\/strong> Open X-Embodiment Collaboration. \u201cOpen X-Embodiment: Robotic Learning Datasets and RT-X Models.\u201d <em>arXiv<\/em>, 2023.<br \/>\n<a href=\"https:\/\/arxiv.org\/abs\/2310.08864\">Open X-Embodiment \u2014 arXiv<\/a><\/p>\n<p id=\"reference-2\"><strong>[2]<\/strong> Grauman, Kristen, et al. \u201cEgo4D: Around the World in 3,000 Hours of Egocentric Video.\u201d <em>CVPR<\/em>, 2022.<br \/>\n<a href=\"https:\/\/arxiv.org\/abs\/2110.07058\">Ego4D \u2014 arXiv<\/a><\/p>\n<p id=\"reference-3\"><strong>[3]<\/strong> Jiang, Yunfan, et al. \u201cRoboMIND: Benchmark on Multi-embodiment Intelligence Normative Data for Robot Manipulation.\u201d <em>arXiv<\/em>, 2024.<br \/>\n<a href=\"https:\/\/arxiv.org\/abs\/2412.13877\">RoboMIND \u2014 arXiv<\/a><\/p>\n<p id=\"reference-4\"><strong>[4]<\/strong> Khazatsky, Alexander, et al. \u201cDROID: A Large-Scale In-the-Wild Robot Manipulation Dataset.\u201d <em>arXiv<\/em>, 2024.<br \/>\n<a href=\"https:\/\/arxiv.org\/abs\/2403.12945\">DROID \u2014 arXiv<\/a><\/p>\n<p id=\"reference-5\"><strong>[5]<\/strong> Belkhale, Suneel, Yuchen Cui, and Dorsa Sadigh. \u201cData Quality in Imitation Learning.\u201d <em>arXiv<\/em>, 2023.<br \/>\n<a href=\"https:\/\/arxiv.org\/abs\/2306.02437\">Data Quality in Imitation Learning \u2014 arXiv<\/a><\/p>\n<p id=\"reference-6\"><strong>[6]<\/strong> Zitkovich, Brianna, et al. \u201cRT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.\u201d <em>Proceedings of the 7th Conference on Robot Learning<\/em>, 2023.<br \/>\n<a href=\"https:\/\/proceedings.mlr.press\/v229\/zitkovich23a.html\">RT-2 \u2014 PMLR<\/a><\/p>\n<p id=\"reference-7\"><strong>[7]<\/strong> Octo Model Team, et al. \u201cOcto: An Open-Source Generalist Robot Policy.\u201d <em>Robotics: Science and Systems<\/em>, 2024.<br \/>\n<a href=\"https:\/\/arxiv.org\/abs\/2405.12213\">Octo \u2014 arXiv<\/a><\/p>\n<p id=\"reference-8\"><strong>[8]<\/strong> Perit. \u201cPhysical Intelligence.\u201d <em>Perit<\/em>, 2026.<br \/>\n<a href=\"https:\/\/perit.ai\/physical-intelligence\">Perit \u2014 Physical Intelligence<\/a><\/p>\n<p id=\"reference-9\"><strong>[9]<\/strong> Kim, Moo Jin, et al. \u201cOpenVLA: An Open-Source Vision-Language-Action Model.\u201d <em>Proceedings of the 8th Conference on Robot Learning<\/em>, 2024\/2025.<br \/>\n<a href=\"https:\/\/arxiv.org\/abs\/2406.09246\">OpenVLA \u2014 arXiv<\/a><\/p>\n<\/div>\n<\/div>\n<\/div>\n<\/div>\n<\/div>\n<\/div>\n<\/div>\n<\/div>\n<\/div>\n<\/div>\n<\/section>\n<\/div>\n<\/div>\n<\/div>\n<\/div>\n<\/div>\n<\/div>\n<\/div>\n<\/div>\n<\/div>\n<\/div>\n<\/section>\n<\/div>\n<\/div>\n<\/div>\n<\/div>\n<\/div>\n<\/div>\n<\/div>\n<\/section>\n<\/div>\n","protected":false},"excerpt":{"rendered":"<p>Why Humanoid Robots Need a Different Kind of Training Data Humanoid robots are learning to work in environments designed for people, but teaching them to do so requires more than\u2026<\/p>\n","protected":false},"author":1,"featured_media":806,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[18],"tags":[],"class_list":["post-781","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-physical-intelligence"],"_links":{"self":[{"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/posts\/781","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/comments?post=781"}],"version-history":[{"count":5,"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/posts\/781\/revisions"}],"predecessor-version":[{"id":808,"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/posts\/781\/revisions\/808"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/media\/806"}],"wp:attachment":[{"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/media?parent=781"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/categories?post=781"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/tags?post=781"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}