What Training Data Does a Humanoid Robot Actually Need?

Cloudpano
September 30, 2026
•
5 min read
Share this post

What training data does a humanoid robot actually need?

A humanoid robot needs layered training data: web and video data for general understanding, egocentric human video for broad manipulation experience, simulation for balance and control, and robot data (teleoperation or autonomous runs) to fit its own body. Figure reports that pretraining on human video raised zero-shot success in 30 unseen homes from 9% to 56%, but each company in this guide still fine-tunes on robot data.

Key Takeaways

  • Humanoid training data comes in layers. NVIDIA describes a pyramid: web data and human video at the base, synthetic data in the middle, real robot data at the top.
  • Human video is now the pretraining layer that scales. Figure's Index pretraining raised zero-shot success across 30 unseen homes from 9% to 56%.
  • Robot data is still required. 1X fine-tunes on 70 hours of robot data after 900 hours of egocentric human video; Figure's first Helix used about 500 hours of teleoperation.
  • Simulation handles balance and whole-body control. Agility trains Digit's motor-control model in simulation for decades of simulated time over three or four days.
  • More hours alone isn't the goal. Diversity of tasks, objects and environments per hour, plus filtering and labels, decides what the data is worth.

What Training Data Does a Humanoid Robot Actually Need?

If you are building a humanoid, the hardest question is often not the model but the data mix. In September 2026, Figure reported that pretraining on human video raised its robot's zero-shot success in 30 unseen homes from 9% to 56%, with the robot-specific data held fixed (Figure, Helix 2.5).

That result explains a lot of the current buying spree, which we covered in why robotics labs are buying egocentric video at scale. But it doesn't mean human video is all a humanoid needs. Every company in this guide still trains on its own robot data, and most lean on simulation too. Below is what each layer of humanoid robot training data does, using only what Figure, 1X, Agility, Apptronik and NVIDIA have published.

What kinds of training data does a humanoid robot use?

Humanoid teams combine four kinds of data: web-scale data, human video, simulation or synthetic data, and real robot data. Each one teaches something the others can't, and they are usually used in that order, from broad to specific.

NVIDIA's GR00T N1 paper describes this as a "data pyramid." Web data and human videos form the base, synthetic data from physics simulation or neural models forms the middle, and data collected on real robot hardware sits at the top (NVIDIA, GR00T N1). The paper says the lower layers give broad visual and behavioral priors, while the top layer grounds the model in real robot execution.

Agility Robotics frames it in terms of where the data comes from. It names three sources: generating data itself through teleoperation, simulation, and public internet data, which it notes does not yet exist for robot movement (Agility, Agility and AI).

  • Web-scale data: teaches language, objects and general visual understanding, usually through a pretrained vision-language model.
  • Human video: teaches how people move through tasks, handle objects and adapt to new rooms.
  • Simulation and synthetic data: teaches balance and control, and multiplies a small number of demonstrations.
  • Robot data: teaches the model its own body, cameras and joints, and grounds everything else in real execution.

Why does a humanoid need its own robot data?

Because only robot data records what the robot's own body did. Human video shows a hand closing on a mug; robot data shows which joints moved, how hard the hand pushed, and what the robot's cameras saw while it happened.

Agility describes its teleoperation setup plainly. Engineers wear VR headsets, see through Digit's cameras and use two controllers to show it a task. While they do, Digit records camera feeds, joint positions, force readings and end-effector location (Agility, Agility and AI).

This is still how many humanoid models start. Figure's first Helix model was trained on about 500 hours of teleoperated data collected across multiple robots and operators. That data drove a 35-degree-of-freedom upper body at 200 Hz, from individual fingers to head and torso (Figure, Helix).

Some companies are scaling robot-side collection directly. Apptronik opened an expanded, nearly 90,000-square-foot Robot Park in Austin in June 2026, where Apollo 2 robots collect data through a mix of teleoperation and autonomous work. Similar workflows run at Google DeepMind and at customers including Mercedes-Benz and GXO, and the data trains Google DeepMind's Gemini Robotics models (Apptronik).

The limit is speed. Agility says teleoperation is the most expensive way to get data and is bound by space and time, because a person has to drive a real robot for every hour.

What does simulation teach a humanoid?

Simulation is best at the parts of control that are easy to score and dangerous to practice on real hardware, like staying balanced. It can also turn a handful of demonstrations into thousands of training examples.

Agility's whole-body controller for Digit is a small neural network with fewer than one million parameters. It is trained in NVIDIA's Isaac Sim for decades of simulated time over three or four days, is learned purely in simulation, and transfers to the real robot without extra training (Agility, Whole-Body Control Foundation Model). Agility says it can then learn dexterous manipulation skills on top of that layer.

NVIDIA used simulation to multiply data instead. Its team generated 780,000 simulated trajectories, which it says equal 6,500 hours of human demonstration data, in 11 hours. It also used video-generation models to turn 88 hours of in-house teleoperation data into about 827 hours of new video, roughly a 10x increase (NVIDIA, GR00T N1).

Figure uses simulation too. The fast control part of the original Helix used a vision backbone pretrained entirely in simulation (Figure, Helix).

Why are humanoid companies pretraining on human video?

Because human video can be collected far faster than robot data, and published results now show it transfers to humanoids. It gives a model broad experience of real homes and workplaces before it ever touches a robot.

Figure's Helix 2.5 is the clearest test. Figure trained two policies on identical task data. One started from scratch; the other was pretrained on Index, Figure's dataset of human behavior. In blind trials across 30 Bay Area homes the robot had never seen, the scratch policy succeeded 9% of the time and the Index-pretrained policy 56% of the time (Figure, Helix 2.5).

Figure reports two more effects. Helix 2.5 needed half as much task-specific data as an earlier Helix 02 behavior. And loss fell predictably as Index data doubled, which Figure calls a human-to-humanoid transfer scaling law.

What makes the video useful is its range. Per 1,000 hours, Figure says Index contains 373 unique tasks, 1,146 unique manipulated objects and 116 unique environments. By September 2026, Figure said Index was generating roughly 35 minutes of new human experience every second (Figure, Introducing Index; Figure, Helix 2.5).

1X uses human video differently, inside a world model for its NEO humanoid. It starts from a 14-billion-parameter video model, trains on 900 hours of egocentric human video, then fine-tunes on 70 hours of robot data. In 1X's tests, adding the egocentric human data improved video generation quality on new tasks (1X, World Model).

What can't human video teach a humanoid robot?

Human video can't show a robot its own body. It carries no joint positions or force readings, and a human hand is not a robot hand, so every company above adapts the model with robot data afterward.

1X states the reason directly: its 70 hours of robot data adapt the model to NEO's visual appearance and kinematics. Figure still trains each Helix 2.5 behavior on task-specific data, just less of it. NVIDIA keeps real robot data at the top of its pyramid for the same reason.

There are coverage limits too. 1X notes that its NEO post-training data was 98.5% pick-and-place. Human video can widen what a model has seen, but the robot-side data still defines what the robot has practiced.

This is why human demonstration data collection is a pretraining and coverage tool, not a replacement for robot data. The open question is how far the robot share can shrink. We'll compare the two sources directly in a future post on human video versus teleoperation.

How much humanoid robot training data do you need?

There is no single number. The published amounts vary by orders of magnitude, depending on whether a team is pretraining a foundation model or adapting one to a task.

Here is what the companies above report:

  • Figure Helix (Feb 2025): about 500 hours of teleoperated robot data, which Figure says is under 5% of the size of previous VLA datasets (Figure).
  • 1X world model (Jan 2026): 900 hours of egocentric human video plus 70 hours of robot data, on top of a video model trained on web-scale video (1X).
  • NVIDIA GR00T N1 (Mar 2025): 88 hours of in-house teleoperation expanded to about 827 hours with generated video, plus 780,000 simulated trajectories (NVIDIA).
  • Figure Index (Sep 2026): about 35 minutes of new human experience every second, used to pretrain Helix 2.5 from random initialization (Figure).
  • Agility motor cortex (Aug 2025): decades of simulated time, generated in three or four days (Agility).

The practical takeaway: robot data budgets in these reports run from tens to hundreds of hours, while Figure's Index rate works out to roughly 50,000 hours of human video a day. Most teams won't build the pretraining layer themselves. They need targeted data for the environments, objects and tasks their model handles badly.

How should you plan robot training data collection?

Start from the gap, not the volume. Work out which layer your model is weak in, then collect for that layer.

  • Weak at control or balance: look at simulation and reinforcement learning first.
  • Weak in new rooms or with new objects: add human video with more environments, objects and task variations per hour.
  • Weak at a specific task on your robot: collect teleoperated demonstrations on your own hardware.
  • Weak at rare situations: collect the edge cases and failures on purpose rather than waiting for them.

Quality checks matter at every layer. Figure runs its human video through five stages before training: filtering, fraud review, deduplication, rebalancing and annotation. It drops clips too similar to data it already accepted (Figure, Introducing Index). Embodied AI training data is only as useful as the pipeline that cleans and labels it.

Where does Firsthand fit, and what can't we do?

Firsthand, CloudPano's AI training data team, collects the human video layer. We run egocentric video data collection to a written spec, at three capture tiers: a phone in a chest mount, a head-mounted rig with an IMU, or synchronized head, wrist and depth cameras on one hardware clock. Annotation layers include actions, hand pose, objects and transcripts, and we deliver in formats such as RLDS and LeRobot. Faces, plates and screens are blurred and checked by a reviewer before delivery.

Here is what we can't do. Human demonstration is our default, so we don't supply the robot-side joint and force data that only your own hardware can record; if you need teleoperated robot trajectories, that has to be scoped as its own program. We don't offer simulation or synthetic data. And the results in this article (Figure's 9% to 56%, 1X's ablations, NVIDIA's data pyramid) were measured on those companies' own data and models. None of them used Firsthand footage, so they show what human video can do, not what our data will do for your robot.

The bottom line

A humanoid robot needs more than one kind of data. Web data gives it general understanding, human video gives it broad real-world experience, simulation teaches it balance and control, and robot data teaches it its own body. The newest results show human video is the layer that scales, but every published recipe still ends on the robot. If you are planning robot training data collection, find the layer where your model is weakest and spend there first.

🚀 Your All‑In‑One Virtual Experience Stack
🎬
PhotoAIVideo
Turn photos into scroll‑stopping AI videos.
Get Started →
🏡
Pictastic
Instantly stage listings with AI.
Try Staging →
🌀
CloudPano
Create stunning 360° tours in minutes.
Launch Tour →
💰
VirtualTourProfit
Build a profitable virtual tour business.
Learn More →
🤝
CloudPano Reseller
Resell AI visual software without building it.
Become a Reseller →
📹
iFirstHand
Custom first‑person video & sensor data for AI & robotics.
Get Data →
🏗️
AI Floor Plan Builder
Generate detailed floor plans with AI.
Build Now →
📐
3D Measure
Capture accurate floor plans & 3D measurements.
Measure Now →
🧠
AI Training Data
Custom AI training data services.
Learn More →

Frequently Asked Questions

How are humanoid robots being trained?

Most humanoids are trained in layers. Teams start from models pretrained on web data, add human video for broad experience, use simulation to learn balance and control, then fine-tune on robot demonstrations collected by teleoperation. Figure, 1X, Agility and Apptronik have each described a version of this mix.

How much training data is enough?

It depends on the job. Adapting a pretrained model to one robot can take tens to hundreds of hours: 1X used 70 hours of robot data, and Figure's first Helix used about 500. Pretraining a foundation model takes far more, which is why Figure collects human video continuously.

How do robots collect data?

Robots record their own camera feeds, joint positions and force readings while a person drives them by teleoperation or while they work on their own. Apptronik runs fleets doing both at its Robot Park. Many teams also add human video recorded by people, not robots.

How to collect data for AI training?

Start with a written spec: the tasks, environments, objects and labels your model needs. Collect against it, then filter for quality, remove duplicates and add annotations. For robotics, plan which mix of human video, simulation and robot demonstrations fills your model's actual gaps.

Can you give me an example of human-generated data?

A first-person video of someone folding towels, filmed from a head-mounted camera, is human-generated data. For robotics it is often labeled with hand poses, objects and actions. Figure's Index dataset and the egocentric video 1X used for its world model are both examples.

Sources

Share this post
Cloudpano

Choose The Right 360° Camera

Insta360 ONE RS 1-Inch 360 Edition

  • Compact, ready to go anywhere

  • Interchangeable lens that’s upgradeable

  • Dual 1-inch sensors for improved clarity and low light performance

  • Dynamic range and 6K 360° capture

  • 360° photo resolution at 21MP

Learn More

Insta360 X4

  • 8K 360° video recording for ultra-detailed visuals.

  • 4K single-lens mode for traditional wide-angle shots.

  • Invisible selfie stick effect for drone-like perspectives.

  • 2.5-inch touchscreen with Gorilla Glass protection.

  • Waterproof up to 33ft for underwater shooting.

Learn More

Ricoh Theta Z1

  • 360° photo resolution in 23MP

  • Slim design at 24 mm thick

  • Built-in image stabilization for smooth video capture.

  • Internal 19GB storage for photo and video storage.

  • Wireless connectivity for remote control and sharing.

Learn More

Ricoh Theta X

  • 60MP 360° still images for high-resolution photography.

  • 5.7K 360° video recording at 30fps.

  • 2.25-inch touchscreen for intuitive control.

  • USB Type-C port for fast charging and data transfer.

  • MicroSD card slot for expandable storage.

Learn More
Property Marketing
Allows potential buyers to explore properties in detail from anywhere, enhancing the real estate marketing process.
Automotive Spins
Create an interactive virtual showroom and engage affluent digital buyers with live 360º video calls, all through the CloudPano mobile app for a complete automotive sales solution.
Interactive Floor Plans
Create 2D and 3D floor plans with measurements in 4 minutes or less, all from your phone. Download the Floor Plan Scanner app and get your first scan free.

360 Virtual Tours With CloudPano.com. Get Started Today.

Try it free. No credit card required. Instant set-up.

Try it free
Latest posts

See our other posts

Interviews, tips, guides, industry best practices, and news.

What Training Data Does a Humanoid Robot Actually Need?

Humanoid robot companies now train on four kinds of data: web data, human video, simulation and robot demonstrations. This guide explains what each layer teaches, what Figure, 1X, Agility, Apptronik and NVIDIA have published about their mixes, and how much data teams report using.
Read post

Why Robotics Labs Are Buying 1,000,000 Hours of Egocentric Video

Robotics labs are now pretraining on a million hours of first-person human video. This guide explains what egocentric video data for robotics is, what Dyna Robotics' 20%-to-53% scaling result shows, why teleoperation can't keep up, and what makes the footage worth buying.
Read post

Why Better Data Matters More Than Bigger Models

Bigger models attract attention, but the data they learn from determines which patterns they can recognize and where they may fail. This guide explains what “better data” means, why dataset quality matters, and how AI teams can improve performance by finding gaps, fixing labels, and evaluating against real-world conditions.
Read post