
A humanoid robot needs layered training data: web and video data for general understanding, egocentric human video for broad manipulation experience, simulation for balance and control, and robot data (teleoperation or autonomous runs) to fit its own body. Figure reports that pretraining on human video raised zero-shot success in 30 unseen homes from 9% to 56%, but each company in this guide still fine-tunes on robot data.
If you are building a humanoid, the hardest question is often not the model but the data mix. In September 2026, Figure reported that pretraining on human video raised its robot's zero-shot success in 30 unseen homes from 9% to 56%, with the robot-specific data held fixed (Figure, Helix 2.5).
That result explains a lot of the current buying spree, which we covered in why robotics labs are buying egocentric video at scale. But it doesn't mean human video is all a humanoid needs. Every company in this guide still trains on its own robot data, and most lean on simulation too. Below is what each layer of humanoid robot training data does, using only what Figure, 1X, Agility, Apptronik and NVIDIA have published.

Humanoid teams combine four kinds of data: web-scale data, human video, simulation or synthetic data, and real robot data. Each one teaches something the others can't, and they are usually used in that order, from broad to specific.
NVIDIA's GR00T N1 paper describes this as a "data pyramid." Web data and human videos form the base, synthetic data from physics simulation or neural models forms the middle, and data collected on real robot hardware sits at the top (NVIDIA, GR00T N1). The paper says the lower layers give broad visual and behavioral priors, while the top layer grounds the model in real robot execution.
Agility Robotics frames it in terms of where the data comes from. It names three sources: generating data itself through teleoperation, simulation, and public internet data, which it notes does not yet exist for robot movement (Agility, Agility and AI).
Because only robot data records what the robot's own body did. Human video shows a hand closing on a mug; robot data shows which joints moved, how hard the hand pushed, and what the robot's cameras saw while it happened.
Agility describes its teleoperation setup plainly. Engineers wear VR headsets, see through Digit's cameras and use two controllers to show it a task. While they do, Digit records camera feeds, joint positions, force readings and end-effector location (Agility, Agility and AI).
This is still how many humanoid models start. Figure's first Helix model was trained on about 500 hours of teleoperated data collected across multiple robots and operators. That data drove a 35-degree-of-freedom upper body at 200 Hz, from individual fingers to head and torso (Figure, Helix).
Some companies are scaling robot-side collection directly. Apptronik opened an expanded, nearly 90,000-square-foot Robot Park in Austin in June 2026, where Apollo 2 robots collect data through a mix of teleoperation and autonomous work. Similar workflows run at Google DeepMind and at customers including Mercedes-Benz and GXO, and the data trains Google DeepMind's Gemini Robotics models (Apptronik).
The limit is speed. Agility says teleoperation is the most expensive way to get data and is bound by space and time, because a person has to drive a real robot for every hour.
Simulation is best at the parts of control that are easy to score and dangerous to practice on real hardware, like staying balanced. It can also turn a handful of demonstrations into thousands of training examples.
Agility's whole-body controller for Digit is a small neural network with fewer than one million parameters. It is trained in NVIDIA's Isaac Sim for decades of simulated time over three or four days, is learned purely in simulation, and transfers to the real robot without extra training (Agility, Whole-Body Control Foundation Model). Agility says it can then learn dexterous manipulation skills on top of that layer.
NVIDIA used simulation to multiply data instead. Its team generated 780,000 simulated trajectories, which it says equal 6,500 hours of human demonstration data, in 11 hours. It also used video-generation models to turn 88 hours of in-house teleoperation data into about 827 hours of new video, roughly a 10x increase (NVIDIA, GR00T N1).
Figure uses simulation too. The fast control part of the original Helix used a vision backbone pretrained entirely in simulation (Figure, Helix).
.png)
Because human video can be collected far faster than robot data, and published results now show it transfers to humanoids. It gives a model broad experience of real homes and workplaces before it ever touches a robot.
Figure's Helix 2.5 is the clearest test. Figure trained two policies on identical task data. One started from scratch; the other was pretrained on Index, Figure's dataset of human behavior. In blind trials across 30 Bay Area homes the robot had never seen, the scratch policy succeeded 9% of the time and the Index-pretrained policy 56% of the time (Figure, Helix 2.5).
Figure reports two more effects. Helix 2.5 needed half as much task-specific data as an earlier Helix 02 behavior. And loss fell predictably as Index data doubled, which Figure calls a human-to-humanoid transfer scaling law.
What makes the video useful is its range. Per 1,000 hours, Figure says Index contains 373 unique tasks, 1,146 unique manipulated objects and 116 unique environments. By September 2026, Figure said Index was generating roughly 35 minutes of new human experience every second (Figure, Introducing Index; Figure, Helix 2.5).
1X uses human video differently, inside a world model for its NEO humanoid. It starts from a 14-billion-parameter video model, trains on 900 hours of egocentric human video, then fine-tunes on 70 hours of robot data. In 1X's tests, adding the egocentric human data improved video generation quality on new tasks (1X, World Model).

Human video can't show a robot its own body. It carries no joint positions or force readings, and a human hand is not a robot hand, so every company above adapts the model with robot data afterward.
1X states the reason directly: its 70 hours of robot data adapt the model to NEO's visual appearance and kinematics. Figure still trains each Helix 2.5 behavior on task-specific data, just less of it. NVIDIA keeps real robot data at the top of its pyramid for the same reason.
There are coverage limits too. 1X notes that its NEO post-training data was 98.5% pick-and-place. Human video can widen what a model has seen, but the robot-side data still defines what the robot has practiced.
This is why human demonstration data collection is a pretraining and coverage tool, not a replacement for robot data. The open question is how far the robot share can shrink. We'll compare the two sources directly in a future post on human video versus teleoperation.
There is no single number. The published amounts vary by orders of magnitude, depending on whether a team is pretraining a foundation model or adapting one to a task.
Here is what the companies above report:
The practical takeaway: robot data budgets in these reports run from tens to hundreds of hours, while Figure's Index rate works out to roughly 50,000 hours of human video a day. Most teams won't build the pretraining layer themselves. They need targeted data for the environments, objects and tasks their model handles badly.
Start from the gap, not the volume. Work out which layer your model is weak in, then collect for that layer.
Quality checks matter at every layer. Figure runs its human video through five stages before training: filtering, fraud review, deduplication, rebalancing and annotation. It drops clips too similar to data it already accepted (Figure, Introducing Index). Embodied AI training data is only as useful as the pipeline that cleans and labels it.
Firsthand, CloudPano's AI training data team, collects the human video layer. We run egocentric video data collection to a written spec, at three capture tiers: a phone in a chest mount, a head-mounted rig with an IMU, or synchronized head, wrist and depth cameras on one hardware clock. Annotation layers include actions, hand pose, objects and transcripts, and we deliver in formats such as RLDS and LeRobot. Faces, plates and screens are blurred and checked by a reviewer before delivery.
Here is what we can't do. Human demonstration is our default, so we don't supply the robot-side joint and force data that only your own hardware can record; if you need teleoperated robot trajectories, that has to be scoped as its own program. We don't offer simulation or synthetic data. And the results in this article (Figure's 9% to 56%, 1X's ablations, NVIDIA's data pyramid) were measured on those companies' own data and models. None of them used Firsthand footage, so they show what human video can do, not what our data will do for your robot.
A humanoid robot needs more than one kind of data. Web data gives it general understanding, human video gives it broad real-world experience, simulation teaches it balance and control, and robot data teaches it its own body. The newest results show human video is the layer that scales, but every published recipe still ends on the robot. If you are planning robot training data collection, find the layer where your model is weakest and spend there first.
Most humanoids are trained in layers. Teams start from models pretrained on web data, add human video for broad experience, use simulation to learn balance and control, then fine-tune on robot demonstrations collected by teleoperation. Figure, 1X, Agility and Apptronik have each described a version of this mix.
It depends on the job. Adapting a pretrained model to one robot can take tens to hundreds of hours: 1X used 70 hours of robot data, and Figure's first Helix used about 500. Pretraining a foundation model takes far more, which is why Figure collects human video continuously.
Robots record their own camera feeds, joint positions and force readings while a person drives them by teleoperation or while they work on their own. Apptronik runs fleets doing both at its Robot Park. Many teams also add human video recorded by people, not robots.
Start with a written spec: the tasks, environments, objects and labels your model needs. Collect against it, then filter for quality, remove duplicates and add annotations. For robotics, plan which mix of human video, simulation and robot demonstrations fills your model's actual gaps.
A first-person video of someone folding towels, filmed from a head-mounted camera, is human-generated data. For robotics it is often labeled with hand poses, objects and actions. Figure's Index dataset and the egocentric video 1X used for its world model are both examples.

Compact, ready to go anywhere
Interchangeable lens that’s upgradeable
Dual 1-inch sensors for improved clarity and low light performance
Dynamic range and 6K 360° capture
360° photo resolution at 21MP

8K 360° video recording for ultra-detailed visuals.
4K single-lens mode for traditional wide-angle shots.
Invisible selfie stick effect for drone-like perspectives.
2.5-inch touchscreen with Gorilla Glass protection.
Waterproof up to 33ft for underwater shooting.

360° photo resolution in 23MP
Slim design at 24 mm thick
Built-in image stabilization for smooth video capture.
Internal 19GB storage for photo and video storage.
Wireless connectivity for remote control and sharing.

60MP 360° still images for high-resolution photography.
5.7K 360° video recording at 30fps.
2.25-inch touchscreen for intuitive control.
USB Type-C port for fast charging and data transfer.
MicroSD card slot for expandable storage.
.png)
.png)

Try it free. No credit card required. Instant set-up.

