
Robotics labs are buying egocentric (first-person) video because it is the robot-relevant data that scales, and the first million-hour results show it pays off. Dyna Robotics pretrained a model on more than 1 million hours of human video and watched its average score across 14 robot tasks climb from 20% to 53% as the data grew from 1,000 hours. Robot-collected data can't grow that fast, so labs record people instead.
In August 2026, Dyna Robotics published a model pretrained on more than one million hours of first-person human video — by its own estimate, roughly 170 years of continuous waking experience. About two weeks later, Figure AI said its data app was taking in 30 minutes of video every second.
If you build robot policies, you know why these numbers matter. Every hour of robot demonstration data has to be produced on purpose, by a person driving a robot, in a place where a robot already is. That makes robot data slow, expensive, and narrow. The labs above are betting that real-world video data of people doing ordinary tasks can fill the gap, and they now have results to back the bet.
Egocentric video is footage recorded from the point of view of the person doing a task, usually from a head-mounted or chest-mounted camera. For robotics, it is valuable because it shows hands, objects, and contact from roughly the angle a robot's own cameras would see them.
Most of Dyna's million-hour corpus is head-mounted, first-person recordings of people doing everyday manipulation such as cooking, tidying, folding, and assembling, according to its Dyna-2 report. Dyna then extracts 3D hand-pose tracks from that footage and uses wrist poses and grip signals as stand-in robot actions.
That is the key difference from general video. Egocentric video data for robotics is video plus enough structure (hand pose, task labels, object labels) for a model to learn what the hands did, not just what the scene looked like.
For scale: Ego4D, the best-known research dataset in this space, holds 3,670 hours of video from 923 participants in 9 countries. Dyna's corpus is about 270 times larger.
Because Dyna showed that robot performance keeps improving all the way up to a million hours of human video, with no robot data in pretraining. That is the first published scaling law of its kind at this size.
Dyna built nested training sets of exactly 1,000, 10,000, 100,000, and 1,000,000 hours, so each bigger set only added data. It pretrained one model per set on human video only, then fine-tuned each on the same small robot dataset (at most 10 hours per task) and tested them blind on 14 real tasks.

Two details stand out. First, some skills appeared only at the top rung: no checkpoint below a million hours turned the lockbox key. Second, the bottle cap task was fine-tuned on roughly 10 minutes of robot demonstrations, yet still improved as human video grew.
Dyna also compared its new model against its previous one, Dyna-1, at customer sites neither model had seen. Both passed close to 100% in-house. On site, Dyna-1 passed 46% of production checks and Dyna-2 passed 87%.
Not every task improved in a straight line. Rope tying peaked at 100,000 hours, and food scooping dropped at the top rung. The average trend is clear, but it is an average.
They do, but teleoperation can't grow fast enough on its own. A person has to drive a robot in real time, in a place where the robot is installed, so hours accumulate slowly and environments repeat.
Skild AI put the case bluntly in a January 2026 post: even a global workforce running robots around the clock could not reach language-model scale through teleoperation alone. Skild says it can teach its model new skills from video plus less than an hour of robot data.
An independent study points the same way. In HumanScale (arXiv, June 2026), researchers pretrained models on equal amounts of egocentric human video or teleoperated robot data. The egocentric models showed 24% lower validation loss on real-robot action prediction and 52.5% and 90% higher success on in-distribution and out-of-distribution robot tasks.

The clearest buyers are robot foundation-model companies that have published their data volumes. Some build collection pipelines in-house; Dyna says its corpus comes from data partners as well as its own operation.

Volume gets the headlines, but both Dyna and Figure describe heavy filtering and labeling before any video reaches a model. Embodied AI training data is only as good as the pipeline it passes through.
What the two companies say they check:
There is one encouraging finding for unlabeled footage. Dyna reports that adding video with no action labels, used only to teach the model to predict what happens next, still improved results on robot data the model had never seen. Dyna's summary: video is a new scaling axis.
You probably don't need to. The million-hour results come from pretraining a foundation model; most teams fine-tune one, or need targeted data for a gap their model has.
The same reports show how little task data the final step can take. Dyna fine-tuned bottle-cap untwisting on about 10 minutes of robot demonstrations. Skild reports learning new skills with less than an hour of robot data. Where outside data helps most is coverage: environments, objects, and people your current data doesn't include.
That is where robotics training data collection services fit. At Firsthand, CloudPano's AI training data team, we run egocentric video data collection to a written spec, from a chest-mounted phone up to synchronized head, wrist, and depth cameras, with faces, plates, and screens blurred before delivery.
What we can't do is promise you Dyna's results. The scaling numbers in this article are Dyna's, measured on Dyna's data and models; Firsthand footage wasn't part of those studies. Phone-tier capture is the fastest way to reach broad geography, but if you need gaze or hand-pose-grade signals, plan for the head-rig or multi-camera tier instead.
Robotics labs are buying egocentric video by the million hours because it is the first robot-relevant data source shown to keep improving robot performance at that scale. Dyna's 20%-to-53% result is the evidence; Figure's $1B data-and-compute commitment is the market response. The open question for most teams isn't whether human video helps, but which footage closes their model's specific gaps — and whether it is clean, diverse, and labeled well enough to use.
Egocentric data is information recorded from the viewpoint of the person doing a task, usually first-person video from a head- or chest-mounted camera. Robotics teams use it to teach models how hands and objects interact, often extracting hand poses as stand-in robot actions.
An egocentric dataset is a collection of first-person recordings, usually labeled with actions, objects, or hand poses. Ego4D, the best-known research example, holds 3,670 hours of video. Dyna Robotics pretrained its Dyna-2 model on more than one million hours.
It means recruiting people to film real tasks from their own point of view, following a written spec. The footage is checked for quality and consent, bystanders are blurred, and labels are added before it is delivered as training data.
Embodied AI is AI that perceives and acts in the physical world through a body, such as a robot arm or humanoid. Because it handles real objects, its training data must show physical interaction, which is why labs pretrain on human manipulation video.
Robots collect data through their own cameras and sensors, and through teleoperation, where a person drives the robot to record demonstrations. Many labs now add human video too. Figure AI, for example, pays app users to film everyday tasks.

Compact, ready to go anywhere
Interchangeable lens that’s upgradeable
Dual 1-inch sensors for improved clarity and low light performance
Dynamic range and 6K 360° capture
360° photo resolution at 21MP

8K 360° video recording for ultra-detailed visuals.
4K single-lens mode for traditional wide-angle shots.
Invisible selfie stick effect for drone-like perspectives.
2.5-inch touchscreen with Gorilla Glass protection.
Waterproof up to 33ft for underwater shooting.

360° photo resolution in 23MP
Slim design at 24 mm thick
Built-in image stabilization for smooth video capture.
Internal 19GB storage for photo and video storage.
Wireless connectivity for remote control and sharing.

60MP 360° still images for high-resolution photography.
5.7K 360° video recording at 30fps.
2.25-inch touchscreen for intuitive control.
USB Type-C port for fast charging and data transfer.
MicroSD card slot for expandable storage.
.png)
.png)

Try it free. No credit card required. Instant set-up.

