Why Robotics Labs Are Buying 1,000,000 Hours of Egocentric Video

CloudPano
September 30, 2026
•
5 min read
Share this post

Why are robotics labs buying 1,000,000 hours of egocentric video?

Robotics labs are buying egocentric (first-person) video because it is the robot-relevant data that scales, and the first million-hour results show it pays off. Dyna Robotics pretrained a model on more than 1 million hours of human video and watched its average score across 14 robot tasks climb from 20% to 53% as the data grew from 1,000 hours. Robot-collected data can't grow that fast, so labs record people instead.

Key Takeaways

  • Dyna Robotics' Dyna-2 was pretrained on 1M+ hours of egocentric video. Its average robot task score rose 20% → 28% → 45% → 53% as pretraining grew from 1,000 to 1,000,000 hours.
  • Some skills only appear at scale: no Dyna-2 checkpoint up to 100,000 hours could turn a lockbox key; the 1M-hour model did it 90% of the time.
  • Figure AI built its own collection app, Index, after it says vendors couldn't meet its throughput, diversity, and quality needs. It has paid Creators $15M and committed to spend over $1B on data and compute in the next 12 months.
  • Teleoperation gives precise robot actions but is slow to collect. Skild AI and the June 2026 HumanScale study both point to human video as the scalable, more diverse pretraining source.
  • Raw hours aren't the product. Labs filter, deduplicate, and extract hand poses, and they pay for diversity: new homes, objects, and ways of doing a task.

‍

Why Robotics Labs Are Buying 1,000,000 Hours of Egocentric Video

In August 2026, Dyna Robotics published a model pretrained on more than one million hours of first-person human video — by its own estimate, roughly 170 years of continuous waking experience. About two weeks later, Figure AI said its data app was taking in 30 minutes of video every second.

If you build robot policies, you know why these numbers matter. Every hour of robot demonstration data has to be produced on purpose, by a person driving a robot, in a place where a robot already is. That makes robot data slow, expensive, and narrow. The labs above are betting that real-world video data of people doing ordinary tasks can fill the gap, and they now have results to back the bet.

What is egocentric video data for robotics?

Egocentric video is footage recorded from the point of view of the person doing a task, usually from a head-mounted or chest-mounted camera. For robotics, it is valuable because it shows hands, objects, and contact from roughly the angle a robot's own cameras would see them.

Most of Dyna's million-hour corpus is head-mounted, first-person recordings of people doing everyday manipulation such as cooking, tidying, folding, and assembling, according to its Dyna-2 report. Dyna then extracts 3D hand-pose tracks from that footage and uses wrist poses and grip signals as stand-in robot actions.

That is the key difference from general video. Egocentric video data for robotics is video plus enough structure (hand pose, task labels, object labels) for a model to learn what the hands did, not just what the scene looked like.

For scale: Ego4D, the best-known research dataset in this space, holds 3,670 hours of video from 923 participants in 9 countries. Dyna's corpus is about 270 times larger.

Why is 1,000,000 hours the number everyone is talking about?

Because Dyna showed that robot performance keeps improving all the way up to a million hours of human video, with no robot data in pretraining. That is the first published scaling law of its kind at this size.

Dyna built nested training sets of exactly 1,000, 10,000, 100,000, and 1,000,000 hours, so each bigger set only added data. It pretrained one model per set on human video only, then fine-tuned each on the same small robot dataset (at most 10 hours per task) and tested them blind on 14 real tasks.

Average score is each task's result as a share of its maximum, averaged across tasks. Source: Dyna Robotics, Dyna-2 report, August 2026.

‍

Two details stand out. First, some skills appeared only at the top rung: no checkpoint below a million hours turned the lockbox key. Second, the bottle cap task was fine-tuned on roughly 10 minutes of robot demonstrations, yet still improved as human video grew.

Dyna also compared its new model against its previous one, Dyna-1, at customer sites neither model had seen. Both passed close to 100% in-house. On site, Dyna-1 passed 46% of production checks and Dyna-2 passed 87%.

Not every task improved in a straight line. Rope tying peaked at 100,000 hours, and food scooping dropped at the top rung. The average trend is clear, but it is an average.

Why don't labs just collect more teleoperation data?

They do, but teleoperation can't grow fast enough on its own. A person has to drive a robot in real time, in a place where the robot is installed, so hours accumulate slowly and environments repeat.

Skild AI put the case bluntly in a January 2026 post: even a global workforce running robots around the clock could not reach language-model scale through teleoperation alone. Skild says it can teach its model new skills from video plus less than an hour of robot data.

An independent study points the same way. In HumanScale (arXiv, June 2026), researchers pretrained models on equal amounts of egocentric human video or teleoperated robot data. The egocentric models showed 24% lower validation loss on real-robot action prediction and 52.5% and 90% higher success on in-distribution and out-of-distribution robot tasks.

The catch is real. Skild notes that video carries no forces, torques, or touch, and a human hand moves differently from a robot arm. Most labs therefore pretrain on human video and fine-tune on a small amount of robot data, which is the recipe Dyna used.

‍

Who is buying egocentric video at this scale?

The clearest buyers are robot foundation-model companies that have published their data volumes. Some build collection pipelines in-house; Dyna says its corpus comes from data partners as well as its own operation.

Figure's note is the most telling for anyone selling data. The company says it tried buying data first, but vendors "couldn't hit the throughput, diversity, or quality bar" its model required. The demand is there; the bar is high.Skild's revenue figure doesn't prove anything about video by itself. It does show that robot foundation models are now earning real money, which funds the data budgets behind this trend.

What makes a million hours of video actually useful?

Volume gets the headlines, but both Dyna and Figure describe heavy filtering and labeling before any video reaches a model. Embodied AI training data is only as good as the pipeline it passes through.

What the two companies say they check:

  • Hand-pose quality. Dyna keeps only episodes whose 3D hand tracks pass its quality bar, and says not every recording setup produces hand poses that do.
  • Fraud and duplicates. Figure's pipeline runs five stages: filtering, fraud review, deduplication, rebalancing, and annotation. It discards clips too similar to data it already accepted.
  • Diversity per hour. Figure reports that every 1,000 hours of Index data contains 373 unique tasks, 1,146 unique manipulated objects, and 116 unique environments.
  • Labels. Figure generates layered text captions for every episode; Dyna derives wrist trajectories and grip signals from hand pose.

There is one encouraging finding for unlabeled footage. Dyna reports that adding video with no action labels, used only to teach the model to predict what happens next, still improved results on robot data the model had never seen. Dyna's summary: video is a new scaling axis.

What does this mean if you can't collect a million hours?

You probably don't need to. The million-hour results come from pretraining a foundation model; most teams fine-tune one, or need targeted data for a gap their model has.

The same reports show how little task data the final step can take. Dyna fine-tuned bottle-cap untwisting on about 10 minutes of robot demonstrations. Skild reports learning new skills with less than an hour of robot data. Where outside data helps most is coverage: environments, objects, and people your current data doesn't include.

That is where robotics training data collection services fit. At Firsthand, CloudPano's AI training data team, we run egocentric video data collection to a written spec, from a chest-mounted phone up to synchronized head, wrist, and depth cameras, with faces, plates, and screens blurred before delivery.

What we can't do is promise you Dyna's results. The scaling numbers in this article are Dyna's, measured on Dyna's data and models; Firsthand footage wasn't part of those studies. Phone-tier capture is the fastest way to reach broad geography, but if you need gaze or hand-pose-grade signals, plan for the head-rig or multi-camera tier instead.

The bottom line

Robotics labs are buying egocentric video by the million hours because it is the first robot-relevant data source shown to keep improving robot performance at that scale. Dyna's 20%-to-53% result is the evidence; Figure's $1B data-and-compute commitment is the market response. The open question for most teams isn't whether human video helps, but which footage closes their model's specific gaps — and whether it is clean, diverse, and labeled well enough to use.

‍

‍

‍

Frequently Asked Questions

What is egocentric data in robotics?

Egocentric data is information recorded from the viewpoint of the person doing a task, usually first-person video from a head- or chest-mounted camera. Robotics teams use it to teach models how hands and objects interact, often extracting hand poses as stand-in robot actions.

What is an egocentric dataset?

An egocentric dataset is a collection of first-person recordings, usually labeled with actions, objects, or hand poses. Ego4D, the best-known research example, holds 3,670 hours of video. Dyna Robotics pretrained its Dyna-2 model on more than one million hours.

What does egocentric video data collection mean?

It means recruiting people to film real tasks from their own point of view, following a written spec. The footage is checked for quality and consent, bystanders are blurred, and labels are added before it is delivered as training data.

What exactly is embodied AI?

Embodied AI is AI that perceives and acts in the physical world through a body, such as a robot arm or humanoid. Because it handles real objects, its training data must show physical interaction, which is why labs pretrain on human manipulation video.

How do robots collect data?

Robots collect data through their own cameras and sensors, and through teleoperation, where a person drives the robot to record demonstrations. Many labs now add human video too. Figure AI, for example, pays app users to film everyday tasks.

Sources

Share this post
CloudPano

Choose The Right 360° Camera

Insta360 ONE RS 1-Inch 360 Edition

  • Compact, ready to go anywhere

  • Interchangeable lens that’s upgradeable

  • Dual 1-inch sensors for improved clarity and low light performance

  • Dynamic range and 6K 360° capture

  • 360° photo resolution at 21MP

Learn More

Insta360 X4

  • 8K 360° video recording for ultra-detailed visuals.

  • 4K single-lens mode for traditional wide-angle shots.

  • Invisible selfie stick effect for drone-like perspectives.

  • 2.5-inch touchscreen with Gorilla Glass protection.

  • Waterproof up to 33ft for underwater shooting.

Learn More

Ricoh Theta Z1

  • 360° photo resolution in 23MP

  • Slim design at 24 mm thick

  • Built-in image stabilization for smooth video capture.

  • Internal 19GB storage for photo and video storage.

  • Wireless connectivity for remote control and sharing.

Learn More

Ricoh Theta X

  • 60MP 360° still images for high-resolution photography.

  • 5.7K 360° video recording at 30fps.

  • 2.25-inch touchscreen for intuitive control.

  • USB Type-C port for fast charging and data transfer.

  • MicroSD card slot for expandable storage.

Learn More
Property Marketing
Allows potential buyers to explore properties in detail from anywhere, enhancing the real estate marketing process.
Automotive Spins
Create an interactive virtual showroom and engage affluent digital buyers with live 360º video calls, all through the CloudPano mobile app for a complete automotive sales solution.
Interactive Floor Plans
Create 2D and 3D floor plans with measurements in 4 minutes or less, all from your phone. Download the Floor Plan Scanner app and get your first scan free.

360 Virtual Tours With CloudPano.com. Get Started Today.

Try it free. No credit card required. Instant set-up.

Try it free
Latest posts

See our other posts

Interviews, tips, guides, industry best practices, and news.

Why Robotics Labs Are Buying 1,000,000 Hours of Egocentric Video

Robotics labs are now pretraining on a million hours of first-person human video. This guide explains what egocentric video data for robotics is, what Dyna Robotics' 20%-to-53% scaling result shows, why teleoperation can't keep up, and what makes the footage worth buying.
Read post

Why Better Data Matters More Than Bigger Models

Bigger models attract attention, but the data they learn from determines which patterns they can recognize and where they may fail. This guide explains what “better data” means, why dataset quality matters, and how AI teams can improve performance by finding gaps, fixing labels, and evaluating against real-world conditions.
Read post

How to Turn Listing Photos Into Videos With an API

Learn how to turn listing photos into property videos with an API. This guide explains the complete workflow, including photo selection, image uploads, AI video generation, asynchronous job processing, clip assembly, branding, output formats, error handling, and automated real estate video workflows.
Read post