Methods · objection handler
“Why not just teleoperate?”
It is the first question a robotics team asks about egocentric data, and it deserves a straight answer rather than a brochure. This page collects what the published literature actually reports, names every source, and states what each one says it does not show.
If human egocentric video is a good training substrate, that is true for every company collecting it — including the best-funded ones, who have far more of it than we do. Nothing on this page is a reason to choose FourthHuman. It exists to clear one objection so we can have the conversation that actually matters, which is what we emit and whether you can verify where it came from.
What the literature reports
None of these are our measurements. We have not reproduced any of them. Two of the four are peer-reviewed; two are recent preprints, marked as such. Every figure below belongs to the paper beside it.
EgoMimic
Kareer et al. · ICRA 2025 · arXiv:2410.24221
Peer-reviewed
Scaling Imitation Learning via Egocentric Video. Reports that adding one hour of human hand data was more valuable than adding one hour of robot data, across three real-world manipulation tasks.
The mechanism is throughput, and it is worth being precise about: an hour of human capture yielded roughly 1,400 demonstrations against roughly 135 from an hour of teleoperation — EgoMimic's figures, not ours. That is a collection-rate difference, not a claim that one human demonstration is worth more than one robot demonstration.
EMMA
Zhu et al. · IEEE RA-L 2026 · arXiv:2509.04443
Peer-reviewed
Scaling Mobile Manipulation via Egocentric Human Data. Trained mobile-manipulation policies from human egocentric data co-trained with static robot teleoperation data, avoiding mobile teleoperation specifically, and reported performance comparable to or better than Mobile ALOHA at equivalent data-collection time.
On the Handover Wine task the human-data policy reached 82% against 52%; on Table Service the two were comparable. Both are EMMA's figures. This does not replace teleoperation — it removes the mobile half of it while still requiring static robot data.
HumanEgo
Wang et al. · 2026 preprint · arXiv:2605.24934
Preprint · not peer-reviewed
Zero-Shot Robot Learning from Minutes of Human Egocentric Videos. Reports 92.5% average success across four real-world tasks from 30 minutes of human video per task — HumanEgo's figure — exceeding a matched-time teleoperation baseline.
The baseline is a policy trained on 30 minutes of teleoperation. This is evidence about the very-low-data regime — nobody claims a half-hour teleop policy is strong — and not evidence that human video beats teleoperation at scale.
HumanScale
Ma et al. · 2026 preprint · arXiv:2606.20521
Preprint · not peer-reviewed
Egocentric Human Video Can Outperform Real-Robot Data for Embodied Pretraining. Argues egocentric human video, given a filtering and labelling pipeline, can outperform real-robot data for pretraining.
We quote no figures from this paper. Reviewing it, we could not establish from the text what its headline comparison is actually against, and it contains no ablation showing the filtering pipeline — on which the whole claim is conditioned — was necessary. We cite the argument and withhold the numbers.
What this does not prove
These are the authors’ own words, not our caveats. A page that cited only the favourable half would be advertising.
The embodiment gap is real and unsolved. HumanEgo states transfer “remains challenging due to the embodiment gap between human and robot in both visual appearance and kinematics”. EMMA’s first stated limitation is that its approach breaks down “when the kinematic or viewpoint differences between human and robot embodiments become substantial”.
Coordination-heavy tasks still need robot data. EMMA: “the direct transfer of human navigation strategies is insufficient for tasks requiring tightly coupled arm–base coordination.”
Precision has a floor. HumanEgo: “few-shot learning plateaus at ∼1 cm precision; reaching sub-centimeter accuracy on contact-rich tasks will likely require reinforcement-learning refinement or simulation-based fine-tuning.”
The perception stack is a chain of failure points. HumanEgo: “The pipeline chains several off-the-shelf perception modules whose failures cascade into the policy,” and it depends on stereo hand tracking — “monocular substitutes drop real-world success sharply.”
Two of these four have not been refereed. HumanEgo and HumanScale are preprints published within the last ninety days. Neither has been through review, and neither should be weighted like the two that have.
Even the favourable papers still need robot data. HumanScale’s own proposed method is to pretrain on egocentric video and then “adapt with a small amount of labeled real-robot data for action-space alignment”. None of this work claims robot data becomes unnecessary.
The case against
We are an egocentric data company, so this literature is commercially convenient for us. That makes it more important to publish what cuts the other way, not less.
Human capture records what was observed, not what a controller should do. Surendran et al. (arXiv:2606.26603) find that handheld human-style capture yields trajectories that “lack dynamic feasibility in contact-sensitive phases, where tracking observed trajectories at high stiffness produces large, unsafe contact forces” — and that naively mixing such data with teleoperated data performed worse than the handheld data alone. Their study is of handheld rigs rather than head-mounted capture; we think the concern extends, but that extension is our inference and not their finding.
The advantage is scoped to the low-robot-data regime. Li et al. (arXiv:2606.06627) — themselves reporting a gain from co-training on human video — state it as “an absolute success rate gain of 29.7% in the low-robot-data regime”, and find that “even with accurate hands, the inherent motion gap hinders transfer unless the vision and policy networks specialize to each embodiment.”
Large-scale visual pretraining has failed scrutiny before. Hansen et al. (arXiv:2212.05749, ICML 2023) found a learn-from-scratch baseline “surprisingly competitive” with methods built on frozen representations from large vision corpora. Older, and about frozen visual features rather than action-labelled pretraining — but a standing reminder that “more video” has not always meant better control.
Where we think this lands
Searching for a reproduction failure or a direct rebuttal of these results, we found none — the 2026 literature does lean this way. But the defensible reading is narrower than the headlines: the advantage is real, concentrated where robot data is scarce, weakest on contact-rich force-sensitive tasks, and every leading result still requires robot action data for action-space alignment. “Egocentric video replaces teleoperation” is not a claim this literature supports, and it is not one we make.
So why us?
Nothing above answers that, and it is not supposed to. If the category argument holds, it holds for everyone in it. Two things are ours specifically, and both are things you can check rather than take on trust:
The reference implementation
One real clip, every schema layer, downloadable. Published with its known failure modes and what it does not prove.
The consent ledger
Per-clip provenance you can audit yourself: the hash pre-image is published, so you can recompute it rather than trust a badge.
Citations verified against primary sources 13-08-2026. Every figure on this page belongs to the cited paper and is not a FourthHuman measurement.