Egocentric capture · India
The physical world is humanity’s largest dataset.
We build the infrastructure that makes it learnable — captured in India to your brief, annotated before it reaches you, and every clip carries a consent receipt whose hash you can recompute yourself.
Why this is hard
AI learned from the internet. The physical world was never on it.
A model can read every sentence ever published and watch every video ever uploaded, and still not know how much force a drawer needs. Physical intelligence is not written down anywhere — it is performed, by people, and then it is gone.
Language models
Billions of words
Already collected, already public. The corpus existed before the models did.
Vision models
Billions of frames
Images and video scraped from an internet that was already there. Again, found rather than made.
Physical AI
No corpus
Nothing equivalent exists. Egocentric capture of ordinary human work has to be recorded on purpose, by someone, with consent.
No figure is given for the third column because there is no measured one to give — not ours and not anyone’s. The asymmetry is the point.
Check it yourself
This is the whole mechanism. There is nothing else behind it.
Every capture session writes one row into an append-only ledger. Eleven fields join with a pipe, that string is hashed, and the hash of the row before it is one of the eleven — so changing any earlier row breaks every hash after it. The record below is an example, because a real one names a real person. The digest is not an example: it is produced by the same function that runs in production.
The eleven fields, in the order they are joined
- loading the published example…
The pre-image — those eleven, joined by |
loading…
SHA-256 of that string
loading…
That is the entire claim, and it does not require us to be honest — only to be consistent, which is checkable. So check it: the button below hashes the pre-image above in your browser, with the Web Crypto API your browser ships, and compares the result to the digest we published.
For Capture Partners
₹250–400 an hour, paid on QA approval, auditable per session.
We pay Capture Partners directly by UPI on QA approval — no agent in the middle, no deduction — and we publish the range instead of quoting it privately per engagement. Those are the terms partners will be paid on, not a payout rail already running: no payment provider is integrated and no payout has ever settled. Every session carries a consent receipt with a SHA-256 audit signature that the Capture Partner can revoke. Onboarding and consent infrastructure are built and testable today; the payout endpoint records an intent to pay and nothing more.
Planned pilot hubs: Bangalore, Mumbai, Delhi-NCR, Hyderabad, Chennai, Kolkata and Pune. That is a rollout plan, not an active Capture Partner count — we do not have one to quote yet.
All seven are in India, and that is where the cost structure we sell comes from — so it is where capture runs today. Other countries are on the roadmap, not in operation. If a brief needs a specific geography, ask: we would rather tell you we cannot source it than have you discover that at delivery.
Method
Provenance you can check, not a badge you have to trust.
Every capture session produces a digital consent receipt persisted with a SHA-256 audit signature, revocable by the Capture Partner under the data-deletion rights in India's DPDP Act 2023. A buyer audits it per clip. That is the difference we sell: competitors sell a badge attesting to a process; we hand you the receipt and the hash that verifies it.
Every clip passes automated face redaction server-side at ingestion, before any human review or delivery. The detector is YuNet — a DNN, chosen because egocentric faces are off-axis, motion-blurred and partial, the angles a classical cascade misses. It is not Meta EgoBlur. The stage fails closed: if the detector cannot load, the clip fails the stage rather than passing through, so there is no unredacted path to delivery. Licence-plate redaction is not implemented — we would rather name the gap than let you assume it is covered.
We hold no third-party attestation: no SOC 2 report, no ISO 27001 certificate. Those require an external audit we have not commissioned, and we will link the report here when we do.
Stage 1
Acquire
A Capture Partner records the brief on a head-mounted camera, in the environment the task actually happens in. Consent is taken before the first frame and written to the ledger.
Stage 2
Enrich
Automated 3D annotation — 21 hand keypoints per hand, temporal action segmentation, dense captions. No manual labelling step, which is why the cost scales with capture rather than with annotators.
Stage 3
Prepare
Quality scoring, multimodal alignment and DPDP-aligned face redaction. Redaction runs server-side at ingestion, before any human sees the footage, and fails closed.
Stage 4
Validate
Export to the format you asked for, then read back by an independent verifier before delivery. A clip with no consent link is never delivered.
Frames in the reference clip
1,843
Published in full, at 30 fps. Every number on this page that describes the engine comes from that one file.
Mean per-hand detection confidence
0.87
Arithmetic mean over the kinematic track. The minimum is published too, on the reference-implementation page.
Delivery formats shipping
5
Each has a writer and an independent verifier that reads the export back before delivery.
Hours of capture delivered
0
None
The Capture Partner network is in pilot. Volume is agreed in a brief, not drawn from stock.
Companies we can name as customers
0
None
We are pre-revenue. There are no logos on this page because there is nobody to put on it.
Third-party attestations held
0
None
We hold no SOC 2 report and no ISO 27001 certificate, and say so rather than implying one.
What we make
Two hands, measured through sixty-one seconds of real work.
Scroll, and the path below draws from the published annotation of our reference clip — one point per sampled frame, both hands, relative to the camera. It is the same file you can download and check, and the readout only ever prints a value that is in it.
Capture Catalogue
One annotation standard, across every domain we scope.
Each domain below is a capture specification you can commission, and each ships against FH-Ego v1 — so a dataloader you write once works across every delivery. Volumes are target volumes agreed in a brief, not stock we hold. Availability is stated per domain, and it is not the same for any two.
Tier 1 · Session
What the clip is
Scene, task label, frame count, frame rate, usability verdict — and the provenance block: clip id, content hash, cohort, QA verdict, and which fields were derived rather than measured.
Tier 2 · Action
What happens, and when
Frame bounds per action segment, hand laterality (left / right / both), scene semantics, camera motion, a verb-noun action code and a dense natural-language description.
Tier 3 · Kinematics
Where the hands are
21 3D joint landmarks per hand, camera-relative wrist position, frame-to-frame velocity, and a per-hand detection confidence — so you can filter on quality rather than trust it.
Bimanual Household Tasks
Chopping, stirring, plating and other two-handed kitchen sequences, specified for capture in Indian household kitchens.
Target volume 42,500 clips / 350 hrs
Pilot hub forming
HealthcareSpecialized Caregiving & Nursing
Elder-care assistance, patient transfers and medical-tool handling, with face redaction applied before any human review.
Target volume 18,200 clips / 150 hrs
Scoped on demand
ManufacturingPrecision Factory Assembly
Garment stitching, component soldering and torque-wrench fastening. This is the domain the published reference clip comes from — a person at an industrial sewing workstation.
Target volume 35,000 clips / 290 hrs
Pipeline validated
WarehouseWarehouse & Logistics
Box lifting, barcode scanning, bin packing and shelf stocking in active distribution centres.
Target volume 50,000 clips / 410 hrs
Pilot hub forming
CleaningCleaning & Maintenance
Floor mopping, window wiping, waste disposal and surface disinfection in commercial premises.
Target volume 22,000 clips / 180 hrs
Scoped on demand
SimulationGame Environment & Simulator Data
Simulated environment telemetry, 3DGS scans and controllable physics for sim-to-real co-training.
Target volume 12,000 scenes / 10,000 hrs
Research phase
Why egocentric
A head-mounted camera sees the task from where the hands are.
Teleoperation gives you a robot's own proprioception, and it costs one person driving one rig for every hour of demonstration it produces. Egocentric capture gives you human hands solving the task in the environment the task actually happens in, and it scales with people rather than with rigs. There is a real case against it — no joint torques, no action labels for free, a domain gap you have to close — and the published work arguing both sides is a category argument, not a FourthHuman result. We set it out on the method page, including the part that does not favour us.
Price
We publish what reaches the person holding the camera.
We pay Capture Partners ₹250–400 an hour, direct by UPI on QA approval, with no agent in the middle and no deduction. We publish that band on a public page rather than quoting it privately per engagement — so it is the one cost inside a FourthHuman quote you can check before you have spoken to us.
What moves the number · 1
The environment
An Indian household kitchen and an active distribution centre are not the same shoot. Access, scheduling and how many people have to consent all change with the room.
What moves the number · 2
The session
How long a session runs, how many of them, and how many Capture Partners the brief needs to reach the volume you are asking for.
What moves the number · 3
The annotation tier
Tier 1 is session metadata. Tier 3 is 21 3D joints per hand with per-hand detection confidence, and it costs more to produce and more to verify. You pick the tier in the brief.
A brief is scoped first and priced against that scope. There is no rate card yet, and we would rather say so than publish a figure the first real engagement would break. The structural reason a bespoke brief is affordable here at all is the India–US wage differential — which is also why we publish our input cost instead of hiding it inside a margin.
FAQ
Frequently asked questions
What kind of training data do you produce?
How do you handle privacy and data protection?
How much do you pay your Capture Partners?
What annotation schema do you use?
What data formats do you deliver datasets in?
lerobot package), HDF5 (EgoDex-style paired .hdf5+.mp4 — ALOHA / RoboMimic-lineage labs), and the open FH-Ego v1 JSON schema. Zarr v2 (UMI / diffusion-policy layout) and TFRecord (RLDS-style step schema; not tfds.load-compatible) are produced on request. Every export is read back by an independent verifier before delivery. We do not offer MCAP or WebDataset — MCAP is a robotics logging container no training dataloader consumes, and WebDataset is an internal sharding optimization, not a distribution format.
What makes you different from a general annotation vendor?
How does your pricing compare to US-based teleoperation data collection?
Get started
Tell us what you’re training.
We’ll scope the dataset.
Describe the policy and the embodiment you are training it for. We scope the capture specification, price it against your volume and format, and come back with what we can commit to — including the parts we cannot. One of the two founders reads every one of these. There is no queue, because there is nobody to queue behind.
What a brief costs is scoped before it is priced — how we price, and what moves the number.