Egocentric capture · India

The physical world is humanity’s largest dataset.

We build the infrastructure that makes it learnable — captured in India to your brief, annotated before it reaches you, and every clip carries a consent receipt whose hash you can recompute yourself.

FRAME 0276 / 1843 · loading…

The annotation engine runs today on a clip we publish in full; the Capture Partner network is in pilot.

Why this is hard

AI learned from the internet. The physical world was never on it.

A model can read every sentence ever published and watch every video ever uploaded, and still not know how much force a drawer needs. Physical intelligence is not written down anywhere — it is performed, by people, and then it is gone.

Language models

Billions of words

Already collected, already public. The corpus existed before the models did.

Vision models

Billions of frames

Images and video scraped from an internet that was already there. Again, found rather than made.

Physical AI

No corpus

Nothing equivalent exists. Egocentric capture of ordinary human work has to be recorded on purpose, by someone, with consent.

No figure is given for the third column because there is no measured one to give — not ours and not anyone’s. The asymmetry is the point.

For Capture Partners

₹250–400 an hour, paid on QA approval, auditable per session.

We pay Capture Partners directly by UPI on QA approval — no agent in the middle, no deduction — and we publish the range instead of quoting it privately per engagement. Those are the terms partners will be paid on, not a payout rail already running: no payment provider is integrated and no payout has ever settled. Every session carries a consent receipt with a SHA-256 audit signature that the Capture Partner can revoke. Onboarding and consent infrastructure are built and testable today; the payout endpoint records an intent to pay and nothing more.

Planned pilot hubs: Bangalore, Mumbai, Delhi-NCR, Hyderabad, Chennai, Kolkata and Pune. That is a rollout plan, not an active Capture Partner count — we do not have one to quote yet.

All seven are in India, and that is where the cost structure we sell comes from — so it is where capture runs today. Other countries are on the roadmap, not in operation. If a brief needs a specific geography, ask: we would rather tell you we cannot source it than have you discover that at delivery.

Method

Provenance you can check, not a badge you have to trust.

Every capture session produces a digital consent receipt persisted with a SHA-256 audit signature, revocable by the Capture Partner under the data-deletion rights in India's DPDP Act 2023. A buyer audits it per clip. That is the difference we sell: competitors sell a badge attesting to a process; we hand you the receipt and the hash that verifies it.

Every clip passes automated face redaction server-side at ingestion, before any human review or delivery. The detector is YuNet — a DNN, chosen because egocentric faces are off-axis, motion-blurred and partial, the angles a classical cascade misses. It is not Meta EgoBlur. The stage fails closed: if the detector cannot load, the clip fails the stage rather than passing through, so there is no unredacted path to delivery. Licence-plate redaction is not implemented — we would rather name the gap than let you assume it is covered.

We hold no third-party attestation: no SOC 2 report, no ISO 27001 certificate. Those require an external audit we have not commissioned, and we will link the report here when we do.

Stage 1

Acquire

A Capture Partner records the brief on a head-mounted camera, in the environment the task actually happens in. Consent is taken before the first frame and written to the ledger.

Stage 2

Enrich

Automated 3D annotation — 21 hand keypoints per hand, temporal action segmentation, dense captions. No manual labelling step, which is why the cost scales with capture rather than with annotators.

Stage 3

Prepare

Quality scoring, multimodal alignment and DPDP-aligned face redaction. Redaction runs server-side at ingestion, before any human sees the footage, and fails closed.

Stage 4

Validate

Export to the format you asked for, then read back by an independent verifier before delivery. A clip with no consent link is never delivered.

Raw frame 150 of the reference clip, from the wearer's viewpoint: the left hand aligns and organizes fabric items on a workstation surface beside a sewing machine — the action the published annotation records for this segment. The same frame with FH-Ego v1 annotation drawn on: the left hand's 21 tracked joints, the connecting skeleton, its wrist triad axes, and a label giving detection confidence and the wrist position vector. Only the left hand is annotated in this frame. A header burnt into the image reads: Egocentric Bimanual Workstation Interaction, frame 150 of 1843, action organize_items.

50% annotated · raw at left, FH-Ego v1 overlay at right

Reference clip, sewing workstation · 21 3D joints per hand, left/right classified · the overlay is the engine's own output, untouched

Frames in the reference clip

1,843

Published in full, at 30 fps. Every number on this page that describes the engine comes from that one file.

Mean per-hand detection confidence

0.87

Arithmetic mean over the kinematic track. The minimum is published too, on the reference-implementation page.

Delivery formats shipping

5

Each has a writer and an independent verifier that reads the export back before delivery.

Hours of capture delivered

0

None

The Capture Partner network is in pilot. Volume is agreed in a brief, not drawn from stock.

Companies we can name as customers

0

None

We are pre-revenue. There are no logos on this page because there is nobody to put on it.

Third-party attestations held

0

None

We hold no SOC 2 report and no ISO 27001 certificate, and say so rather than implying one.

What we make

Two hands, measured through sixty-one seconds of real work.

Scroll, and the path below draws from the published annotation of our reference clip — one point per sampled frame, both hands, relative to the camera. It is the same file you can download and check, and the readout only ever prints a value that is in it.

FRAME  / 1843 · · · speed · conf · loading…

Sampled every tenth frame by the pipeline, so the path is drawn point-to-point between real measurements and nothing between them is invented. Positions and velocities are derived — they are the tracking model's own output, not an instrumented ground truth.

Capture Catalogue

One annotation standard, across every domain we scope.

Each domain below is a capture specification you can commission, and each ships against FH-Ego v1 — so a dataloader you write once works across every delivery. Volumes are target volumes agreed in a brief, not stock we hold. Availability is stated per domain, and it is not the same for any two.

Delivery formats We deliver LeRobot v2.1, HDF5 (EgoDex-style paired files) and the open FH-Ego v1 JSON schema; Zarr v2 and TFRecord (RLDS-style) are produced on request. Every format is verified by an automated export bench before we quote it. We do not offer MCAP or WebDataset — MCAP is a robotics logging container rather than a training format, and labs shard WebDataset internally; we would rather tell you which is which.

Tier 1 · Session

What the clip is

Scene, task label, frame count, frame rate, usability verdict — and the provenance block: clip id, content hash, cohort, QA verdict, and which fields were derived rather than measured.

Tier 2 · Action

What happens, and when

Frame bounds per action segment, hand laterality (left / right / both), scene semantics, camera motion, a verb-noun action code and a dense natural-language description.

Tier 3 · Kinematics

Where the hands are

21 3D joint landmarks per hand, camera-relative wrist position, frame-to-frame velocity, and a per-hand detection confidence — so you can filter on quality rather than trust it.

Why egocentric

A head-mounted camera sees the task from where the hands are.

Teleoperation gives you a robot's own proprioception, and it costs one person driving one rig for every hour of demonstration it produces. Egocentric capture gives you human hands solving the task in the environment the task actually happens in, and it scales with people rather than with rigs. There is a real case against it — no joint torques, no action labels for free, a domain gap you have to close — and the published work arguing both sides is a category argument, not a FourthHuman result. We set it out on the method page, including the part that does not favour us.

Price

We publish what reaches the person holding the camera.

We pay Capture Partners ₹250–400 an hour, direct by UPI on QA approval, with no agent in the middle and no deduction. We publish that band on a public page rather than quoting it privately per engagement — so it is the one cost inside a FourthHuman quote you can check before you have spoken to us.

What moves the number · 1

The environment

An Indian household kitchen and an active distribution centre are not the same shoot. Access, scheduling and how many people have to consent all change with the room.

What moves the number · 2

The session

How long a session runs, how many of them, and how many Capture Partners the brief needs to reach the volume you are asking for.

What moves the number · 3

The annotation tier

Tier 1 is session metadata. Tier 3 is 21 3D joints per hand with per-hand detection confidence, and it costs more to produce and more to verify. You pick the tier in the brief.

A brief is scoped first and priced against that scope. There is no rate card yet, and we would rather say so than publish a figure the first real engagement would break. The structural reason a bespoke brief is affordable here at all is the India–US wage differential — which is also why we publish our input cost instead of hiding it inside a margin.

FAQ

Frequently asked questions

What kind of training data do you produce?
We build training data for Physical AI: egocentric video, 3D hand pose, and manipulation trajectories. Four stages take a brief to a dataset — Acquire (egocentric video capture via our Capture Partner network in India), Enrich (automated 3D annotation with 21-joint hand keypoints, temporal action segmentation), Prepare (quality scoring, multimodal alignment, DPDP-aligned anonymization), and Validate (benchmark curation, multi-format export). We deliver annotated datasets in LeRobot v2.1, HDF5 (EgoDex-style paired files) and the open FH-Ego v1 JSON schema, with Zarr v2 and TFRecord (RLDS-style) available on request. We do not offer MCAP or WebDataset — a logging container and an internal sharding layout, neither of which a training dataloader consumes.
How do you handle privacy and data protection?
Every clip passes automated face redaction server-side at ingestion, before any human review or delivery. The detector is YuNet, a DNN we chose because egocentric faces are off-axis, motion-blurred and partial — the angles a classical cascade misses. The stage fails closed: if the detector cannot load, the clip fails the stage rather than passing through, so there is no unredacted path to delivery. Licence-plate redaction is not yet implemented — we would rather name the gap than let you assume it is covered. Every capture session produces a digital consent receipt persisted with a SHA-256 audit signature, revocable by the Capture Partner under the data-deletion rights in India's DPDP Act 2023. To be precise about what else we do not yet have: we hold no third-party attestation — no SOC 2 report, no ISO 27001 certificate. Those require an external audit we have not commissioned, and we will link the report here when we do.
How much do you pay your Capture Partners?
We pay ₹250–400/hr ($3–4.80/hr) and we publish the range instead of quoting it privately per engagement. We have not measured what comparable capture work pays in India, so we do not claim a multiple against it — read the band and judge it against what you know. Pay is by UPI on QA approval, direct to the Capture Partner with no agent in the middle and no deduction. Those are the terms partners will be paid on, not a payout rail already running: no payment provider is integrated and no payout has ever settled. Ethical compensation is not charity: it buys retention, and retention is what makes capture quality consistent across a long collection run.
What annotation schema do you use?
We use the FourthHuman Ego Schema v1 — a 3-tier annotation schema: Tier 1 (Session) — task_id, scene, task_label, usability flag. Tier 2 (Action) — frame bounds, hand laterality (left/right/both), scene_semantic, camera_motion, verb-noun action code, dense NL description. Tier 3 (Kinematics) — 21 3D joint landmarks per hand, camera-relative Cartesian position vector P=[X,Y,Z], and frame-to-frame velocity, each hand carrying a detection confidence for the frame. All exported as structured JSON alongside video streams.
What data formats do you deliver datasets in?
We deliver LeRobot v2.1 (pi0/openpi and GR00T pipelines — our output loads with the real lerobot package), HDF5 (EgoDex-style paired .hdf5+.mp4 — ALOHA / RoboMimic-lineage labs), and the open FH-Ego v1 JSON schema. Zarr v2 (UMI / diffusion-policy layout) and TFRecord (RLDS-style step schema; not tfds.load-compatible) are produced on request. Every export is read back by an independent verifier before delivery. We do not offer MCAP or WebDataset — MCAP is a robotics logging container no training dataloader consumes, and WebDataset is an internal sharding optimization, not a distribution format.
What makes you different from a general annotation vendor?
We work only on Physical AI, not general-purpose annotation. We combine a Capture Partner network in India (ethical pay at ₹250–400/hr, a rate we publish rather than quote privately) with automated 3D annotation pipelines. The India–US wage differential is the structural cost advantage; we publish our own input rate rather than a comparison against a US teleoperation figure we have not independently verified. Unlike general-purpose BPO annotation vendors, we own the full stack: Capture Partner recruitment, capture hardware, QA/privacy pipeline, and AI annotation. Delivery is LeRobot v2.1, HDF5 (EgoDex-style) and the open FH-Ego v1 JSON schema, with Zarr v2 and TFRecord (RLDS-style) on request. We don't sell raw footage — we sell annotated, privacy-cleared, format-ready training data.
How does your pricing compare to US-based teleoperation data collection?
We publish our own input rate — ₹250–400/hr to the Capture Partner — instead of quoting a US teleoperation benchmark we have not independently verified. The India–US wage differential is the structural tailwind and it is real; what we won't do is put a number on someone else's cost base and present it as measured. Ask for a quote against your capture spec and compare it with what you pay today. We capture value by delivering annotated, format-ready data, not raw footage.

Get started

Tell us what you’re training.
We’ll scope the dataset.

Describe the policy and the embodiment you are training it for. We scope the capture specification, price it against your volume and format, and come back with what we can commit to — including the parts we cannot. One of the two founders reads every one of these. There is no queue, because there is nobody to queue behind.

What a brief costs is scoped before it is priced — how we price, and what moves the number.

All fields are required unless marked optional.

We reply from a named person, not a queue.

LeRobot v2.1, HDF5 and FH-Ego v1 JSON are delivered; Zarr and TFRecord are produced on request.

Volume, timeline, the constraint that matters most — whatever tells us whether we can help.

What happens when you submit

When you press Send this brief, your response opens as a pre-filled email addressed to [email protected] and can be sent directly to our team.

Your message is not submitted to or stored on our website. It goes directly to our team for personal review, with no automated acknowledgement or sales queue.

We aim to respond within 1–2 business days. If you have not heard from us by then, please feel free to follow up. We look forward to hearing from you.