Physical AI training data

Egocentric capture, specified to a brief, for models that have to reason about contact, force and spatial layout rather than read about them.

What FourthHuman does

What a brief captures

A capture specification pairs first-person video with the telemetry a policy needs — IMU streams and a 3D camera trajectory — and every clip is annotated against the open FH-Ego v1 schema. Nothing here is stock: the brief defines the task, the environment and the rig before anything is recorded.

How provenance works

Every clip carries a SHA-256 consent receipt in a hash-chained, DPDP-aligned ledger, so a buyer re-verifies one clip instead of trusting a badge. Automated face redaction (YuNet) runs server-side at ingestion, before any human review, and fails closed — it refuses to emit a clip rather than emit one un-redacted. Licence plates are not covered. We hold no third-party attestation — no SOC 2, no ISO 27001.

What a brief costs to scope

Scoping is a conversation, not a price list: environment, session length, sensor rig and annotation tier each move the number. Capture Partners are paid transparently at a published band — ₹250–400/hr of accepted capture, by UPI on QA approval, no agent in the middle and no deduction. Those are the terms partners will be paid on, not a payout rail already running: the Capture Partner network is in pilot build-out and no payment provider is integrated.

Signals a physical-AI brief specifies

Physical AI needs cause, effect, force and spatial layout — material compliance, weight distribution, tool mechanics. None of it survives a text prompt or a web scrape. Below are the streams a capture specification names, and the layers the annotation pipeline derives from them. Model-generated output is derived, never measured, and is labelled that way wherever it appears.

Streams and derived layers — specified per brief
SignalWhat it is
Egocentric video60fps high-resolution first-person capture, on the rig the brief specifies
IMU telemetryHigh-frequency accelerometer and gyroscope streams, time-aligned to the video
Camera trajectory3D camera pose over time (SLAM), so motion is metric rather than relative
Hand pose21-joint bimanual 3D hand skeletons (MediaPipe) Derived
Wrist trajectoryCamera-relative wrist vectors, executable by an actuator Derived
Action boundariesTemporal verb-noun segments marking where one action ends and the next begins Derived

Commission a brief

Tell us the task, the environment and the policy you are training, and we scope a capture brief against it. Stated plainly: we are pre-revenue with no customers, one validated end-to-end annotation run, and a Capture Partner network in pilot build-out. The catalogue lists capture specifications you can commission, with target volumes and their real availability status.

World models need contact and layout, not just pixels. Tell us which physical property your model is failing on and we will scope capture against that, rather than against a volume target.