About FourthHuman
The annotation engine works. The Capture Partner network is in pilot.
FourthHuman is a two-founder company in India that turns first-person video of real physical work into annotated training data for humanoid and VLA models. This page states what we have built, what we have not, and how the people who record for us are paid — in that order, because that is the order a buyer needs it in.
We compete on two things and it is worth naming them plainly. The first is provenance you can verify — every clip carries a consent receipt whose hash you can recompute yourself, rather than a badge attesting that someone reviewed a process. The second is bespoke capture at an Indian cost structure: you define the brief, review a sample, then scale, at a labour cost base a US or EU supplier cannot match. Data quality is not on that list. Every provider claims it, so it distinguishes nobody.
Standing
What we have proven, and what we have not
Pre-revenue. Zero customers. One validated end-to-end annotation run. Every row below resolves to a committed artifact or to a deliberate zero, and the zeros are published on purpose — they are the rows most vendors leave out.
| What | Standing | Where it comes from |
|---|---|---|
| End-to-end annotation runs | 1 · Measured | One committed clip, annotated all the way through: 1,843 frames, 16 action segments, 5 timeline tracks. It is on the reference implementation page and you can read the JSON. |
| 3D joint confidence on that run | 0.8712 mean · Derived | Arithmetic mean over 294 kinematic frames — an aggregate of the tracking model's own confidence output, which makes it derived, never measured. The floor is 0.508 and we publish the floor too, because worst case is what a buyer is actually assessing. |
| Delivery formats implemented | 5 · Measured | Detected in the services directory by the writer or library each format requires — not asserted. An export bench reads every one back before we quote it. |
| Backend stores still on placeholder data | 0 · None | Every store behind the engine is durable. This row exists because it was not always zero, and the count is re-derived rather than remembered. |
| Commissioning labs served | 0 · None | We have no customers. No company is named as one anywhere on this site, and none will be until one agrees to be. |
| Third-party attestations | 0 · None | No SOC 2. No ISO 27001. No audit of any kind. What we offer instead is a mechanism you can check per clip, not a badge attesting to a process. |
| Capture hours in hand | 0 · None | The Capture Partner network is in pilot. Catalogue entries state a target volume and an availability status; none of them states inventory. |
| QA acceptance rate | Not measured | Producing this number is the whole point of the phase we are in. Until a real cohort has been reviewed there is no rate, so we quote none. |
Method
What the engine actually does
Six stages between a brief and a delivered dataset. Nothing here is aspirational — this is the path the one validated run took, and the path a commissioned brief would take.
01 · Brief
A capture specification, agreed first
A brief specifies the task, the environment, the target volume and the delivery format. Nothing is recorded before it is agreed, which is why we sell bespoke capture rather than a shelf.
02 · Consent
A receipt before a frame
Consent is written per task, before capture, into an append-only hash-chained ledger. It can be revoked at any time, and revocation deletes — the mechanism is documented.
03 · Ingestion
Faces removed before any human looks
Automated face redaction runs server-side at the ingestion boundary, before annotation and before a reviewer opens the file. The detector is YuNet. The stage fails closed.
04 · Annotation
FH-Ego v1, three tiers, open
3D hand keypoints, temporal action segments and object interactions, emitted against our own schema. The schema is published, so you can read what a field means before you buy a file full of them.
05 · QA
Derived is labelled derived
Each clip carries a verdict, and anything a model produced is marked
derived in the file itself. Model output is never presented as
measurement — not to a buyer, not on this page.
06 · Delivery
Releases that recompute
A delivery is cut as an immutable, content-addressed release: the release id recomputes from the manifest, so you can check at any time whether every clip in it is still intact.
We deliver LeRobot v2.1, HDF5 (EgoDex-style paired files) and the open FH-Ego v1 JSON schema; Zarr v2 and TFRecord (RLDS-style) are produced on request. Every format is verified by an automated export bench before we quote it. We do not offer MCAP or WebDataset — MCAP is a robotics logging container rather than a training format, and labs shard WebDataset internally; we would rather tell you which is which.
Ethical sourcing
How Capture Partners are paid, and why the rate is published
We pay ₹250–400 per hour of accepted capture, by UPI, on QA approval — direct, with no agency in between. We publish the band instead of quoting it privately per engagement. Those are the terms partners will be paid on, not a rail already running: no payment provider is integrated and no payout has settled yet.
This is not charity. Paying well, and publishing what we pay, buys retention — and retention is what keeps capture consistent across a long collection run — a partner who stays learns the brief. It is also the reason we call them Capture Partners rather than collectors: the noun is meant to do argumentative work, not decorate.
Publishing the band costs us the ability to quietly pay less. That is the point. A rate you can read on a public page is a rate a partner can hold us to, and a buyer commissioning a brief can see exactly what share of their budget reaches the person holding the camera.
At the floor — ₹250 / hr
₹37,500
To the Capture Partner, by UPI, on QA approval.
At the ceiling — ₹400 / hr
₹60,000
The top of the band we publish.
Method — hours × the published ₹250–400/hr band. Arithmetic on a stated rate, not a payroll record: the network is in pilot.
Not a price. It is what the Capture Partner receives. What a brief costs a Commissioning lab is quoted against that brief, and we do not publish a rate card we have not settled.
Not a comparison. This page used to show an “industry average” wage and a multiple against it. Neither was sourced, so both are gone.
Not a payroll record. No capture has been paid at scale yet. The band is policy; the total is arithmetic.
The company
Two founders, and no one else yet
There is no sales team, so the person who answers your first email is one of the two people below. We think that is worth saying plainly rather than implying a bench we do not have.
Jayesh
Co-founder · engine and infrastructure
Jayesh owns everything between a recorded clip and a delivered dataset: the ingestion boundary and its face-redaction stage, the 3D annotation pipeline (MediaPipe), the FH-Ego v1 schema, the export bench that verifies a format before we quote it, and the hash-chained consent ledger. When this page says the engine works, the run it works on is his.
Pipelines · Schema · Consent ledger
Krishna
Co-founder · market and method
Krishna owns what we sell and how we prove it: the capture catalogue and how a brief is scoped, competitor and pricing research, Commissioning-lab relations, and Capture Partner recruitment. He is also the reason this page says “zero customers” rather than something more flattering — the content-integrity rules this site is built under came out of an audit he ran on our own copy.
Catalogue · Research · Partner relations
Redaction before review
Face redaction that fails closed
Sourcing egocentric video inside homes, clinics and factories means recording people who never agreed to be in a training set. Under India's Digital Personal Data Protection (DPDP) Act 2023, data principals hold consent and deletion rights, and the Act sets penalties of up to ₹250 crore for a breach of the security-safeguards obligation.
That is the exposure we engineer against. Every clip passes automated face redaction at the ingestion boundary — before annotation, before any reviewer opens the file. The detector is YuNet, a DNN we chose because egocentric faces are off-axis, motion-blurred and partial, which is exactly where a classical cascade fails. Every session is bound to a consent receipt carrying a SHA-256 audit signature that the Capture Partner can revoke.
Covered: human faces, by YuNet, server-side at ingestion, before any human review.
Fails closed: if the detector cannot load, the clip fails the stage rather than passing through. There is no unredacted path to delivery.
Not covered: licence plates. Not implemented. We name the gap rather than let it be assumed away.
Not EgoBlur: this is our own stage, not Meta's EgoBlur, and we do not restate EgoBlur's accuracy as ours.
Next
Bring us a brief
The useful first conversation is a capture specification: the task, the environment, the volume and the format you train against. We will tell you what we can scope, and what we cannot.