Humyn Labs
Website:
humynglobal.com
Company:
https://www.linkedin.com/company/humyn-labs
Industries: Software Development and Artificial Intelligence
Job details:
About Humyn Labs
Humyn Labs converts real human action - sound, sight, motion, and touch - into training signal for physical AI. We run verified field capture across 20+ countries in India, Southeast Asia, Latin America, and the Middle East: the real-world environments where physical AI deploys, not the labs where it is built.
Where each modality stands. Sound is proven: 100,000 hours across 33 languages, revenue-generating today, with BRIDGE - our published benchmark evaluating 15 ASR models across 22 languages and 7 voice parameters. Sight is the current bet: 500,000 hours across 10+ countries, residential and commercial, with a live labeling stack (6DoF pose, hand-skeletal tracking, IMU/SLAM) and a discard rate below 15%, down from 40–60%. Motion is integrating: joint physics and navigation, captured through the sight pipeline. Touch is the frontier: teleoperated grippers with force and tactile sensors, built as a byproduct of motion collection - not a separate build.
The physical AI data problem has four stages of increasing depth: foundation-building (scale, format, diversity), task generalisation (domain-aware labeling), deployment readiness (environment-specific evaluation), and continuous refinement (correction-episode capture from deployed robots). The technology leader we hire will build a platform that serves all four - simultaneously, at scale.
The architecture you will build
This is not a workflow tool. It is not a B2C application. It is not a systems integration project. It is an open data platform with two sides that must compound each other - and a live delivery engine that pays for it today.
The arc: Phase 1 (now) - Humyn is the first supplier, proving the pipelines directly: sound done, sight in production, motion integrating. Phase 2 (next) - Humyn publishes the fusion spec (timestamps, hardware, SOPs) and certified external networks join the supply side. Phase 3 (destination) - the full two-sided platform: any certified network's pipelines fuse into one signal; any lab buys it. You are hired in Phase 1 to build the architecture Phases 2 and 3 run on.
Supply side - collection infrastructure that scales non-linearly
An open architecture that allows multiple verified collectors, domain experts, and operational partners across 20+ countries to contribute multi-modal sensory data through Humyn's verification and labeling pipeline. Every modality runs the same four stages - collection, validation, processing, labelling - with human-in-the-loop QC across all of it. The platform must scale without Humyn owning every collection operation: contributors bring data; Humyn's pipeline processes, verifies, labels, and routes it. In Phase 2, the fusion spec you harden becomes the certification standard every external network must meet to plug in.
Demand side - open evaluation platform for research labs
An open interface that allows all four buyer cohorts - foundation model builders, humanoid builders, world model companies, voice AI platforms - to benchmark their models against Humyn's held-out evaluation sets: to identify which tasks fail, in which environments, under which conditions. The gap between what a model can do and what the evaluation reveals it cannot becomes a structured data brief that Humyn fulfils. BRIDGE for sound is the working proof of this model. The physical AI equivalent - deployment-readiness benchmarking for robots - is what you will build.
The capture stack - where the signal is won or lost
Quality is decided at capture time, not in post-processing. You will own the hardware and software discipline that field operations execute across 20+ countries: sensor selection and evaluation (stereo depth, global-shutter RGB, IMU specifications), hardware time-sync and calibration workflows, on-device validation, and the per-unit economics of running capture fleets at scale. You will make the build-vs-buy calls on rigs and drive the hardware roadmap as specs from frontier labs tighten - including the touch path: teleoperated grippers with force/tactile sensing, designed as an extension of motion collection.
This is not future platform work. Live contracts with frontier labs carry hard hardware specifications and delivery deadlines today. You own throughput, spec compliance, and discard rate on those deliveries from day one.
The processing and fusion stack
The pipeline that converts raw multi-modal capture into embodiment-aligned training signal: 6DoF head pose, 21-keypoint hand-skeletal tracking, IMU/SLAM spatial reconstruction, dense natural-language action labeling, verifiable provenance records, and format-agnostic delivery in MCAP, RLDS, and LeRobot v3. Fusion is the core conversion: many synchronized streams become one multi-modal signal per domain - the same spec across every modality, geography, and environment. This stack is what separates a raw-footage commodity supplier from a verified-signal platform priced an order of magnitude higher. You will own it, extend it across modalities, and make it the processing standard researchers depend on.
Real-time field QA infrastructure
Feedback that arrives in post-processing is useless - the session is over. You will maintain and extend the real-time quality validation layer that operates at collection time across 20+ countries: frame-drift detection, signal-to-noise monitoring, domain compliance checking, and instant feedback to field contributors in their own language. This is how Humyn got from 40–60% discard rates to below 15%. You will drive it further.
Who we are looking for
Multi-modal data platform builder - not an application developer
- You have built production systems that process, label, and deliver data at scale to ML researchers or model training pipelines - not consumer-facing products, not internal workflow tools, not system integrations.
- You understand what it means to produce a training signal, not just store data - and why the difference matters.
Platform architect - not a pipeline builder
- You have worked closely with ML researchers or foundation model teams as an infrastructure partner, not a service provider.
- You understand what researchers need to evaluate model efficacy, why they distrust benchmarks they did not help design, and how to build evaluation infrastructure they will cite rather than ignore.
- You think in two-sided platforms, not linear pipelines: more supply enriches the evaluation standard; better evaluation generates better data briefs; better data briefs attract more supply.
- You have built open APIs or developer-facing infrastructure that scaled without proportional headcount growth.
Capture-literate - not post-processing only
- You have real fluency with the physical capture stack: camera intrinsics and extrinsics, stereo depth, IMU fusion, hardware time-sync, SLAM.
- You have debugged a calibration or sync problem in the field - not just read about one.
- You can hold your own in a hardware specification negotiation with a frontier-lab research team.
Multi-modal depth - not single-modality expertise
- You think natively across sound, sight, and motion, and you understand why synchronised multi-modal data from the same session is worth more than the sum of its modalities.
- You are not a computer vision engineer who tolerates audio, or an ASR specialist who considers video an edge case.
- You have worked with researchers consuming training data and understand what they actually need: format, provenance, evaluation metadata - not just pixels or waveforms.
Research infrastructure mindset - not enterprise software
- Your mental model for the demand-side platform is closer to Hugging Face or OpenReview than to Salesforce or ServiceNow.
- You want to build the infrastructure researchers cite, not the software enterprises procure.
Experience signals
- Built ML data pipelines that delivered training data to model research teams - at scale, in production, with quality guarantees.
- Hands-on with multi-sensor capture systems: sensor selection, calibration, hardware synchronisation, field deployment.
- Worked with multi-modal data across at least two of: video/egocentric, audio/speech, IMU/motion, tactile.
- Designed or contributed to open platform or API architectures that enabled non-linear growth on the supply or demand side.
- Built evaluation or benchmarking infrastructure used by external researchers or model teams.
- Worked directly with ML researchers as an infrastructure partner - you understand what researchers actually ask for, how they consume training data, and what makes a benchmark useful versus ignored.
- Understanding of data provenance, quality frameworks, and verifiable data lineage.
- Experience scaling distributed collection operations: contributor networks, field QA, real-time validation.
- Physical AI, robotics, or embodied AI data experience is a strong signal - but not required if the capture and platform depth is there.
What success looks like in 18 months
- Live frontier-lab contracts delivered on spec and on time, with discard rate driven below today's sub-15% benchmark as volume scales.
- The fusion spec is published - timestamps, hardware, SOPs - and at least one certified external network is contributing verified multi-modal data through Humyn's pipeline without Humyn managing the collection operation.
- The demand-side evaluation platform has at least three foundation model labs using it to benchmark physical AI deployment readiness.
- The processing stack extends to motion data with the same quality guarantees as sight, and the touch capture path - teleoperated grippers with force/tactile sensing - is designed and in early collection.
- Humyn is cited by at least one foundation model lab as their evaluation-infrastructure standard for physical AI deployment readiness.
Click on Apply to know more.