LH2 AI Labs
Website:
lh2.fr
Company:
https://www.linkedin.com/company/lh2-ai-labs
Seniority: Mid-Senior level
Industries: Artificial Intelligence
Job details:
Job Title: Founding Engineer - Data Products
Location: On-site, Bengaluru
Employment Type: Full-Time
About LH2 AI Labs
Built by second-time founders who have built and sold companies before, LH2 AI Labs is building the post-training infrastructure for frontier AI models.
We bring private, high-quality institutional datasets and vetted domain experts into frontier AI pipelines across verticals such as coding, computer use, agentic workflows, medical, audio, and more.
For AI to keep progressing, it needs high-quality training data drawn from real production use cases. The public web has already been crawled and trained on—there is limited new signal left there. That is where we come in.
Our vision is to create a world where frontier models can access high-quality data on tap, the same way they access compute today.
About the role
We're building the platform that turns raw enterprise data — Slack, Jira, GitHub, email, docs, tickets — into structured, machine-trainable products for AI labs. You'll own the data foundation end to end: how we ingest from many enterprise sources, normalize the mess into one common shape, resolve identities across systems, and assemble the structured "decision" units we sell.
This is a hands-on founding role on a small team. You'll make the core architectural calls, set the engineering patterns the rest of the team builds on, and get a real product shipping inside customer cloud environments fast.
Responsibilities:
- Own the ingestion → normalization → entity-resolution → product-assembly pipeline architecture
- Build and operate production connectors across multiple SaaS/API sources (starting Slack, Jira, GitHub; expanding to Google Workspace, Teams, Notion, Confluence, CRM)
- Implement reliable incremental sync, pagination, retries, backfills, and checkpointing
- Design the medallion data model (raw → normalized entities → sellable products) and the transformation layers between them
- Build tiered entity resolution — authoritative and deterministic cross-system identity joins first, probabilistic matching later — with confidence and evidence retained
- Assemble structured, citation-backed "task/decision" units from resolved data, seeded from concrete outcomes (merged PRs, closed tickets)
- Handle schema evolution, deduplication, late/deleted data, and per-tenant workflow differences via config, not code forks
- Build pipelines that are idempotent, replayable, observable, and reproducible (versioned output manifests)
- Establish data-engineering standards and patterns for the team
- Work closely with the Privacy/PII founding engineer so the pipeline hands off cleanly at the trust boundary
Must-have:
- Graduate or postgraduate degree from a Tier 1 college.
- 6+ years of data/software engineering, including owning a production pipeline that ingested from multiple external APIs with incremental sync, retries, and schema drift (this is the non-negotiable – not "I've used a pipeline," but "I built and operated one")
- Strong Python and SQL
- Strong ETL/ELT experience across multiple SaaS platforms and APIs
- Deep understanding of OAuth, pagination, rate limits, incremental/delta sync, backfills, idempotency, and failure recovery
- Experience designing data models, normalization layers, and entity resolution across heterogeneous sources
- Experience with a modern warehouse/lakehouse (Snowflake, BigQuery, Databricks, Redshift, or equivalent) and orchestration (Airflow, Dagster, Prefect, or equivalent)
- Comfortable owning architectural decisions and setting patterns in a small founding team
- Experience operating data workloads in cloud environments, ideally customer VPC / private-cloud deployments
Nice-to-have:
- Connector platforms (Airbyte, Meltano, Fivetran) and building custom API connectors
- Entity-resolution tooling (Splink, Zingg) and record-linkage fundamentals
- Identity/directory/HR data sources as an ER seed
- Data lineage/catalog tooling (OpenLineage or similar)
- Handling large volumes of semi-structured/unstructured data
- Familiarity with AI training-data workflows (SFT, RLHF, evals) and what labs actually buy
Click on Apply to know more.