Website:
lh2.ai
Job details:
Job Title: Founding Engineer - Data Products
Location: On-site, Bengaluru
Employment Type: Full-Time
About the role
We're building the platform that turns raw enterprise data — Slack, Jira, GitHub, email, docs, tickets — into structured, machine-trainable products for AI labs. You'll own the data foundation end to end: how we ingest from many enterprise sources, normalize the mess into one common shape, resolve identities across systems, and assemble the structured "decision" units we sell.
This is a hands-on founding role on a small team. You'll make the core architectural calls, set the engineering patterns the rest of the team builds on, and get a real product shipping inside customer cloud environments fast.
Responsibilities
- Own the ingestion → normalization → entity-resolution → product-assembly pipeline architecture
- Build and operate production connectors across multiple SaaS/API sources (starting Slack, Jira, GitHub; expanding to Google Workspace, Teams, Notion, Confluence, CRM)
- Implement reliable incremental sync, pagination, retries, backfills, and checkpointing
- Design the medallion data model (raw → normalized entities → sellable products) and the transformation layers between them
- Build tiered entity resolution — authoritative and deterministic cross-system identity joins first, probabilistic matching later — with confidence and evidence retained
- Assemble structured, citation-backed "task/decision" units from resolved data, seeded from concrete outcomes (merged PRs, closed tickets)
- Handle schema evolution, deduplication, late/deleted data, and per-tenant workflow differences via config, not code forks
- Build pipelines that are idempotent, replayable, observable, and reproducible (versioned output manifests)
- Establish data-engineering standards and patterns for the team
- Work closely with the Privacy/PII founding engineer so the pipeline hands off cleanly at the trust boundary
Must-have skills
- 6+ years of data/software engineering, including owning a production pipeline that ingested from multiple external APIs with incremental sync, retries, and schema drift (this is the non-negotiable – not "I've used a pipeline," but "I built and operated one")
- Strong Python and SQL
- Strong ETL/ELT experience across multiple SaaS platforms and APIs
- Deep understanding of OAuth, pagination, rate limits, incremental/delta sync, backfills, idempotency, and failure recovery
- Experience designing data models, normalization layers, and entity resolution across heterogeneous sources
- Experience with a modern warehouse/lakehouse (Snowflake, BigQuery, Databricks, Redshift, or equivalent) and orchestration (Airflow, Dagster, Prefect, or equivalent)
- Comfortable owning architectural decisions and setting patterns in a small founding team
- Experience operating data workloads in cloud environments, ideally customer VPC / private-cloud deployments
Nice-to-have skills
- Connector platforms (Airbyte, Meltano, Fivetran) and building custom API connectors
- Entity-resolution tooling (Splink, Zingg) and record-linkage fundamentals
- Identity/directory/HR data sources as an ER seed
- Data lineage/catalog tooling (OpenLineage or similar)
- Handling large volumes of semi-structured/unstructured data
- Familiarity with AI training-data workflows (SFT, RLHF, evals) and what labs actually buy
Click on Apply to know more.