About the role
Founding Engineer — Data Engineer
Full-time · Founding team · ~4-6 years of experience
Compensation. - 30 LPA with 2% equity.
About the Role
[company intro]
You own the data substrate and batch runtime of a new product we're building from scratch. This is a data-specialist seat, not a general backend seat: routine app development gets guided in weekly reviews; what can't be substituted is your depth in data modeling, set-based computation and pipeline correctness.
Stack: Python, PostgreSQL, Django, server-rendered frontend (htmx).
What You'll Do
Third-party data ingestion done deeply — webhooks drop silently, so reconciliation against source-of-truth totals is not optional
Batch computation over large datasets as set-based SQL, inside fixed nightly windows
Multi-tenant isolation as a reliability property — one tenant's bad data must never degrade another's run
Backfill and replay as designed capabilities, not emergency scripts
Append-only, auditable records — every automated decision must be reconstructable
Vendor API integrations built for failure: circuit breakers, retries, reconciliation
The ops substrate: alerting, partitioning, idempotent jobs
Qualifications
Mandatory
Around 5–8 years building and operating production systems — years are a proxy, the bullets below are the bar — with deep Postgres chops: schema design, query plans, batch performance at real scale
You've owned data pipelines in production — ingestion, transformation, reconciliation, serving (not just app-layer CRUD against an ORM); moving and reshaping data is the job you've actually done, whatever your resume calls it
You think in sets, not loops — multi-step SQL transformations, window functions, statistical aggregates (percentiles, distributions) and incremental computation are your native mode, and you've turned raw event or order data into cohort, retention or funnel metrics
Production Python backend experience irrespective of the framework — you keep business logic cleanly separated from framework details, which is exactly why the framework doesn't matter
Tests are how you build, not what you add later — batch jobs and reconciliations especially, where a bug means silently wrong data for weeks
War stories involving queues backing up, webhooks silently dropping, or reconciliations catching data loss no one noticed
Strong written communication — this is a documentation-driven team
Comfortable owning outcomes between weekly reviews — your tests and written decisions carry the quality bar in between, not supervision
Preferred
Production Django, e-commerce platform experience, streaming/event-architecture experience (Kafka-class systems), familiarity with probabilistic/approximate data structures (HyperLogLog, Bloom filters, t-digest), comfort building the occasional server-rendered screen.
Note
The first iterations run on plain old Postgres, architected excellently — a deliberate bet, not a constraint. The heavier machinery (Spark-class compute, warehouses, distributed runtimes, etc.) isn't ruled out, just deferred until scale earns it; the roadmap already holds real algorithmic depth: incremental/delta computation, mergeable sketches, analytical query engines.
This page is fully interactive when JavaScript is enabled. Please enable JavaScript to apply or browse related roles.