Website:
Job details:
About the role
Apeiro Digital builds sovereign health data infrastructure for enterprise and governments and
ministries of health across the world. Our stack spans a FHIR-native EMR, a national logistics and
supply chain platform, pharma track and trace, and a Health Information Exchange. The data team
underpins all of these — turning transactional health data into real-time analytics, ML-assisted
forecasting, and regulatory intelligence.
This role exists at the intersection of data engineering and applied data science. You will own both
production pipelines and the models that run on top of them.
What You Will Do
– Build and maintain production services and APIs that serve analytics and ML model outputs
at scale across multicountry GCC deployments.
– Design and implement ETL/ELT pipelines feeding analytical and ML workloads — including
supply chain demand forecasting, claims adjudication signals, and population health
indicators.
– Develop predictive and statistical models for health supply chain commodity flows, eClaims
payer analytics, and population health use cases.
– Collaborate with product and domain teams to translate clinical and operational questions
into well-scoped data problems with measurable success criteria.
– Own model monitoring, retraining pipelines, and data quality frameworks — not just initial
deployment.
– Contribute to and extend the data lakehouse architecture built on Apache Iceberg,
ClickHouse, and Parquet — with Dagster orchestrating the pipeline.
– Participate in architecture decisions around the analytics data stack, including ClickHouse
DR, lakeFS as Iceberg REST catalog, and dbt-clickhouse transformations.
– Write clean, testable, production-grade code — not research notebooks passed to an
engineering team.
What We Are Looking For
– 6–9 years of total experience with a demonstrable split between data science and backend
or data engineering — not purely one or the other.
– Strong Python (non-negotiable): Pandas, NumPy, scikit-learn, SQLAlchemy, and FastAPI or
equivalent async frameworks.
– Solid SQL and hands-on experience with columnar or analytical databases — ClickHouse
strongly preferred; BigQuery, Snowflake, or DuckDB accepted.– Production model deployment experience — you have shipped models to live environments
and maintained them, not just trained and handed off.
– Familiarity with Kafka or event-driven architectures is a meaningful plus.
– Experience with dbt or equivalent transformation tooling and an understanding of medallion
or lakehouse architecture patterns.
– Experience in health data - FHIR, claims, supply chain, or logistics — is a strong
differentiator but not a hard requirement.
– Comfort working in a B2G context where data sovereignty, auditability, and schema
compliance are first-class concerns.
Nice to Have
– Exposure to GS1 or EPCIS data formats for pharmaceutical or medical device traceability.
– Experience with Dagster, Prefect, or Airflow for pipeline orchestration.
– Familiarity with lakeFS, Apache Iceberg, or open table format ecosystems.
Why This Role
– You will be a technical authority, not a ticket-taker. Architecture input is expected, not
tolerated.
– Apeiro is part of IHC Group, giving you the security of a large holding company with the
pace of a product startup.
Click on Apply to know more.