Website:
getbiostack.com
Job details:
Location: India
Experience: 3+ years
Employment Type: Full-time
Hiring Priority: Urgent
About BioStack
BioStack is building the data and environment layer for medical AI. We work with large-scale healthcare datasets across clinical records, imaging, pathology, labs, genomics, and longitudinal patient data.
We are urgently hiring a Data Engineer who can take ownership of large datasets from ingestion through final delivery.
What You’ll Do
- Ingest and process large datasets from hospitals, clinics, labs, research organizations, and data partners.
- Build pipelines for structured and unstructured healthcare data.
- Clean, normalize, transform, merge, and validate large datasets.
- Work with datasets containing millions of records and files.
- Build ETL/ELT pipelines that are fast, reliable, and reproducible.
- Automate data-quality checks for missing, duplicate, inconsistent, or corrupted records.
- Design schemas and standardized formats for heterogeneous datasets.
- Build workflows for securely transferring and processing large datasets.
- Work with ML, clinical, and operations teams to turn raw data into AI-ready datasets.
- Maintain documentation around schemas, transformations, lineage, and dataset versions.
- Troubleshoot broken pipelines and unexpected data issues quickly.
- Take end-to-end ownership of data delivery.
Requirements
- 3+ years of professional experience in data engineering, backend engineering, infrastructure, or a related role.
- Strong Python skills.
- Strong SQL skills.
- Experience building production data pipelines.
- Experience processing large datasets using tools such as Spark, Ray, DuckDB, Polars, Pandas, or similar technologies.
- Experience with AWS, GCP, or Azure.
- Experience with object storage such as S3.
- Strong understanding of ETL/ELT, schemas, transformations, and data validation.
- Ability to work with messy, incomplete, and inconsistent real-world data.
- Strong debugging and problem-solving skills.
- Ability to work independently and move quickly.
Strongly Preferred
- Experience with PII/PHI detection, redaction, or de-identification.
- Experience working with healthcare or sensitive datasets.
- Familiarity with HIPAA or healthcare data privacy requirements.
- Experience building automated redaction or anonymization pipelines.
- Experience with EHR, clinical, imaging, or biomedical datasets.
Nice to Have
- Familiarity with FHIR, HL7, DICOM, ICD, CPT, LOINC, or SNOMED.
- Experience processing medical imaging datasets.
- Experience with clinical text or NLP pipelines.
- Experience preparing datasets for ML or LLM training.
- Experience working at an early-stage startup.
Who Will Do Well Here
You should be comfortable receiving a raw dataset with millions of files, inconsistent formatting, missing fields, duplicate records, and sensitive information—and figuring out how to make it usable.
We value people who are highly hands-on, move fast, and take ownership without waiting for perfectly defined requirements.
- This is an urgent requirement. Candidates who can join immediately or within a short notice period will be prioritized.
Click on Apply to know more.