Website:
combinehealth.ai
Job details:
Senior AI/ML Engineer
CombineHealth | Bangalore (HSR Layout) | Hybrid, 4 days/week in-office
About CombineHealth
CombineHealth builds AI employees for healthcare revenue cycle management — agents that automate medical coding, billing, eligibility checks, denial management, and appeals for hospitals and physician groups across the US. Our platform runs in production at 97%+ coding accuracy across live customers, integrated with Epic, Cerner, athenahealth, and other major EHRs.
Before You Apply — Read This First
4 days/week in our HSR Layout, Bangalore office. Remote or fully-hybrid-light setups won't work for this team.
Regular overlap with US business hours. Our customers, most of our leadership, and key stakeholders are US-based. You'll need flexibility for calls and collaboration that land in US morning/afternoon (India evening/night) on a recurring basis, not occasionally.
If either of these is a dealbreaker, save yourself and us the time — this isn't the role.
What You'll Actually Do
This is an agent-engineering role. Our systems are LLM-based agents, and the hard part isn't training models — it's making autonomous, multi-step systems reliable and auditable on regulated, high-stakes data. You'll build the agents and own getting them into production reliably. Roughly:
- Build AI agents that read unstructured clinical and claims documentation and produce structured, auditable outputs — codes, claim decisions, denial root causes, appeal drafts
- Design the agent architecture: multi-step orchestration, tool/function calling, retrieval, and context engineering — plus the prompt and evaluation loop that makes it reliable, not just good in a demo
- Build eval harnesses and regression suites with hard accuracy bars. "The eval moved" is the unit of progress here — if you can't measure whether a change helped, it didn't happen
- Own reliability and observability in production: tracing, logging, latency, cost, failure handling — and yes, the pager when your agent breaks at 11pm US time
- Design confidence-threshold and human-in-the-loop review systems — this is regulated, high-stakes data; "mostly right" isn't good enough, and every decision needs a traceable rationale
- Work directly with product and RCM domain experts to translate messy real-world payer behavior into agent requirements
- Own your systems end-to-end: if it breaks in production, you're the one who gets paged, not someone else
What We Need
- 5+ years shipping production software, with real depth building LLM-based systems — agents, RAG, structured extraction. Production systems with real users and real data, not POCs or research. Tenure in classical ML or CV is not what we're screening for.
- Strong engineering fundamentals. You write production-grade Python, not notebooks. You design systems, handle failure paths, and care about latency and cost.
- Fluency in the LLM application stack: prompt and context engineering, tool/function calling, retrieval, and — most important — eval-driven development. You know how to measure whether a change actually improved the system.
- Comfortable owning your system in production end-to-end: deployment, observability, tracing, debugging live issues, cloud (AWS/GCP/Azure). We're not expecting a dedicated infra specialist, but you own your agent in production — you don't throw it over a wall.
- Experience with unstructured/messy real-world text where the output has to be accurate and explainable. Healthcare is a bonus, not a requirement — legal, financial services, insurance, or other high-stakes-accuracy domains count.
- Demonstrated ability to work independently and make judgment calls. We will not micromanage you, and we need people who don't need it.
- Comfortable in a fast-moving startup environment: ambiguity, shifting priorities, and building things that don't have a textbook answer yet.
Nice to Have (Not Required)
- Exposure to healthcare claims, medical coding (CPT/ICD-10/HCPCS), or RCM workflows
- Experience in a regulated data environment (HIPAA, SOC 2, or similar)
- Fine-tuning, distillation, or building smaller task-specific models where a general LLM is too slow or expensive
What Success Looks Like in 6 Months
- You own at least one production model/agent end-to-end — from data to deployment to monitoring
- You've meaningfully improved an accuracy, latency, or automation-rate metric on a live customer workflow
- You can independently debug a production issue in your system without pulling in a senior engineer to hold your hand
How to Apply
Please click on the link and fill in the details: https://forms.gle/6Dr8Y3Dp5uPrnn6V6
Click on Apply to know more.