LogixHealth
Website:
logixhealth.com
Job details:
Job Title: Data Engineer / Senior Data Engineer
Location: Bangalore / Coimbatore
Experience: 5+ years
Job Type: (Hybrid, Fulltime)
Immediate joiners or notice period of less than 10 days are needed
About LogixHealth
At LogixHealth, we're transforming Revenue Cycle Management (RCM) through technology, data engineering, automation, and AI. We're building solutions to simplify complex healthcare operations, improve accuracy, and accelerate the journey from clinical documentation to payment.
We're bringing together modern data platforms, intelligent automation, and AI-driven capabilities to reshape how healthcare organizations manage medical coding, claims processing, and billing.
We're looking for a passionate Data Engineer who wants to solve meaningful engineering challenges, build systems at scale, and help shape the future of autonomous healthcare operations.
The Role
As a Data Engineer, you'll be part of a globally distributed engineering team building the data foundation for next-generation healthcare technology. You'll design and develop scalable data pipelines, distributed processing systems, and governed lakehouse architectures using Databricks, Apache Spark, Python, SQL, and Azure.
Your work will directly support analytics, AI applications, autonomous medical coding, intelligent billing, and large-scale operational automation. This isn't just about moving data. It's about engineering the intelligence and infrastructure that will power the next generation of healthcare operations.
What You'll Build
- The data foundation for autonomous medical coding
Build scalable data pipelines that bring together clinical records, medical documentation, coding information, and historical claims. Enable AI systems to interpret clinical context, recommend codes, and support increasingly autonomous coding workflows.
- Intelligent and autonomous billing
Engineer the data infrastructure behind intelligent claim generation, validation, submission, and exception handling. Help build systems that reduce manual intervention, improve accuracy, and accelerate the revenue cycle.
- A modern, self-service healthcare data platform
Design a governed data ecosystem that connects healthcare clients, operational systems, and analytics workloads. Make trusted data accessible to engineers, analysts, AI applications, and business teams.
- AI-ready data and agentic workflows
Prepare high-quality, well-structured data for LLMs, AI agents, automation engines, and decision-support applications. Help establish the data pipelines, feedback loops, and monitoring needed to make intelligent automation dependable.
Key Responsibilities:
- Architect and engineer: Design, develop, and operate scalable data pipelines and distributed processing systems using Databricks, Apache Spark, Python, SQL, and Delta Lake.
- Build the lakehouse: Develop Bronze, Silver, and Gold data layers, reusable ingestion frameworks, incremental processing, and CDC pipelines.
- Optimize at scale: Engineer high-performance Spark workloads through partitioning, caching, Adaptive Query Execution, join optimization, and skew handling.
- Enable AI and automation: Deliver trusted, timely data for autonomous coding, billing, claims processing, analytics, and AI-driven decision-making.
- Establish data governance: Implement Unity Catalog, access controls, lineage, data quality checks, and healthcare data protection practices.
- Build production-grade engineering: Establish CI/CD, automated testing, infrastructure as code, observability, and operational reliability.
- Collaborate across teams: Work with data engineers, AI/ML engineers, architects, product teams, and business stakeholders to translate complex requirements into robust solutions.
- Continuously improve: Identify opportunities to simplify architecture, eliminate duplication, reduce processing costs, and increase automation.
- Own production outcomes: Troubleshoot complex data issues, improve pipeline reliability, and ensure solutions meet performance, scalability, and availability requirements.
Qualifications:
To perform this job successfully, an individual must be able to perform each duty satisfactorily.
The requirements listed below are representative of the knowledge, skills, and/or ability required. Reasonable accommodation may be made to enable individuals with disabilities perform the duties.
Required Qualifications
Education and Experience
- Bachelor's degree or higher in Computer Science, Engineering, or a related discipline, or equivalent practical experience.
- 5+ years of hands-on experience building production-grade data engineering solutions.
- Strong experience with Apache Spark and Databricks in cloud environments.
- Experience delivering and supporting large-scale data processing systems, ideally handling terabytes of data.
Technical Expertise
Apache Spark and PySpark
- Expert-level PySpark or Scala, DataFrames, Spark SQL, and distributed processing.
- Strong understanding of Spark execution, performance tuning, debugging, and optimization.
- Experience with partitioning, caching, Adaptive Query Execution, joins, and data skew handling.
- Experience with large-scale batch processing and Structured Streaming.
Databricks and Delta Lake
- Strong experience with Databricks notebooks, Jobs, Workflows, and production deployments.
- Expertise in Delta Lake, including ACID transactions, schema evolution, optimization, and incremental processing.
- Experience with Delta Live Tables (DLT) and pipeline design.
- Knowledge of Unity Catalog, data governance, access control, and lineage.
- Experience designing Medallion architectures (Bronze, Silver, Gold).
Programming and Data Engineering
- Strong Python and SQL skills.
- Experience designing reusable, modular, maintainable, and testable data pipelines.
- Strong understanding of ETL/ELT patterns, CDC, data quality, and data modeling.
- Experience building reliable ingestion frameworks and handling schema changes.
Cloud and DevOps
- Experience with cloud data platforms, preferably Microsoft Azure.
- Working knowledge of Azure Blob Storage / ADLS and cloud-native data engineering.
- Experience with Git, CI/CD, automated testing, and production deployment practices.
- Familiarity with monitoring, logging, and troubleshooting distributed systems.
Good to Have
- Experience with Azure Data Factory, Event Hubs, or other data integration technologies.
- Experience with Airflow or external workflow orchestration platforms.
- Familiarity with SQL Server, PostgreSQL, MySQL, or NoSQL databases.
- Exposure to AI/ML pipelines, LLM applications, RAG, agentic AI, or AI-assisted automation.
- Understanding of healthcare data, medical claims, revenue cycle operations, HIPAA, or other regulated data environments.
- Experience with Terraform or other infrastructure-as-code tools.
- Experience building data products and self-service analytics platforms.
Why Join Us?
- Build technology that matters. Your work will contribute to modernizing healthcare revenue operations and reducing the complexity of administrative workflows.
- Work on the next generation of RCM. Help transform a traditionally manual, complex industry through intelligent automation, autonomous coding, and AI-powered billing.
- Build for AI, not just analytics. Work on the underlying infrastructure that makes autonomous coding and billing systems possible.
- Solve meaningful engineering challenges. Work with complex healthcare data, distributed systems, high-volume pipelines, and demanding reliability requirements.
- Own real production systems. Contribute to architecture, implementation, optimization, deployment, and operational excellence.
- Grow with a modern technology stack. Build deep expertise in Databricks, Spark, Azure, data governance, and AI-enabled engineering.
- Collaborate globally. Work with cross-functional teams across engineering, AI, product, and healthcare operations.
The Engineer We're Looking For
We're looking for someone who enjoys solving hard engineering problems, takes ownership of production systems, and thinks beyond individual pipelines.
You should be comfortable challenging inefficient processes, designing for scale, writing clean and reliable code, and collaborating with teams to deliver measurable outcomes.
You don't need to be an AI researcher. But you should be excited about how data engineering, automation, and AI can change the way entire industries operate.
If you want to build the data infrastructure that enables healthcare operations to move from manual processing toward intelligent, increasingly autonomous workflows, we'd love to meet you.
Click on Apply to know more.