Data Unveil
Website:
dataunveil.com
Job details:
Position Title: Data Scientist — Patient Outcomes & Next Best Action
Experience: 3+ Years
Location: Hyderabad, Telangana
Hire Type: Full Time, On-site
Start Date: Immediate
About the role
We build the analytics and intelligence layer that our life sciences clients and their field teams use to act on patient data — hub enrollments, benefits and prior authorization activity, dispense history, and field engagement, brought together into one view of the patient journey.
We're hiring a Data Scientist to own the predictive models behind that layer. The core question is not "how did the program perform last month?" — it's "what should happen next for this patient, and who should do it?" You'll build the risk, propensity, and recommendation models that tell a Field Reimbursement Manager which cases are about to stall, which patients are drifting toward discontinuation, and which intervention is most likely to work.
This is not a dashboard-building role and it is not a data engineering role. Pipelines and data aggregation sit upstream.
What you'll do
- Model the patient journey end to end — enrollment through benefits verification, prior authorization, first fill, and ongoing therapy — identifying the stages, delays, and drop-off points that predict downstream outcomes.
- Build and own production models for adherence and discontinuation risk, therapy-switch and lapse propensity, time-to-first-fill, and prediction of which access cases will stall in copay or prior authorization.
- Develop Next Best Action recommendations for field reimbursement and field teams: which patients or cases to prioritize, which intervention to take, and when — with the reasoning attached, not just a score.
- Extend the same approach to HCP-level targeting, segmentation, and channel propensity — which provider to engage, on what theme, through which channel.
- Design the explainability surface: surface top contributing factors for every score in language a field user can act on, and make sure the explanation actually reflects the model.
- Build the measurement side honestly — define what a "successful" recommendation is, instrument the feedback loop from field outcomes back into model quality, and evaluate lift against realistic baselines rather than against doing nothing.
- Own the full model lifecycle: feature definition and lineage, training and versioning, deployment, drift and performance monitoring, and scheduled retraining.
- Work directly with product and client-facing teams to turn ambiguous asks into scoped analytical problems with defensible answers.
Technical skills
Statistical foundations
- Survival & time-to-event concepts: Understanding Kaplan-Meier curves and basic hazard models to properly handle right-censoring and patient drop-off timelines.
- Sampling & stratification: Group-based splitting and temporal stratification to prevent data leakage across time periods, patient clusters, or healthcare practices during model validation.
- Observational data & bias: Practical awareness of confounding, selection bias, and immortal time bias when modeling non-randomized patient journey data.
- Bayesian intuition: Understanding prior probabilities and updating estimates to handle small-sample or rare-event scenarios.
- Core inference & baseline testing: Solid foundation in hypothesis testing, confidence intervals, and proper handling of missing-not-at-random data to validate model lift against realistic baselines.
Machine learning
- Supervised learning at production quality: Gradient boosting (XGBoost, LightGBM, CatBoost), regularized regression, random forests; disciplined hyperparameter tuning, custom loss function optimization, cross-validation schemes that respect temporal ordering, and calibration of predicted probabilities (Platt scaling, isotonic regression).
- Class imbalance: Resampling, class weighting, threshold optimization, and choosing evaluation metrics (PR-AUC, lift at k) appropriate to rare-event prediction.
- Explainability & field trust (Crucial): Hands-on experience with SHAP, permutation importance, and partial dependence to surface transparent, actionable drivers behind every score so field users trust the outputs.
- Feature engineering on longitudinal data: Window aggregations, gap and recency features, trajectory shape features, point-in-time correctness, and rigorous avoidance of target leakage from post-outcome fields.
- Recommendation & ranking: Learning-to-rank, contextual bandits, or heuristic-based action selection under field-capacity limits.
- Sequence & unsupervised modeling: Clustering (k-means, HDBSCAN) for patient/HCP segmentation, and using sequence features or embeddings where they earn their complexity over tabular data.
ML engineering & tooling
- Python at production standard — pandas, NumPy, scikit-learn, statsmodels, and survival libraries (e.g., lifelines); clean, tested, reviewable code rather than notebook sprawl.
- SQL fluency against large relational datasets — window functions, complex joins, CTEs, query performance awareness; comfort building point-in-time correct training sets from transactional history.
- Cloud ML deployment (AWS preferred) — containerized model serving, batch and near-real-time scoring, model registry and versioning, orchestration (Airflow, Step Functions, or equivalent), and CI/CD for model artifacts.
- Monitoring in production — data drift and concept drift detection, population stability index, feature distribution monitoring, retraining triggers, and rollback strategy when a model degrades.
- Version control and reproducibility — Git, experiment tracking (MLflow, Weights & Biases, or equivalent), and environments where a result from six months ago can be reproduced.
Applied AI
- Familiarity with LLM application patterns — RAG, structured-output prompting, text-to-SQL — and where an LLM is the wrong tool for a scoring problem.
- Generating narrative explanations grounded in model output, and evaluating those narratives for faithfulness rather than fluency.
- Awareness of evaluation practice for generative systems: golden datasets, rubric-based scoring, and regression testing of prompts.
Experience & background
Required
- 3+ years applying machine learning and statistics to production problems, not just analysis notebooks.
- Advanced degree in Computer Science, Data Science, Statistics, Economics, or a related quantitative field — or equivalent demonstrated depth.
- Experience building models on patient- or member-level longitudinal data.
- Ability to explain a model's reasoning and its limits to a non-technical stakeholder without hedging into uselessness.
Strongly preferred
- Healthcare, specialty pharmacy, or life sciences experience — hub services, patient support programs, field reimbursement workflows, benefits investigation and prior authorization, or claims.
- Experience building models that drive human decisions rather than automated ones, and designing for human-in-the-loop review.
- Exposure to patient-level data governance — de-identification, tokenization, HIPAA constraints, and consent and aggregation rules — and comfort designing analyses within those limits.
- Experience contributing to product requirements alongside engineering and design teams.
What success looks like
- First 90 days — You understand the patient journey data end to end, know which signals are reliable and which are artifacts of reporting lag, and have a first risk model validated against historical outcomes.
- First year — Field teams work their queue in the order your models suggest, trust the explanations enough to repeat them to clients, and there's measurable evidence that acting on the recommendations changes outcomes.
Why this role
Specialty patient data is messy in ways that make it genuinely interesting: fragmented sources, inconsistent reporting cadence, small populations, censored outcomes, and real consequences on every case. You'll have unusual latitude to decide how the modeling problem gets framed, and your work reaches the people making the calls rather than sitting in an internal report.
Industry
- IT Services and IT Consulting
Employment Type
Click on Apply to know more.