FirstHive
Website:
firsthive.com
Job details:
Designation : Data Science Engineer
Location : Bengaluru
Experience : 4 - 6 years
Function : AI & Data Science
Role Description
We are seeking a Data Science Engineer to build and deploy production ML models and AI features for our CDP platform. You will work in a small, high-ownership AI & Data Science team building customer segmentation models, entity resolution algorithms, predictive analytics, NLP capabilities, and LLM-powered automation that directly impact how enterprise clients understand and engage with their customers. This is a hands-on engineering role you build models that ship to production, not notebooks that stay in research.
Key Responsibilities
- Build and deploy customer segmentation and clustering models (K-Means, DBSCAN, hierarchical) at scale.
- Develop entity resolution algorithms fuzzy matching, blocking strategies, probabilistic scoring to unify customer profiles across disparate data sources.
- Build predictive models churn prediction, conversion propensity, next-best-action recommendations using classification and regression (XGBoost, Random Forest, logistic regression).
- Design and build LLM-powered features schema mapping automation, natural language querying, AI-driven insight generation using prompt engineering, RAG pipelines, and structured output extraction.
- Build NLP capabilities text embeddings, semantic similarity, entity extraction, text classification using transformers (BERT or similar).
- Write complex SQL for feature engineering window functions, sessionization, time-series aggregation, customer behavior features from raw event data on Snowflake/BigQuery.
- Integrate ML models into the core platform via APIs (FastAPI) for real-time and batch inference.
- Own model lifecycle in production monitoring, drift detection, retraining, versioning.
- Work with data engineering teams to ensure clean, structured training data and feature pipelines.
Experience And Skills
- Python ML stack : scikit-learn, Pandas, NumPy, XGBoost. This is 70% of the work.
- LLM / GenAI : prompt engineering, RAG fundamentals, embeddings, vector similarity search, calling LLM APIs (Claude, OpenAI, or similar) with structured outputs. Not fine-tuning effective use of APIs.
- SQL : complex feature extraction queries on analytical databases. Window functions, sessionization, time-series aggregation, cost-aware query patterns. Not basic SELECT.
- NLP : text embeddings (sentence-transformers or similar), named entity recognition, text classification, semantic search.
- Model deployment : FastAPI or Flask, Docker containerization, REST API serving.
- Entity resolution / record linkage fuzzy matching (Levenshtein, Jaro-Winkler), blocking strategies, probabilistic matching across multiple fields.
- Model evaluation precision/recall trade-offs, cross-validation, A/B testing.
Good To Have
- LangChain, vector databases (Pinecone, FAISS).
- MLflow or experiment tracking.
- Kafka consumers for real-time scoring.
- Snowflake ML / BigQuery ML.
- Time-series forecasting (Prophet, ARIMA).
- Customer analytics or MarTech platform experience.
(ref:hirist.tech)
Click on Apply to know more.