- Design, build, and maintain balanced Data, ML, and AI engineering pipelines to automate data provisioning, model deployment, and enterprise workflow execution.
- Implement and support robust LLMOps and MLOps practices across data, machine learning, and Generative AI systems to automate model evaluation, monitoring, and CI/CD workflows.
- Manage, optimise, and scale Databricks workspace configurations, clusters, and jobs for enterprise data processing and AI/ML workloads.
- Collaborate cross-functionally with data engineers, ML engineers, software engineers, and product leads to design, deploy, and scale data pipelines, feature stores, and ML serving systems into production.
- Implement proactive incident response, event instrumentation, and self-healing mechanisms to detect and remediate system anomalies or data/model quality issues.
- Provide day-to-day operational support, infrastructure upgrades, capacity planning, and cloud resource optimization for data platforms and ML infrastructure.
- Work closely with IT DevOps, SRE, and security teams to enforce governance, data lineage, compliance, and enterprise CI/CD deployment standards.
- Promote engineering best practices, conduct code reviews, participate in on-call rotation support, and contribute to knowledge sharing across teams.
- Develop hands-on Proofs of Concept (POCs) for modern data platforms, feature stores, and real-time streaming tools in collaboration with product and analytics teams.
- Maintain agility towards evolving technology stacks across DataOps and MLOps platforms to continually modernise infrastructure.
Candidate Attributes
- Strong problem-solving mindset
- Exceptional cross-functional communication
- A track record of driving collaborative DevOps/DataOps/MLOps culture
- Solid understanding of modern DevOps, MLOps, and DataOps methodologies, including CI/CD automation, model governance, and observability tools (e.g., MLflow, Weights & Biases, LangSmith)
- Hands-on experience with core cloud data & ML services on AWS (e.g., SageMaker, Glue, EMR, Athena, S3)
- Strong expertise in Databricks management, optimisation, and workspace administration
- Strong understanding of database fundamentals, replication, relational/NoSQL databases, and vector databases (e.g., Pinecone, FAISS, Milvus, Weaviate)
- Proven experience building CI/CD automation pipelines for containerised Python, Java, or Scala microservices and ML serving systems
- Proficient with Git version control and standard branching workflows
- Practical experience deploying, monitoring, and debugging distributed data pipelines, ETL workflows, and model deployment systems
- Familiarity with configuration management and provisioning tools (e.g., Terraform, CloudFormation, Ansible)
- Scripting proficiency in one or more languages (Python, Bash, or JavaScript)
- Practical experience with containerization using Docker and exposure to Kubernetes orchestration
- Solid grasp of machine learning lifecycle, deep learning frameworks (PyTorch, TensorFlow), NLP/CV concepts, and feature engineering workflows
- 2 – 5 years of experience across DataOps, MLOps, ML Engineering, or Data Engineering in enterprise cloud environments