Website:
dsksolutions.net
Job details:
About us
We are a UK and India-based technology consultancy delivering IT strategy, architecture, and programme management for Tier 1 and Tier 2 enterprises across Spacetech, Telecommunications, Retail, Financial services and Utilities. Our clients include global satellite operators and major UK telecoms and retail brands, and our work is built on industry standards including TM Forum ODA, MEF, TOGAF, and ARTS. This role sits on our India delivery team, working directly with UK-based enterprise clients on live production systems.
Role summaryWe're looking for an MLOps Engineer to build and operate the infrastructure that takes UK clients' machine learning models from notebook to reliable production service — training pipelines, model registries, deployment automation, and monitoring. This is an infrastructure and platform role: you'll operationalise models that data science teams build, working from the same DevOps toolset and delivery model already proven on our infrastructure engagements, not a data science or research role.
Key responsibilities• Design and build ML training and deployment pipelines (data ingestion through to served model endpoint), with full CI/CD automation.
• Own model registry and versioning practices (MLflow, SageMaker Model Registry, or equivalent) so every production model is traceable to its training run and data.
• Containerise and deploy models to production using Kubernetes, ensuring scalable, cost-efficient inference serving (batch and real-time).
• Build monitoring and alerting for model performance in production — data drift, prediction drift, latency, and resource utilisation — not just infrastructure uptime.
• Automate retraining triggers and rollback procedures so degraded models are caught and addressed without manual intervention.
• Manage infrastructure-as-code for ML platforms (Terraform) across the cloud provider the client standardises on (AWS, Azure, or GCP).
• Partner directly with client-side data science teams to translate research code into production-grade, observable services.
• Own cost visibility and optimisation for compute-heavy training and inference workloads.
What we're looking forRequired
• 5-8 years in DevOps/platform/SRE engineering, with 2+ years specifically building and operating ML infrastructure or pipelines (not just general backend platform work).
• Deep, hands-on expertise in one major cloud platform's ML services (AWS SageMaker, Azure ML, or GCP Vertex AI).
• Real production Kubernetes and Docker experience, including deploying and scaling model-serving workloads specifically.
• Hands-on experience with an ML pipeline/orchestration tool (Kubeflow, Airflow, SageMaker Pipelines, Azure ML Pipelines, or equivalent).
• Terraform/IaC ownership for infrastructure provisioning, and CI/CD pipeline ownership for automated model deployment.
• Strong Python skills sufficient to work directly in data science teams' codebases, not just wrap them in infrastructure.
• Excellent written and spoken English — client-facing, no intermediary.
• Comfortable working a 1pm-10pm IST overlap window to align with UK client hours.
Nice to have
• Experience operationalising LLM/GenAI workloads specifically — vector databases, RAG pipeline infrastructure, model gateway/routing layers.
• Familiarity with feature store technologies (Feast, Tecton, or cloud-native equivalents).
• Experience with GPU infrastructure management and cost optimisation for training workloads.
• Exposure to responsible AI/model governance tooling (model cards, bias/fairness monitoring).
• Prior experience in telecoms, retail, or financial services environments.
What we offer• One of the fastest-growing specialisms in our practice — genuine first-mover exposure to a role most India-based consultancies aren't yet staffing seriously.
• Direct exposure to UK client production ML systems, working alongside data science teams rather than behind a ticket queue.
• A senior, embedded role within a small, high-trust delivery team, not a large offshore bench.
• Remote-first, India-based role with fixed UK-overlap hours.
• An above-industry-average package for working directly with top-tier UK clients.
Practical detailsHow to apply
Please send your CV, current CTC, notice period, and a short answer to the following screening question to recruitment@dsksolutions.net: Describe a production ML model you helped deploy and monitor — what broke after go-live (data drift, latency, cost, or otherwise), and how did your monitoring catch it?
Click on Apply to know more.