CloudKeeper
Website:
cloudkeeper.com
Job details:
Designation: Lead MLOps Engineer (GPU Optimization)
About CloudKeeper
CloudKeeper is a cloud cost optimization partner that combines the power of group buying & commitments management, expert cloud consulting & support, and an enhanced visibility & analytics platform to reduce cloud cost & help businesses maximize the value from AWS, Microsoft Azure, & Google Cloud.
A certified AWS Premier Partner, Azure Technology Consulting Partner, Google Cloud Partner, and FinOps Foundation Premier Member, CloudKeeper has helped 350+ global companies save an average of 20% on their cloud bills, modernize their cloud set-up and maximize value — all while maintaining flexibility and avoiding any long-term commitments or cost.
CloudKeeper hived off from TO THE NEW, a digital technology service company with 2500+ employees and an 8-time GPTW winner.
To know more, please visit - https://www.cloudkeeper.com/
Responsibilities
- Drive R&D and engineering for AI Infrastructure optimization within CloudKeeper's FinOps for AI platform — building the Tuner AI / Commit AI capability on GPU and ML workloads
- Design and build optimization engines for GPU right-sizing, idle shutdown, spot migration with checkpoint/resume automation, inference batching, quantization, and model placement
- Extend the optimization stack to LLM-era workloads — caching, model routing, dynamic batching, prompt optimization, RAG-aware architectures
- Partner with the Lens AI team to translate GPU and ML workload signals into actionable, dollar-quantified optimization recommendations for customers
- Work cross-functionally with product, platform, and customer success teams to ship optimization features end-to-end (data ingestion → optimization engine → customer-facing recommendation)
- Lead technical direction for AI workload optimization, set engineering standards, and mentor the ML / MLOps engineering bench as the AI Infrastructure pillar scales
- (Lead level) Hire, ramp, and grow a team of ML infrastructure engineers as headcount expands
Must Have
- B.E / B.Tech / M.Tech / MCA with 7+ years of hands-on engineering experience
- Production experience with GPU workloads — has measurably optimized GPU utilization, throughput, or cost in a real production environment, not just academic / lab work
- Strong performance engineering background — must come ready with a concrete optimization story including before/after metrics (latency, throughput, or cost reduction)
- Strong Python + Linux + systems fundamentals
- Solid understanding of the ML model lifecycle — training, serving, inference — able to reason about what is running on the GPU and why
- MLOps fluency — model deployment, monitoring, observability, GPU cluster operations
- Hands-on with cloud GPU instances (AWS P5 / G6, Azure ND series, GCP A3, or equivalent) and Kubernetes-based GPU orchestration (EKS / AKS / GKE GPU node pools, Karpenter, Run:ai, NVIDIA GPU Operator, or similar)
- Familiarity with at least one modern LLM inference framework — vLLM, TGI, Triton, SGLang, Ray Serve, or BentoML
- Strong communication skills — able to translate deep technical optimization into customer / business outcomes
- (Lead level) Experience managing or technically leading a team of 3+ engineers
Good To Have
- Deep LLM-era optimization expertise — KV caching, semantic caching, model routing, dynamic batching, quantization (FP16 → INT8 → INT4), model distillation, structured outputs
- Familiarity with LLM workload patterns — RAG, agents, embeddings, vector databases (Pinecone, Weaviate, Qdrant)
- CUDA, NCCL, mixed-precision training and inference
- Experience with managed ML training platforms — SageMaker, Azure ML, Vertex AI, Databricks Mosaic
- Exposure to GPU-native clouds — CoreWeave, Lambda Labs, RunPod, Crusoe
- Open source contributions to ML infrastructure projects — vLLM, llama.cpp, TGI, Ray, Triton, KubeRay
- Adjacent experience in cloud cost optimization / FinOps — Spot.io, ScaleOps, Granulate, CAST AI
- Comfort with Agile methodology and modern engineering practices (CI/CD, code review, observability)
Skills:- Graphics Processing Unit (GPU), MLOps and Large Language Models (LLM)
Click on Apply to know more.