VirtualMaze
Website:
virtualmaze.com
Job details:
We're Hiring | AI Platform Engineer (Model Deployment & Inference)
At VirtualMaze Softsys Pvt. Ltd., we're building scalable AI platforms that power real-world applications across Computer Vision, Large Language Models (LLMs), and intelligent automation. We're looking for passionate engineers who can bridge the gap between AI and infrastructure by deploying, optimizing, and managing production-grade AI inference systems.
As an AI Platform Engineer, you'll work with cutting-edge technologies to deploy AI models at scale, optimize GPU performance, and build cloud-native AI infrastructure that delivers high-performance, low-latency inference.
What You'll Do
✅ Deploy and manage Computer Vision, NLP, and LLM models in production
✅ Build scalable real-time and batch inference services
✅ Optimize AI inference using NVIDIA Triton Inference Server, TensorRT, ONNX Runtime, and TorchScript
✅ Manage GPU-enabled Kubernetes infrastructure with Docker and Helm
✅ Build CI/CD pipelines for automated AI model deployment
✅ Monitor and optimize GPU utilization, latency, throughput, and inference performance using Prometheus and Grafana
✅ Collaborate with AI engineers to deliver reliable and scalable AI solutions
Required Skills
✔ NVIDIA Triton Inference Server
✔ TensorRT & ONNX Runtime
✔ PyTorch or TensorFlow Deployment
✔ Docker, Kubernetes & Helm
✔ Python, Bash & Linux
✔ CUDA Fundamentals & GPU Optimization
✔ REST APIs
Nice to Have
⭐ Experience with vLLM, TensorRT-LLM, NVIDIA NIM
⭐ Ray Serve, KServe, Seldon Core
⭐ Kubeflow or MLflow
⭐ Computer Vision or LLM Deployment
⭐ Model Quantization (FP16/INT8)
⭐ Distributed Inference Systems
📍 Location: Coimbatore & Hyderabad
💼 Experience: 3–6 Years
📧 Send your resume to: hr.cbe@virtualmaze.co.in
📱 WhatsApp: +91 9159359391
Click on Apply to know more.