Website:
nava.com
Job details:
Role & Responsibilities
- Design and deploy low-latency inference pipelines for LLMs, diffusion models, and vision transformers across CPU, GPU, and NPU architectures.
- Optimize model serving stacks using Triton Inference Server, vLLM, or TGI—tuning batch sizes, quantization, and memory layout for peak performance.
- Containerize and orchestrate inference services via Docker and Kubernetes, ensuring high availability and auto-scaling under fluctuating workloads.
- Implement model monitoring, health checks, and A/B testing frameworks to validate performance and drift in production.
- Collaborate with ML Engineers to convert trained models into production-ready formats (ONNX, TensorRT, GGUF) with minimal accuracy loss.
- Build observability dashboards (Prometheus/Grafana) and alerting logic to detect and mitigate inference bottlenecks in real time.
Skills & Qualifications
Must-Have
- Python
- Docker
- Kubernetes
- Triton Inference Server
- ONNX Runtime
- PyTorch
- TensorRT
- Prometheus
- Grafana
- CI/CD (GitHub Actions, GitLab CI)
Preferred
- Experience with vLLM or TGI
- Knowledge of Model Quantization (AWQ, GPTQ, GGUF)
- Familiarity with NVIDIA Triton model ensemble pipelines
Skills: cuda,architecture,building,foundation,ml,management,models,distributed systems,infrastructure,inference
Click on Apply to know more.