Website:
nava.com
Job details:
About Nava
Nava is building next-generation AI infrastructure and inference platforms that power enterprise AI at scale. We are looking for a
Principal Engineer – GPU Orchestration to lead the design and evolution of our GPU orchestration platform, enabling efficient scheduling, resource sharing, and multi-tenant inference workloads.
This role will own the orchestration layer that manages GPU resources for inference and small GPU clusters (2–4 node deployments), ensuring optimal utilization, rapid scaling, workload isolation, and an exceptional customer experience.
What You'll Do
GPU Orchestration & Scheduling
- Own the architecture and implementation of GPU scheduling for inference workloads and small GPU clusters (2–4 node deployments).
- Design intelligent workload scheduling algorithms that maximize GPU utilization while ensuring fairness and predictable performance.
- Manage GPU allocation, quotas, and resource sharing across multiple customers and workloads.
- Continuously optimize scheduling policies to improve efficiency and reduce infrastructure costs.
Kubernetes & Platform Engineering
- Own the Kubernetes orchestration layer for AI workloads.
- Design and maintain GPU-aware scheduling capabilities within Kubernetes.
- Improve cluster lifecycle management, workload placement, and infrastructure automation.
- Partner with Platform Engineering teams to continuously enhance cluster reliability and scalability.
Multi-Tenancy & Resource Isolation
- Design secure multi-tenant GPU environments with strong workload isolation.
- Implement resource quotas, admission controls, namespace isolation, and scheduling policies.
- Ensure consistent customer experience while maintaining high infrastructure utilization.
- Collaborate closely with the Security team on tenancy and access control mechanisms.
Autoscaling & Workload Optimization
- Build fast autoscaling capabilities for inference workloads.
- Reduce cold-start times through intelligent provisioning and pre-warming strategies.
- Optimize model placement based on workload characteristics, GPU availability, and latency requirements.
- Improve workload elasticity while balancing cost and performance.
Performance Engineering - Define and monitor platform KPIs including:
- GPU utilization
- Scheduling latency
- Cluster efficiency
- Autoscaling performance
- Cold-start latency
- Workload throughput
- Drive continuous improvements through benchmarking, performance tuning, and automation.
Cross-Functional Collaboration
- Partner with Compute & Inference Platform, GPU Cluster Engineering, Platform Reliability, AI Infrastructure Security, and Product teams.
- Support onboarding of new AI models and customer workloads.
- Contribute to platform architecture decisions across Nava's AI infrastructure.
Technical Leadership
- Serve as the technical authority for GPU orchestration and workload scheduling.
- Mentor senior engineers and contribute to engineering best practices.
- Drive architectural reviews, technical design discussions, and long-term platform strategy.
Success Metrics
You Will Be Measured On
- GPU utilization across the platform
- Scheduling efficiency and fairness
- Autoscaling responsiveness
- Cold-start latency
- Multi-tenant performance and isolation
- Platform reliability and scalability
- Customer workload performance
- Infrastructure cost optimization
Qualifications
Required Qualifications - 10+ years of experience in distributed systems, cloud infrastructure, Kubernetes, or platform engineering.
- Deep expertise in Kubernetes internals, scheduling, and container orchestration.
- Strong understanding of GPU resource management and AI infrastructure.
- Experience building large-scale scheduling and orchestration systems.
- Expertise in:
- Kubernetes
- Container runtimes
- Distributed systems
- Infrastructure automation
- Resource scheduling
- Multi-tenant platform architecture
- Performance optimization
- Strong software engineering skills in Go, Python, C++, or similar systems programming languages.
- Excellent problem-solving, architecture, and technical leadership skills.
Preferred Qualifications
- Experience with NVIDIA GPUs, CUDA, MIG (Multi-Instance GPU), GPU Operator, or Kubernetes device plugins.
- Familiarity with KServe, Ray Serve, Triton Inference Server, vLLM, Slurm, or similar AI infrastructure technologies.
- Experience with large-scale inference platforms, GPU cloud providers, or HPC environments.
- Knowledge of Kubernetes scheduler extensions, custom controllers, and operator development.
Why Join Nava?
- Build the orchestration platform powering one of the world's leading AI inference platforms.
- Solve challenging problems in GPU scheduling, resource optimization, and distributed systems.
- Work with world-class engineers building cutting-edge AI infrastructure.
- Shape the future of enterprise AI by enabling efficient, scalable, and secure GPU orchestration at global scale.
Skills: customer,architecture,infrastructure,gpu,design,utilization,kubernetes,cluster,orchestration,scheduling,isolation
Click on Apply to know more.