We are seeking a Founding Inference Engineer to join an early-stage AI infrastructure startup based in Bengaluru. This role offers the opportunity to work closely with the founding team and have a significant impact on the performance and cost efficiency of GPU-based AI inference systems.
Responsibilities:
- Own and optimize cost per token for GPU inference workloads.
- Work on vLLM serving, kernel optimization, and throughput improvements on GPUs.
- Maximize tokens processed per unit cost on GPU infrastructure.
- Deep dive into GPU inference performance engineering to improve capacity planning beyond simple metrics.
Requirements:
- 3 to 6 years of experience in GPU inference or performance engineering.
- Strong understanding of GPU kernels, throughput optimization, and inference serving.
- Experience working with large-scale AI infrastructure and cost optimization.
Location: Bengaluru
Compensation: 60 - 90 LPA