Inference Infrastructure Engineer, Serving
Remarkable AI
- Salary
- $275k - $475k
- Experience
- 3+ yrs
- Location
- Palo Alto, California, United States
- Job type
- Full-time
Required skills
- vLLM
- TensorRT-LLM
- Triton
- SGLang
- C++
- CUDA
- Python
About the role
3+ years building low-latency, high-throughput inference serving systems; knowledge of quantization, batching, speculative decoding, KV cache; experience with vLLM/TensorRT-LLM/Triton/SGLang; multi-GPU model parallelism; C++/CUDA/Python; autoscaling and GPU cost optimization.
About Remarkable AI
AI lab building multimodal models for advanced visual reasoning.
This page is fully interactive when JavaScript is enabled. Please enable JavaScript to apply or browse related roles.