nugget.ai
Website:
nugget.ai
Job details:
Job Summary
We are looking for a Voice Systems Engineer to build and optimize real-time, GPU-accelerated voice AI systems using NVIDIA technologies. The role focuses on speech-to-text, text-to-speech, conversational AI, low-latency inference, and integrating NVIDIA speech/AI models into enterprise applications.
Key Responsibilities
- Build real-time voice AI pipelines for speech-to-text, LLM reasoning, and text-to-speech.
- Work with NVIDIA GPU technologies and optimize inference for low latency and high throughput.
- Integrate NVIDIA speech and generative AI models into enterprise AI platforms.
- Work with NVIDIA NIM, NeMo, TensorRT, Triton Inference Server and CUDA/PyCUDA.
- Build streaming voice APIs using Python, FastAPI, WebSockets/WebRTC or similar technologies.
- Optimize GPU utilization, concurrency, batching and real-time inference performance.
- Integrate voice agents with enterprise applications, databases and APIs.
- Package and deploy production-ready voice AI solutions for enterprise environments.
Required Skills
- Strong Python and backend development experience.
- Experience with speech AI / ASR / TTS / conversational AI.
- NVIDIA GPU ecosystem experience: CUDA, TensorRT, Triton, NIM or NeMo.
- Experience with LLMs and real-time inference.
- WebSockets, WebRTC or streaming architectures.
- Docker and Kubernetes.
- Experience optimizing latency, throughput and GPU utilization.
Click on Apply to know more.