We are seeking an experienced AI/ML Engineer to lead the development and ownership of an on-device speech processing stack. This role involves working on a large-scale language model running efficiently on mid-range Android devices, managing a full speech-to-text (STT) and text-to-speech (TTS) pipeline with a target latency of 5–7 seconds, all operating privately on the device.
Key Responsibilities:
- End-to-end ownership of the on-device speech stack including a 3 billion parameter speech language model.
- Optimization of speech pipelines for mobile latency and performance.
- Support for multiple dialects including Hindi, Hinglish, Tamil, Telugu, and Marathi.
- Implementation of guardrails and evaluation harnesses for system robustness.
Requirements:
- 7 to 10 years of experience in AI/ML engineering.
- Proven production experience with on-device large language models, specifically using tools such as llama.cpp, MediaPipe, and TensorFlow Lite.
- Strong expertise in speech processing pipelines.
- Experience optimizing mobile applications for latency and performance.
- Ability to work remotely and collaborate effectively.
Location: India
Compensation: Competitive salary ranging from 70 to 120 LPA.