Website:
Job details:
We are an emerging, innovation-driven organization focused on building advanced voice and conversational AI experiences. Our team works with cutting-edge open source technologies and modern frameworks to deliver human-like speech interfaces across products and platforms. We value experimentation, rapid iteration, and practical problem-solving, encouraging team members to influence architecture, tooling, and product direction. Collaboration, knowledge-sharing, and continuous learning are central to how we work. Joining our team offers the opportunity to help shape next-generation voice agents used by a diverse global audience.
Role Description This is a full-time remote role for an Expert specializing in humanizing and optimizing open source Text-to-Speech (TTS) systems and Kokoro-based voice agents. The role involves designing, fine-tuning, and deploying natural, expressive voices; improving prosody, intonation, timing, and clarity; and optimizing TTS pipelines for quality, latency, and scalability. The expert will evaluate and compare open source TTS models, integrate and customize Kokoro voice agents, and create reusable components and best practices for product teams. Day-to-day tasks include experimentation with model parameters, building evaluation frameworks, collaborating with engineers and product owners, documenting guidelines, and iterating on voice experiences based on user feedback and metrics.
Qualifications
- Strong expertise in Text-to-Speech systems, including open source TTS frameworks, voice synthesis, and prosody modeling.
- Experience customizing and optimizing Kokoro or similar voice agent architectures for realism, responsiveness, and robustness.
- Proficiency in programming and scripting (e.g., Python), including model configuration, pipeline automation, and integration with APIs and services.
- Background in signal processing, speech science, or machine learning as applied to audio, with a focus on voice quality and intelligibility.
- Skills in designing and running subjective and objective voice quality evaluations, including listening tests, metrics, and A/B experiments.
- Ability to collaborate with product, design, and engineering teams to translate user experience goals into technical voice requirements.
- Comfort working in remote, distributed teams, with clear communication, documentation, and self-directed workflow management.
- Bachelor’s or advanced degree in Computer Science, Electrical Engineering, Linguistics, Speech Technology, or related field, or equivalent practical experience.
- Experience deploying voice solutions in production environments (cloud platforms, microservices, or real-time applications) is highly beneficial.
- Familiarity with privacy, accessibility, and ethical considerations in voice and conversational AI is a plus.
Click on Apply to know more.