Website:
avashya.tech
Job details:
About Avashya
Avashya is a new-age, AI-first, platform-led services company, founded by engineers and specialists who spent years building at AWS and Microsoft. AI and our own platforms sit at the core of how we deliver, not bolted on around the edges. Voice AI is one of those platforms, built and owned in-house, and it's the one this role runs and ships.
The Opportunity
There's a wide gap between a voice AI demo and a voice AI product. A demo tolerates a two-second pause and a script that only survives the happy path. A product holds up under concurrent real calls, degrades gracefully when a provider has a bad day, and moves fast enough that the person on the other end forgets they're talking to a machine. You'll own both sides of closing that gap: the platform, and every client implementation built on it.
What You'll Be Doing
The platform
- Extend our real-time voice framework, and when configuration hits its limit, get underneath it and fix the runtime directly.
- Turn what you learn on one client build into something the next client starts with.
The implementations
- Take a client's use case, outbound negotiation, multilingual support, appointment booking, and ship it as a working, production-grade system, from scoping through go-live.
- Tune the pipeline (ASR, LLM, TTS, RAG, tool-calling, orchestration) to the client, the languages, and the call volume in front of you, not a default stack.
- Stay the owner after launch. If it's live, you know why it behaves the way it does.
Live call quality
- Know what every stage of a conversation is allowed to cost, in milliseconds, and hold that line as load and providers shift.
- When a call feels off, find out why fast, before a client has to tell you.
- Gate releases on measured evals, not a provider's word that quality held after their last update.
Tradeoffs and standards
- Different flows want different tradeoffs, speed for a negotiation bot, polish for a support agent. You make that call per client.
- Review code and mentor as the team grows.
Prerequisites :
NOT years of experience, depth. Here's what we'll go deep on:
Architecture
- Cascaded vs. speech-to-speech: where a cascaded ASR/LLM/TTS pipeline wins on control and swapability, where end-to-end wins on latency and prosody, and which fits which use case.
- Orchestration frameworks: real hands-on depth in Pipecat, LiveKit Agents, or comparable, past the quick-start and into their internals and limits.
- Turn-taking: VAD tuning, endpointing, barge-in handling, keeping interruptions natural instead of jarring.
- Transport and telephony: WebRTC, SIP trunking, jitter buffers, codec tradeoffs (Opus, PCM, mu-law), PSTN vs. browser differences.
Latency, end to end
- Budget decomposition: ASR partials, LLM time-to-first-token, TTS time-to-first-audio-byte, network hops, tool-calling round trips, and exactly where the milliseconds go.
- Streaming everything: streaming ASR, LLM, and TTS working in concert, plus techniques (fillers, early commits, speculative synthesis) to keep a conversation alive during a slow tool call.
- Retrieval and tool-calling under budget: grounding responses via RAG or live tool calls without breaking the sense of a live conversation.
Evaluation and guardrails
- Voice-specific evals: ASR accuracy (WER), TTS naturalness, end-to-end conversation quality, not a text eval bolted onto a voice product.
- Latency as a tracked metric: P50/P95/P99 per stage, per provider, over time.
- Guardrails for a live channel: content safety and hallucination control where you can't re-render a bad response, PII handling in transcripts and recordings, consent and data residency for recorded calls.
Compensation
Top-of-market. Pay reflects depth, not tenure.
Click on Apply to know more.