Liquidnitro Games
Website:
liquidnitro.games
Job details:
AI/ML Engineer (Real-Time Conversational & Commentary Systems, Games)
Location: Hyderabad
Type: Full Time
About Us
Liquidnitro Games is India’s flagship live services and game production company, founded by industry veterans with a proven track record in producing massively successful games & live services. For game companies, studios, and publishers worldwide – we offer world-class game development expertise to power creativity, growth, and profitability in their games.
About the role
We are building AI systems that make game characters and commentators feel genuinely alive: they react to what is happening, stay in character, remember what came before, and respond fast enough to feel natural in real time. This is a rare, high-ownership role at the intersection of applied LLM and voice AI and shipping game development. You will architect and build these systems end to end and see them ship in players' hands.
You will lead the AI systems behind two flagship workstreams:
- Live AI commentary. Dynamic, personalized play-by-play from AI host personas that react to gameplay in real time, blend pre-recorded voice-over with generated lines, and build narrative continuity (rivalries, streaks, running gags) across a session.
- Conversational characters. Real-time, voice-driven conversations between players and in-game characters. The character must stay in persona, remember the exchange, respond with near-instant latency, and handle natural interruptions (barge-in) like a real conversation.
Both are latency-critical, persona-critical, and safety-critical. If you are excited by making generative AI feel instant and in-character inside a real-time engine, this is your kind of problem.
What you'll do:
- Architect the stack end to end - LLM orchestration, voice pipeline, session memory, safety, and game-engine integration - as reusable systems that serve more than one title.
- Own the latency budget. Design for perceived-instant response using anticipatory / speculative generation, caching, streaming synthesis, and select-at-the-moment patterns so the AI never lags the action or the conversation.
- Build the real-time voice pipeline: speech-to-text, voice-activity detection, streaming text-to-speech, endpointing, turn-taking, and barge-in / interrupt handling.
- Build the voice layer. Integrate and tune expressive neural text-to-speech and custom or licensed voice models so hosts and characters sound like themselves, always within voice licensing and consent constraints.
- Engineer persona fidelity and emotional range. Prompt design, structured outputs, and memory systems that keep hosts and characters convincingly in-voice and on-brand across long sessions, with prosody and emotion control so delivery matches the moment: hyped on a photo finish, in-mood in conversation.
- Ground the models in canon. Use retrieval and knowledge grounding so a character stays true to its source-material lore and commentary stays accurate to live game state, rather than fabricating.
- Ship a content-safety pipeline: moderation, guardrails, jailbreak resistance, and a deterministic fallback ladder so output is always safe and never stalls or goes off-brand. Document it to production and partner standards.
- Keep it model- and voice-agnostic. Abstract LLM and TTS providers so models and licensed voices can be swapped with zero rearchitecting.
- Integrate with the engine and pipeline: gameplay event hooks, audio middleware, and audio-driven facial animation / lip-sync for generated lines.
- Build evaluation and observability for generated content: quality, persona consistency, latency SLOs, and cost per session.
- Optimize inference cost at scale via caching, tiered generation, distillation, and cloud-vs-on-device trade-offs.
- Plan for localization. Design the pipeline so generated dialogue and voice can extend to multiple languages, including localizing runtime-generated lines and not only pre-scripted ones.
- Collaborate closely with gameplay, audio, narrative / writing, and animation teams, and mentor engineers as the AI group grows.
What we're looking for:
- 5+ years of software engineering, with hands-on production experience building LLM-powered and / or real-time voice / conversational systems (not only demos).
- Strong LLM application skills: prompt engineering, retrieval, function / tool calling, structured outputs; practical familiarity with commercial and open models and with serving / inference stacks.
- Real-time audio / voice experience: ASR / STT, TTS (including expressive and custom voice models), VAD, streaming, and aggressive latency optimization.
- Proficient in Python plus a systems or game language (C++ or C#).
- Cloud and services: designing low-latency APIs and distributed services (AWS preferred); comfort reasoning about latency, caching, and concurrency.
- Content moderation / safety awareness and a track record of building reliable fallbacks around non-deterministic models.
- A product mindset: you prototype fast to answer the risky questions, then productionize what works.
Bonus points:
- Game engine experience - Unreal Engine 5 (Chaos, MetaHuman / audio-driven facial) or Unity - and having shipped a game or a real-time interactive product.
- Audio middleware (Wwise / FMOD) and game-audio integration.
- Dialogue systems, interactive narrative, or NPC / conversational-agent work.
- MLOps, fine-tuning, model distillation, or on-device inference.
- Experience with LLM evaluation frameworks and building eval harnesses.
Click on Apply to know more.