NextBillion.ai
Website:
nextbillion.ai
Job details:
Role Requirements & Job Description
We build and operate V-Assistant (VisionAI + V-Hub) — a production conversational AI platform that answers natural-language questions over live fleet-telematics data (V-Track) and powers an L&D knowledge-base product. It is not a demo: five environments (dev/test/qa/uat/prod), Azure-native infrastructure behind Front Door, versioned releases, and a distributed engineering team of 8–9 shipping ~250–550 commits a month through Azure DevOps PRs.
The technical core is a LangGraph tool-calling agent: query rewriting (CQR) NeMo Guardrails embedding-based tool routing over ~29 domain tools agent loop with Postgres checkpointing formatter pass streamed response with follow-up generation. Around it: pgvector + BM25 hybrid RAG, Azure OpenAI with a homegrown tokens-per-minute budget/throttling subsystem, Langfuse tracing, and an unusually deep quality harness (LLM-as-judge evals via DeepEval, ground-truth datasets, Playwright e2e suites gating releases).
The practical consequence: quality here is an empirical, tested property of a non-deterministic system — regressions arrive via prompts, model versions, tool schemas, and token budgets, not just code. Leading this repo means owning that discipline, not just the codebase.
Role Summary
Own technical direction, delivery, and production quality for the V-Assistant monorepo, while leading a distributed offshore team of 6–10 engineers and QA. This is a hands-on lead role: you review the agent core and prompt pipelines yourself, make the architecture calls, run incidents, and are accountable for what ships to prod — you don’t just track tickets.
Key Responsibilities
1. Agentic system architecture & technical oversight
- Own the design of the agent pipeline: tool routing, guardrails, context/token management, streaming, conversation persistence, follow-up generation.
- Review and guide changes to the LangGraph agent core, prompts, tool schemas, and RAG retrieval — and reason about their blast radius on answer quality.
- Grow the assistant’s capabilities: design and ship new tools that expose additional upstream API functionality, integrating via generated OpenAPI/Swagger clients against the internal telematics REST API.
- Own multi-turn conversation quality: query rewriting (CQR), conversation memory/persistence, follow-up-question generation, and intent routing.
- Make build/buy/framework calls (e.g., LangGraph vs. alternatives, model routing, provider expansion beyond Azure OpenAI) and record them as ADRs.
- Manage scarce LLM capacity deliberately: token budgets, TPM throttling, schema slimming, prompt caching, cost-per-conversation.
2. LLM quality, evals & observability
- Own the eval strategy: ground-truth datasets, DeepEval/LLM-as-judge gates, regression suites per release — and know when a metric delta is signal vs. noise.
- Own tracing/observability (Langfuse) and use it: debug “answers got worse” reports back to the responsible prompt/model/config/data change.
- Keep guardrails (prompt injection, out-of-scope, safety) effective as the tool surface grows, and keep the guardrails config (Colang rails, embedding rails) and the fast-moving LLM dependency stack (LangChain/LangGraph, NeMo Guardrails) current — including the CVE-pinning / security-scanning discipline the codebase already enforces.
3. Production operations, release & incident management
- Own releases across five environments through Azure DevOps pipelines: scope, DB-migration sequencing (Alembic), rollback plan, go/no-go.
- Run production incidents end-to-end: triage scope/severity, localize root cause across agent/API/infra layers, make the hotfix-vs-rollback call (the history shows both — you must be comfortable rolling back), drive prevention.
- Close current operational gaps: restore missing rollback/e2e/promote pipeline definitions, release tagging discipline, environment-parity in IaC.
4. Team leadership & delivery
- Lead a distributed offshore team across time zones: PR/ticket discipline, code-review standards (the repo runs mypy strict + heavy lint — keep that bar), async communication. Day-to-day feature work is executed offshore; the lead sets the technical bar, reviews the critical paths, and owns outcomes.
- De-risk key-person concentration: today one contributor owns ~⅓ of all commits; spread ownership deliberately.
- Prioritize across incidents, roadmap features, and quality debt; be the technical point of contact for product stakeholders and set realistic expectations for what LLM systems can and cannot promise.
Required Qualifications
- 8+ years building and operating production software, with at least 2 years hands-on on LLM-based / agentic systems in production (not prototypes or wrappers): tool-calling agents, multi-step orchestration, streaming UX.
- Deep, demonstrable agentic AI experience — candidate must be able to speak concretely to: agent state/checkpointing and resumability, tool routing when the tool count exceeds the schema token budget, hierarchical token budgets and runaway-loop defenses, guardrails against prompt injection, and graceful degradation. “Built a RAG chatbot” alone does not meet this bar.
- LLM evaluation rigor: has built or owned eval suites (LLM-as-judge, ground-truth regression, per-release gates) and can reason about statistical noise in small eval sets.
- RAG at production depth: embeddings, hybrid retrieval (BM25 + dense, fusion), pgvector or equivalent, chunking/index trade-offs.
- Strong Python (FastAPI, asyncio/httpx, async SQLAlchemy, Pydantic v2, strict typing) — able to credibly review an 800-line agent core, not just read it.
- Experience integrating and wrapping external REST APIs, including OpenAPI/Swagger-generated clients, and hardening those integrations (the upstream data source is a third-party telematics API).
- Cloud + DevOps: Azure preferred (Azure OpenAI, Container Apps, Front Door), Azure DevOps or equivalent CI/CD, Terraform/OpenTofu literacy.
- Production incident and release management track record on a live product.
- 2+ years leading engineers (lead/staff/EM), ideally distributed teams across time zones; strong written communication.
- Working TypeScript/React fluency — enough to review a streaming chat widget and an admin SPA, not necessarily to build them.
Preferred / Nice-to-Have
- LangGraph/LangChain specifically (our stack), NeMo Guardrails, Langfuse/LangSmith, DeepEval or similar eval frameworks.
- Multi-provider LLM experience (Anthropic, Gemini) — provider expansion is on the roadmap.
- Observability depth: structured logging, tracing, token/cost telemetry.
- Domain exposure: telematics, fleet/field-service, logistics, or other operational-data analytics.
What Success Looks Like
- 30 days: Has mapped the agent pipeline, the 29-tool surface, the eval harness, and the five-environment release path; has personally reviewed PRs on the agent core; has a written list of the top quality and operational risks.
- 60 days: Running releases and incidents independently; eval gates enforced on every release; the missing rollback/e2e pipeline definitions restored; review and ownership load visibly spread beyond the current key person.
- 90 days: Measurable movement — fewer production hotfix/rollback cycles, faster root-cause on quality regressions, a costed roadmap for the agent platform (model routing, provider expansion, cost-per-conversation) that stakeholders have signed off on.
Click on Apply to know more.