- Location
- Bengaluru, Karnataka, India
- Job type
- Full-time
Required skills
- LangChain
- Python
- BigQuery
- caching
- GitHub
- incident response
- Jira
- regression
- Slack
- SRE
About the role
GTMfund
Website:
gtmfund.com
Job details:
Responsibilities
- Agent Development: Build and maintain internal agents for PR review, issue triage, testing, go-live validation, customer calls, regression checks, incident response, anomaly detection, and post-mortem drafting. Design each agent with clear success metrics, feedback loops, monitoring, cost tracking, and kill criteria. Create reliable agent workflows with structured outputs, tool use/function calling, retries, fallbacks, and audit trails. Measure adoption, quality, false positives, latency, and time saved for every agent.
- Agent Platform and Integrations: Build integrations with GitHub, Slack, Jira/Linear, Notion, PagerDuty, CI/CD systems, and BigQuery. Own webhook/event-driven triggers for PRs, issues, deploys, alerts, and transcript availability. Create reusable patterns for prompt management, prompt versioning, evaluation datasets, and model selection. Keep LLM usage cost-disciplined through caching, routing by model tier, context management, and per-agent cost visibility.
- SRE AI Safety Net: Build incident response agents that query logs, deploy events, error spikes, and accuracy signals to generate first-line triage within minutes. Build post-mortem and anomaly detection agents using BigQuery-backed reliability data. Design safe runbook automation where irreversible production actions require human approval. Partner with SDET and DevEx engineers to define the quality bar and data foundation for SRE-AI workflows.
Requirements
- 3+ years of software engineering experience with at least 1 year building production LLM-powered systems or agents used by real users.
- Strong Python engineering skills; ability to write maintainable, tested, production-quality services.
- Hands-on experience with LLM APIs, preferably Anthropic Claude, including tool use/function calling, structured outputs, streaming, prompt caching, and multi-turn context handling.
- Experience with at least one agent/orchestration framework such as LangChain, LlamaIndex, DSPy, CrewAI, or strong custom orchestration experience.
- Strong REST API and webhook integration experience across tools such as GitHub, Slack, Jira/Linear, Notion, PagerDuty, or similar.
Click on Apply to know more.
This page is fully interactive when JavaScript is enabled. Please enable JavaScript to apply or browse related roles.