Trigent Software Inc
Website:
trigent.com
Job details:
About the Role
As a member of Cloud Infrastructure team, you will design, build, and operate the AWS foundation. You will own complex infrastructure services end-to-end, drive automation first practices, and help evolve our “paved road” for engineers building on AWS. A key focus for this role is to apply AI/ML and AI agents to Cloud Infrastructure operations – using single and multi-agent systems and orchestration frameworks to automate diagnostics, runbooks, and workflows while maintaining strong guardrails around security, reliability, and cost. You will collaborate closely with Cloud Infrastructure peers, SREs, Security Engineering, and product teams across geographies.
Responsibilities
- Cloud Infrastructure Engineering (AWS)
- Design, implement, and operate highly available, secure, and scalable AWS infrastructure (e.g., VPC, Transit Gateway, EC2, Load Balancing, S3, EBS/EFS/FSx, Route 53, IAM, KMS, Backup).
- Build and maintain infrastructure-as-code using tools such as AWS CDK / CloudFormation, enforcing standards, guardrails, and reusable patterns.
- Develop automation and tooling (primarily in Python/TypeScript) to remove repetitive operational work (provisioning, patching, configuration, cleanup, compliance checks, reporting).
- Contribute to and sometimes lead design reviews, architecture discussions, and RFCs for new or evolving infrastructure services.
- Partner with Security and Compliance to meet security, audit, and regulatory requirements across accounts and regions.
- AI Agents, Orchestration & Multi-Agent Systems
- Identify high-value Cloud Infra workflows (e.g., incident triage, change impact analysis, runbook execution, capacity/cost recommendations) that can be automated using AI agents.
- Design and implement agentic workflows (single and multi-agent) using modern AI orchestration patterns and frameworks (e.g., tool-calling, planners, evaluators, guardrails).
- Integrate agents with existing cloud APIs, observability tools, ticketing systems, and runbooks to provide end-to-end, human-in-the-loop automation.
- Define and enforce safety, security, and approval guardrails for AI-driven actions (RBAC, policy checks, dry-runs, explicit approvals, audit logging).
- Measure and communicate impact of AI automation (MTTR reduction, hours saved, error reduction, cost optimization, improved engineer experience).
- Reliability, Operations & On-Call
- Own the reliability and performance of services you build – from design through deployment and production operations.
- Implement and tune monitoring, logging, alerting, and SLO/SLA dashboards for Cloud Infra services (Datadog/Splunk/CloudWatch or similar).
- Participate in the on-call rotation, lead troubleshooting for complex AWS infrastructure incidents, and drive post-incident reviews and preventative improvements.
- Proactively identify technical debt and reliability risks in infrastructure and drive remediation plans.
- Collaboration, Mentoring & Best Practices
- Act as a technical mentor to Engineer I/II teammates on AWS fundamentals, automation patterns, and AI-driven operations.
- Help define and evolve paved road standards for AWS infrastructure, automation, and AI agent usage across Cloud Infrastructure.
- Contribute to runbooks, design docs, knowledge base articles, and internal training sessions, including AI and automation best practices.
Qualifications
- 5 to 15 years of hands-on experience in Cloud / Infrastructure Engineering, or similar roles, with strong focus on AWS.
- Deep understanding of core AWS services: VPC & networking (subnets, routing, TGW, VPN/Direct Connect, security groups, NACLs), EC2, Auto Scaling, Load Balancing, S3, EBS/EFS/FSx, Route 53, IAM, KMS, CloudWatch/CloudTrail, and Backup.
- Strong experience with Infrastructure-as-Code (AWS CDK, CloudFormation, or Terraform) and Git-based workflows (branching, PR reviews, CI/CD).
- Solid programming skills in at least one language commonly used for infra-automation, such as Python or TypeScript/Node.js.
- Proven track record designing and operating production-grade, multi-account/multi-region AWS environments with a focus on security, reliability, and cost.
- Experience implementing observability for infrastructure services (metrics, logs, traces, alerting, dashboards).
- Demonstrated ability to own complex projects end-to-end: requirements, design, implementation, rollout, and post-launch improvements.
- Strong collaboration and communication skills; comfortable working with distributed teams and multiple stakeholders (SRE, Security, product engineering, leadership).
Required Skills
- Practical experience building or integrating LLM-based solutions (e.g., using OpenAI, Anthropic, Azure/OpenAI-compatible, or similar APIs).
- Hands-on experience with at least one of:
- AI agent frameworks (e.g., LangGraph-style, CrewAI-like, or custom agent patterns), or
- Orchestrating multi-step AI workflows using tools, function calling, or custom planners.
- Ability to translate infra problems into agentic workflows (e.g., “given alerts + context, generate and execute an actionable runbook; escalate with a summarized context if confidence is low”).
- Familiarity with prompt engineering, retrieval-augmented generation (RAG), and techniques to constrain and evaluate AI outputs in production.
Preferred Skills
- Experience in a large-scale SaaS or multi-tenant environment.
- Exposure to cost optimization / FinOps in AWS (right-sizing, storage optimization, data transfer analysis, automated cleanup).
- Experience with security
Click on Apply to know more.