Sheshi AI
Website:
sheshi.ai
Job details:
SITE RELIABILITY ENGINEER
ABOUT SHESHI
Bringing Sanctity to Numbers
Sheshi is a Financial Operating System — the governed infrastructure layer that converts raw financial data into results that organisations can completely trust. Governed, auditable, period-locked, and traceable to the source.
We are headquartered in Bangalore, building an enterprise-grade product for the global market.
Sheshi runs financial infrastructure that enterprises depend on. Uptime, performance, and reliability are not operational nice-to-haves here — they are part of what makes the platform trustworthy. As Site Reliability Engineer, you own that reliability end to end on the cloud environments Sheshi runs on today, building the systems and practices that let the engineering team ship with confidence.
WHY SHESHI
What Makes This Role Worth Doing
Reliability with real consequence. The platform you keep running produces numbers of enterprises build financial decisions on.
Full ownership from day one. You build the monitoring, incident response, and infrastructure practices from the ground up.
On-call built the right way. You shape what on call looks like at Sheshi from the start, rather than working within a process someone else built.
Room to grow with the platform. As Sheshi's infrastructure grows toward hybrid and on-premises deployment, your role and depth grow with it.
RESPONSIBILITIES
What You Will Own
Team Leadership and Management
- Build and manage a growing team of Site Reliability Engineers, overseeing their day-to-day operations, career development, and performance.
- Mentor and coach junior and mid-level engineers, fostering a culture of technical excellence, psychological safety, and continuous learning.
- Resource and capacity planning, aligning the SRE team's roadmap with broader engineering and business goals.
Production Reliability and Uptime
- Own uptime, performance, and availability of the AWS production environment.
- Build and maintain monitoring, alerting, and observability so issues surface before customers feel them.
- Define and track reliability metrics — SLOs, error budgets, and performance baselines.
Kubernetes Infrastructure & Production Reliability
- Own the uptime, availability, and performance of our production Kubernetes clusters running on AWS (EKS).
- Architect, scale, and maintain our containerized application infrastructure to ensure enterprise-grade resilience.
- Build and maintain Kubernetes-native monitoring, alerting, and observability so issues surface before customers feel them.
- Define, track, and report on reliability metrics — SLOs, error budgets, and cluster resource utilization.
Incident Response
- Build and own the incident response process — detection, escalation, resolution, and communication.
- Lead incident response from day one, including shaping on-call coverage as the team grows.
- Run blameless post-incident reviews and drive fixes that address root cause.
Infrastructure Optimisation and Automation
- Continuously assess and optimise AWS infrastructure for performance, scalability, and cost.
- Build and maintain infrastructure as code for provisioning, scaling, and configuration.
- Automate repetitive operational work so the team's time goes toward what matters.
Security and Future Readiness
- Partner with the DevSecOps function to ensure infrastructure meets security and compliance standards.
- Design cloud infrastructure and practices that can extend cleanly to hybrid or on-premises environments when the business requires it.
WHAT WE ARE LOOKING FOR
Your Background and Strengths
Experience
- 4-8 years in Site Reliability Engineering, DevOps, or infrastructure engineering roles. 2+ years of direct people management or formal team leadership experience
- Proven experience owning production reliability for a SaaS platform at meaningful scale.
- Track record of building monitoring, alerting, and incident response practices from the ground up.
Technical Depth
- Strong hands-on AWS expertise — EC2, ECS/EKS, RDS, S3, CloudWatch, IAM, VPC.
- Experience with infrastructure as code — Terraform, CloudFormation, or equivalent.
- Strong command of observability tooling — metrics, logging, tracing, and alerting stacks.
- Comfort with containerisation and orchestration — Docker, Kubernetes, or equivalent.
- Scripting proficiency in Python, Bash, or similar for automation.
How You Work
- You take ownership of production health as though it were your own product.
- You build systems that prevent problems, not just respond to them.
- You communicate incidents and risks clearly and early, without minimising or overstating.
- You are comfortable being the first call when something breaks, and calm under that pressure.
Nice to Have
- Exposure to hybrid or on-premises infrastructure, including virtualisation platforms such as Proxmox.
- Experience in a regulated, financial, or compliance-sensitive SaaS environment.
- Familiarity with Node.js or React application performance characteristics from an infrastructure perspective.
HOW TO APPLY
Write to us at hr.ind@sheshi.ai
Subject line: Site Reliability Engineer — Your Name
Skip the standard cover letter. Tell us — in under 300 words — about a production incident you owned end to end. What broke, how did you respond, and what did you change afterward so it could not happen the same way again.
- sheshi.ai · hr.ind@sheshi.ai · Bangalore, India
Click on Apply to know more.