Job Overview
We are looking for a Senior Site Reliability Engineer to manage and improve the reliability, scalability, availability, and performance of production systems across a hybrid/multi-cloud environment.
Key Responsibilities
Own SLIs, SLOs, error budgets, capacity planning, and reliability improvements.
Manage production Kubernetes / Amazon EKS and Red Hat OpenShift environments.
Work across AWS and IBM Cloud, including hybrid connectivity, networking, and disaster recovery.
Build and maintain infrastructure using Terraform and automate operations using Python and Bash.
Design and improve observability using Prometheus, Grafana, OpenTelemetry, Thanos, and logging platforms.
Lead high-severity incident response, postmortems, and corrective actions.
Implement secure and compliant infrastructure practices across cloud environments.
Mentor engineers and contribute to architecture and reliability standards.
Must-Have Skills
4–6 years of experience in SRE / DevOps / Cloud Infrastructure.
Strong hands-on experience with AWS and Kubernetes.
Experience with EKS, Docker, Helm; OpenShift is highly preferred.
Strong Terraform, Python, and Bash skills.
Hands-on experience with Prometheus and Grafana.
Experience with SLIs, SLOs, incident management, and production troubleshooting.
Good understanding of DNS, networking, TLS, VPN, load balancing, and cloud connectivity.
Exposure to regulated environments such as HIPAA, SOC 2, PCI DSS, ISO 27001, etc.