Mumba Technologies, Inc.
Website:
mumbatech.com
Job details:
Senior Site Reliability Engineer (SRE) – AWS, Kubernetes & Observability
We’re looking for an experienced Senior SRE to drive infrastructure reliability, observability, automation, and AI-powered operations across cloud-native environments. The role will focus on incident management, RCA, AWS, Kubernetes, security, and performance optimization, collaborating with engineering, product, and security teams.
What you’ll do
- Lead incident response, conduct RCA, and ensure remediation actions are completed.
- Build runbooks, playbooks, and escalation frameworks while reducing operational toil.
- Design and manage observability using Prometheus, Grafana, OpenTelemetry, Datadog, Dynatrace, or similar tools.
- Implement intelligent alerting, distributed tracing, dependency mapping, profiling, and RUM.
- Leverage AIOps and AI/LLM tools for anomaly detection, incident triage, predictive alerting, and MTTR reduction.
- Manage AWS infrastructure using Terraform/Pulumi, including cost optimization and capacity planning.
- Own Kubernetes operations, including autoscaling, networking, resource management, and upgrades.
- Drive infrastructure security, vulnerability remediation, and CI/CD integration of SAST, DAST, and SCA.
- Mentor SRE engineers and collaborate on architecture reviews, engineering standards, and on-call improvements.
What we’re looking for
- Strong experience in SRE, DevOps, Cloud Infrastructure, or Platform Engineering.
- Hands-on experience with incident management, RCA, observability, and automation.
- Strong AWS, Kubernetes, and Infrastructure as Code (Terraform/Pulumi) experience.
- Experience with AIOps, AI-assisted operations, or predictive analytics.
- Knowledge of security practices, vulnerability management, and CI/CD pipelines.
- Strong troubleshooting, communication, and technical leadership skills.
- Experience mentoring engineers and supporting production environments.
The Competitive Edge
AEM Administration
- Experience managing AEM Author, Publish, Dispatcher, and AEMaaCS environments.
- Knowledge of AEM health monitoring, Dispatcher optimization, DAM, OSGi configurations, and replication queues.
Exposure to CDN
- Experience with Cloudflare CDN and Workers, including caching, performance optimization, and edge computing.
Click on Apply to know more.