Mumba Technologies, Inc.
Website:
mumbatech.com
Job details:
Key Responsibilities
Reliability & Availability
- Lead incident response, conduct RCAs and ensure action items are tracked to closure
- Build and maintain runbooks, playbooks and escalation frameworks for proactive and reactive response
- Drive toil reduction by identifying repetitive operational work and engineering it away
Observability
- Design and own the full observability stack — metrics, logs, traces and events — using tools like Prometheus, Grafana, OpenTelemetry, Datadog, Dynatrace or similar
- Build intelligent alerting that reduces noise, eliminates alert fatigue and surfaces actionable signals
- Implement distributed tracing and dependency mapping to provide end-to-end visibility across microservices
- Drive adoption of continuous profiling and real user monitoring (RUM) for proactive performance management
AI Adoption in SRE
- Leverage AIOps platforms to enable anomaly detection, predictive alerting and automated root cause analysis
- Implement AI-assisted incident triage — using LLM-powered tools to summarise incidents, suggest fixes and accelerate MTTR
- Build and maintain ML-powered capacity forecasting models to optimise infrastructure spend and prevent resource saturation
Security & Vulnerability Management
- Embed security-as-reliability principles — treating security incidents with the same urgency as availability incidents
- Own issue remediations across infrastructure (OS, containers, dependencies)
- Integrate SAST, DAST and SCA tools into CI/CD pipelines to shift security left
Infrastructure & Platform Engineering
- Design, build and maintain cloud-native infrastructure on AWS using Infrastructure as Code (Terraform, Pulumi) & drive rightsizing, reserved capacity planning and cost anomaly detection
- Own Kubernetes cluster operations — autoscaling, resource management, networking and upgrade strategy
Leadership & Culture
- Mentor and guide junior and mid-level SREs — conducting technical reviews and pair debugging sessions
- Define and evolve SRE team standards, best practices and engineering principles
- Collaborate closely with product, development and security teams as an embedded reliability partner
- Contribute to on-call rotation and drive continuous improvement of on-call experience
- Represent SRE in architecture reviews, sprint planning and cross-functional forums
The Competitive Edge
AEM Administration
- Own end-to-end reliability and availability of AEM environments — Author, Publish, Dispatcher and AEM as a Cloud Service (AEMaaCS) — across dev, staging and production
- Monitor and manage AEM instance health, optimise Dispatchers, Manage DAM, OSGi Configurations, replication queues.
Exposure to CDN
- Experience with Cloudflare - CDN, Workers,
Click on Apply to know more.