Mumba Technologies, Inc.
Website:
mumbatech.com
Company:
https://www.linkedin.com/company/mumba-technologies
Industries: Software Development
Job details:
Job Title: Site Reliability Engineer (SRE)
Job Type: Full Time
Location: Gurgaon (Hybrid)
Job Summary
We are looking for a Senior Site Reliability Engineer (SRE) with 7–10 years of experience to drive reliability, observability, automation, and cloud-native platform engineering. The ideal candidate will have strong hands-on experience with AWS, Kubernetes, Terraform, monitoring/observability, incident management, and AEM environments.
Key Responsibilities
Reliability & Observability
- Lead production incidents, RCA, problem management, and reliability improvements.
- Build and maintain observability using Prometheus, Grafana, Dynatrace, Datadog, OpenTelemetry or similar tools.
- Develop actionable alerting, dashboards, distributed tracing, and performance monitoring.
- Drive automation and toil reduction across production operations.
Cloud & Platform Engineering
- Design and manage highly available AWS infrastructure using Terraform/Pulumi.
- Manage Kubernetes clusters, including upgrades, autoscaling, networking, and resource optimization.
- Drive cloud cost optimization, capacity planning, and infrastructure reliability.
AEM Administration
- Manage reliability and availability of Adobe Experience Manager (AEM) environments across Author, Publish, Dispatcher, and AEM as a Cloud Service.
- Troubleshoot AEM performance, replication queues, OSGi configurations, DAM, and Dispatcher issues.
- Monitor AEM application and infrastructure health across Dev, QA, and Production.
AI & Automation
- Apply AIOps/AI-assisted tools for anomaly detection, incident triage, RCA, and operational automation.
- Leverage LLM-based solutions to improve troubleshooting and reduce MTTR.
Security
- Support vulnerability remediation across OS, containers, dependencies, and cloud infrastructure.
- Integrate SAST, DAST, and SCA practices into CI/CD pipelines.
Leadership & Collaboration
- Mentor junior and mid-level engineers and establish SRE best practices.
- Participate in architecture reviews, on-call rotations, and cross-functional engineering initiatives.
- Partner with development, security, and product teams to improve overall platform reliability.
Required Skills
- 7–10 years of experience in SRE, DevOps, Cloud Engineering, or Platform Engineering.
- Strong AWS and Kubernetes experience.
- Hands-on Terraform/Pulumi experience.
- Strong knowledge of observability, monitoring, alerting, SLO/SLI, and incident management.
- Experience with AEM Administration, including Author/Publish/Dispatcher.
- Experience with CDN technologies, preferably Cloudflare.
- Strong Linux, scripting, troubleshooting, and automation skills.
- Experience with CI/CD and DevSecOps practices.
- Strong communication, problem-solving, and stakeholder-management skills.
Click on Apply to know more.