Role Summary
We are seeking a highly experienced SRE & DevOps Architect to lead and scale our DevOps & Reliability Center of Excellence (CoE). This role will define enterprise-wide DevOps, SRE, and platform engineering standards, drive reliability at scale, and partner with Engineering, Cloud, Security, and Business teams to enable high-performing, resilient, and cost-efficient platforms.
The ideal candidate brings deep technical expertise, strong architectural thinking, and proven CoE leadership across tools, processes, governance, and enablement.
Experience
15–20 years (with 6+ years in architecture & CoE leadership roles)
Key Responsibilities
🔹 Architecture & Strategy
- Define enterprise SRE & DevOps architecture, reference frameworks, and best practices.
- Design scalable, highly available, fault-tolerant platforms across cloud-native and hybrid environments.
- Establish reliability engineering principles including SLOs, SLIs, error budgets, and capacity planning.
- Lead adoption of platform engineering and Internal Developer Platforms (IDP).
🔹 SRE & DevOps CoE Leadership
- Build and operate the SRE/DevOps Center of Excellence.
- Define CoE operating model, governance, maturity models, and success metrics.
- Standardize CI/CD pipelines, IaC, observability, security, and release practices.
- Act as an internal consultant for product and engineering teams.
🔹 Cloud & Infrastructure
- Architect and govern solutions across AWS / Azure / GCP.
- Drive Infrastructure as Code using Terraform, CloudFormation, ARM, or equivalent.
- Enable container platforms using Kubernetes, OpenShift, and service mesh technologies.
- Establish cloud cost optimization and FinOps practices.
🔹 Observability & Reliability
- Architect enterprise observability using Prometheus, Grafana, ELK, Splunk, OpenTelemetry, Datadog, etc.
- Drive proactive monitoring, alerting, incident response, and root-cause analysis.
- Lead blameless postmortems and continuous reliability improvement initiatives.
🔹 CI/CD & Automation
- Define and standardize CI/CD platforms using tools like GitHub Actions, GitLab, Jenkins, Azure DevOps, Argo CD.
- Champion GitOps, pipeline-as-code, and automated quality gates.
- Integrate security (DevSecOps) into pipelines.
🔹 Security & Compliance
- Embed DevSecOps practices across SDLC.
- Work with security teams to enable secrets management, vulnerability scanning, and policy-as-code.
- Ensure compliance with enterprise and regulatory requirements.
🔹 Stakeholder & Leadership Engagement
- Partner with Engineering Heads, Cloud CoE, Security, and Enterprise Architecture teams.
- Mentor DevOps and SRE engineers; build upskilling programs and communities of practice.
- Influence leadership through metrics, ROI, and business outcomes.
Required Skills & Experience
- Must have strong experience in SRE (SLI, SLO, Error Budgets & Incident management), DevOps, Cloud Architecture, and Platform Engineering
- Containers & Orchestration: Docker, Kubernetes, Helm
- CI/CD: Jenkins, GitHub Actions, GitLab, Azure DevOps, Argo CD
- Cloud Platforms: AWS / Azure / GCP
- IaC & Config Management: Terraform, Ansible, CloudFormation
- Observability: Prometheus, Grafana, ELK, Open Telemetry
- Scripting: Python, Shell, Go (preferred)
- Security & DevSecOps toolchains
Architecture & CoE Experience
- Proven experience building and leading DevOps/SRE CoE
- Defining standards, reusable assets, and reference implementations
- Large-scale enterprise transformation experience
- Multi-team and multi-geography delivery exposure
Soft Skills
- Strong stakeholder management and communication skills
- Influencing without authority
- Mentorship and thought leadership mindset
- Ability to balance governance with developer velocity
Certifications (Preferred)
- AWS / Azure / GCP Architect Professional
- CKA / CKAD / CKS
- SRE or DevOps certifications