WaferWire Cloud Technologies
Website:
waferwire.com
Job details:
SRE Manager – Reliability Engineering, Automation & Operational Excellence
Role Overview
Step into a technically hands-on SRE leadership role where you will engineer reliability – not just manage it. Own architecture for monitoring and automation, review infrastructure-as-code with rigor, author operational designs, and personally dive into complex incidents.
Bring an engineering-first mindset: be the leader who sees beyond what is currently working, identifies risks before they become incidents, champions AIOps, and actively expands the team into adjacent platform engineering.
You will lead a 15+ member team supporting large-scale cloud platform reliability across multiple workstreams – driving operational excellence, cloud cost efficiency, and engineering innovation. This is a high-visibility role as the single point of accountability between the vendor team and client engineering leads.
Key Responsibilities
Architecture & Technical Leadership
- Own reliability architecture: drive decisions for monitoring, alerting, automation, and infrastructure tooling; conduct code reviews on IaC, scripts, and configs
- Personally engage in complex incident debugging using KQL, telemetry, and distributed traces to unblock the team and drive rapid root cause resolution
- Author runbook designs, define operational architecture, standardize procedures, and continuously raise the reliability bar
Delivery & Execution
- Drive delivery in Agile/Scrum against contractual commitments: sprint progress, capacity management, delivery forecasting
- Manage on-call rotation with health tracking (burnout, incident volume, team well-being)
- Drive incident management excellence: SLA targets, post-incident reviews, MTTD/MTTR tracking, corrective actions, systemic improvements
- Drive deployment operations: staged rollouts, Blue/Green deployments, deployment validation gates, safe deployment practices across cluster-based infrastructure
Cloud Cost & Operational Efficiency
- Own cloud cost management: Azure spend monitoring, optimization, efficiency practices, burn rate reporting against budget
- Drive AI and automation: AIOps adoption, automated incident detection, intelligent alerting, auto-remediation, Copilot-assisted engineering
Innovation & Growth
- Think beyond current operations: identify reliability risks before incidents, propose innovative automation, expand into adjacent platform engineering
- Solution for new engagements: scoping, estimation, technical proposals, and delivery model shaping for new workstreams
People & Stakeholders
- Drive multi-stakeholder coordination: primary technical interface with client leads, priority translation, platform roadmap awareness, proactive capacity positioning
- Take full ownership: end-to-end accountability for team outcomes, proactive issue resolution, hold team to SLA and operational standards
- Drive recruitment, professional development, and culture: technical interviews, team connects, mentor/coach across levels (intern to principal)
Problem Solvin
- gDeep infrastructure debugging across distributed systems and cloud service degradatio
- nSees beyond current operations to identify risks and propose innovative automatio
- nSystems thinking for reliability architecture trade-off
- sData-driven incident analysis for staffing, process, and tooling decision
- sFinancial awareness for cloud spend management and team value demonstratio
n
Technic
- alSRE/DevOps leadership, operational architecture, distributed systems (Service Fabric, Kubernetes/AK
- S)Incident management, capacity planning, cloud cost optimizati
- onAzure DevOps, Agile/Scrum, infrastructure-as-code (ARM, Bicep, Terrafor
- m)Monitoring/observability (KQL, telemetry, distributed tracing), AIOps, automation tooli
- ngOn-call rotation design and health monitori
- ngDelivery metrics (MTTD, MTTR, deployment frequency, SLA complianc
e)
Required Skills & Qualificati
- onsHands-on SRE or platform engineering management leading 15+ engineers across geograph
- iesProven reliability architecture ownership, code review rigor, and complex infrastructure debugg
- ingTrack record of proactive improvements, solutioning for new engagements, and delivery growth beyond sc
- opeExperience managing contract-based delivery with SLA accountability and operational commitme
- ntsStrong design reviews, incident coordination, executive reporting, and stakeholder communicat
- ionOwnership mindset delivering operational excellence beyond expectati
ons
Preferred Qualificat
- ionsImproving MTTD/MTTR through technical and process cha
- ngesCloud cost optimization and large-scale SaaS operat
- ionsAIOps adoption (auto-remediation, intelligent alerting) in SRE t
- eamsOn-call health management and sustainable support mo
- delsSolutioning and estimation for new reliability engineering engagem
- entsOperational architecture documents, runbook designs, reliability stand
- ardsService Fabric, Kubernetes, or large-scale cluster-based platform operat
ions
Why You'll Love This
- RoleEngineer reliability at scale – own architecture decisions that directly impact platform uptime and perfor
- manceChampion AIOps, automation, and innovation while leading a high-performing team across multiple workst
- reamsShape operational excellence as the trusted engineering partner to client leadership with direct business i
mpact
Click on Apply to know more.