Responsibilities
Lead migration of services to Kubernetes and manage scalable cluster architecture
Maintain infrastructure using Terraform across Azure and GCP environments
Build and improve CI/CD pipelines for reliable and fast deployments
Ensure consistency across development, staging, and production environments
Manage monitoring and observability using Grafana and Prometheus
Track system performance, including p95/p99 latency, and improve reliability
Handle production incidents, drive resolution, and create preventive runbooks
Collaborate with engineering teams to improve deployment and operational efficiency