Insight Global
Website:
insightglobal.com
Company:
https://www.linkedin.com/company/insight-global
Seniority: Mid-Senior level
Industries: Business Consulting and Services
Job details:
Required Skills & Experience
• 5+ years in SRE, infrastructure, or platform engineering with hands-on production Kubernetes.
• Deep AWS and Azure (services, networking, IAM) and strong Kubernetes / EKS.
• Reliability fundamentals — SLIs/SLOs/error budgets, incident response, on-call.
• Observability tooling (Elastic plus Prometheus/Grafana/Datadog).
• Terraform multi-cloud IaC; scripting in Python, Bash, and YAML.
• Cloud migration — moving Azure-native services (Functions, Cosmos DB) toward AWS / cloud-agnostic.
• Policy and access control — OPA/Gatekeeper, AWS IAM, Azure RBAC.
• Kubernetes troubleshooting — capacity planning, performance tuning, load/scalability testing.
Nice to Have Skills & Experience
• Managed Kubernetes (AKS, EKS, GKE) and service mesh
• Argo Rollouts / Flagger progressive delivery
• Cilium / eBPF networking
• Chaos engineering and hybrid-architecture DR patterns
• Kyverno and supply-chain security
• CKA (CKS a plus); Azure/AWS certs
Job Description
Insight Global is seeking a Kubernetes Site Reliability Engineer to keep Ecolab’s new container platform reliable, fast, and secure across Azure and AWS. Applying software engineering to operations, you’ll set service-level objectives, lead incident response and on-call, automate toil, and help migrate Azure-native services toward cloud-agnostic patterns - codifying everything as Infrastructure as Code, building safe delivery pipelines, and hardening to CIS Benchmarks. Success is measured in uptime, fast recovery, and toil removed.
Key Responsibilities:
• Reliability & SLOs: Define and maintain SLIs, SLOs, and error budgets, and use them to drive priorities.
• Incident & on-call: Lead incident response on-call — detect, triage, resolve — then run blameless post-mortems.
• Observability: Build and tune monitoring, logging, tracing, and alerting (Elastic, Prometheus, Grafana, Datadog).
• Cloud migration: Help migrate Azure-native services (Functions, Cosmos DB) toward AWS / cloud-agnostic patterns with failover, replication, and latency tuning.
• Automation & performance: Engineer toil away (remediation, scaling, patching, backups) and run capacity planning, performance tuning, and load/scalability testing.
• Safe delivery: Build and safeguard CI/CD and progressive delivery (Azure DevOps, GitHub Actions, ArgoCD, Argo Rollouts) with automated rollbacks.
• Security & policy: Harden to CIS Benchmarks and apply policy and access controls — OPA/Gatekeeper, AWS IAM, Azure RBAC — remediating Mythos vulnerabilities with security.
Click on Apply to know more.