Concentric AI
Website:
concentric.ai
Job details:
Location: Bengaluru
Internal Team: OnPrem
Years of experience: 2-5
About the Role
We are seeking an SRE to manage release operations and infrastructure reliability across 100+ Kubernetes clusters. You will drive automated deployments, build observability, troubleshoot complex network/system issues, and create production runbooks to minimize downtime. In addition to core reliability work, you will contribute directly to internal coding projects, build custom tooling, and assist customers during escalations.
Key Responsibilities
- Orchestrate progressive rollouts, upgrades, and release automation across a fleet of 100+ Kubernetes clusters.
- Build and tune monitoring, logging, and alerting systems across all environments.
- Lead root-cause analysis for complex outages and write actionable runbooks to automate recovery.
- Write scripts and tools to eliminate toil and automate release pipelines.
- Join customer-facing calls to directly assist in troubleshooting, diagnosing, and resolving complex technical and environment issues.
Requirements
Must-Have
- Kubernetes & Docker: Production experience deploying, managing, and debugging containerized workloads at multi-cluster scale.
- Release Engineering: Proven experience deploying software across large, distributed K8s environments.
- Monitoring & Debugging: Hands-on experience with observability tools (like Prometheus, Grafana, Datadog, ELK) and distributed troubleshooting.
- Networking Fundamentals: Solid grasp of SSH tunneling, HAProxy, HTTP/S, DNS, and REST APIs.
- Scripting/Coding: Proficiency in Bash / Shell Scripting for debugging system issues.
- Runbook Documentation: Track record of writing clear operational playbooks.
Nice-to-Have
- Rancher for fleet-scale Kubernetes management.
- Enterprise auth and storage: LDAP, Kerberos, and Windows SMB.
- Proficiency in Python / Java / UI.
Click on Apply to know more.