zealant consulting group
Website:
zealantgroup.com
Job details:
Location: Bengaluru, India
Work Mode: Hybrid (3 Days/Week Work from Office)
Experience: 4–8 Years
About The Role
We are looking for a highly skilled
DevOps [Google Cloud] Engineer to build, scale, and manage cloud infrastructure, CI/CD pipelines, and platform reliability for a high-growth AI-driven product company. You will collaborate closely with software engineering, data engineering, and AI teams to improve deployment velocity, platform scalability, observability, security, and operational excellence.
This role is ideal for someone passionate about cloud-native technologies, Kubernetes, automation, Infrastructure as Code (IaC), and modern DevOps practices.
Key Responsibilities
Cloud Infrastructure
- Design, build, and manage scalable cloud infrastructure on Google Cloud Platform (GCP).
- Manage multiple environments including Development, Staging, Production, Preview, and Playground.
- Develop reusable Infrastructure-as-Code modules using Terraform.
- Ensure high availability, disaster recovery, backups, scalability, and secure infrastructure.
- Own platform operations and act as the primary escalation point for infrastructure issues.
Kubernetes & Platform Operations
- Manage production Kubernetes clusters (GKE) and containerized microservices.
- Support databases, background jobs, queues, and distributed systems from an infrastructure perspective.
- Partner with engineering teams to improve performance, scalability, resiliency, and production debugging.
CI/CD & Release Engineering
- Design, maintain, and optimize CI/CD pipelines using GitLab CI and ArgoCD.
- Implement automated deployments, rollback strategies, blue-green deployments, and canary releases.
- Build preview environments and self-service deployment workflows.
- Improve developer productivity through automation and standardized infrastructure.
Observability & Reliability
- Implement end-to-end monitoring using OpenTelemetry.
- Build dashboards, alerts, and monitoring solutions using tools like Prometheus, Grafana, Signoz, ELK, and Loki.
- Define and maintain SLIs, SLOs, SLAs, and error budgets.
- Lead incident management, root cause analysis, and post-incident reviews.
- Drive reliability engineering and production readiness across services.
Security & Compliance
- Implement security best practices across cloud infrastructure and Kubernetes.
- Manage IAM, RBAC, secrets management, encryption, and network security.
- Integrate security scanning and compliance checks into CI/CD pipelines.
- Ensure secure handling of sensitive application and customer data.
- Support compliance initiatives including SOC2, HIPAA, or similar security frameworks.
MLOps & AI Infrastructure
- Support deployment and scaling of AI/LLM applications.
- Manage GPU-backed workloads, autoscaling, and cost optimization.
- Standardize model deployment, versioning, rollout, and monitoring strategies.
- Optimize infrastructure costs through FinOps best practices.
Required Skills - 3–5 years of experience in DevOps, Platform Engineering, or Site Reliability Engineering (SRE).
- Strong hands-on experience with Google Cloud Platform (GCP).
- Expertise in Docker and Kubernetes (GKE).
- Experience building and maintaining CI/CD pipelines using GitLab CI and ArgoCD.
- Strong knowledge of Terraform for Infrastructure as Code.
- Experience with Istio Service Mesh.
- Hands-on experience with monitoring and observability tools including:
- OpenTelemetry
- Prometheus
- Grafana
- ELK
- Loki
- Signoz
- Strong Linux administration skills.
- Scripting experience using Bash and Python.
- Experience managing production systems with high availability requirements.
- Excellent troubleshooting, debugging, and incident management skills.
Good to Have
- Experience with MLOps or LLMOps.
- Experience managing GPU workloads.
- Knowledge of AI infrastructure and model serving.
- Experience with workflow engines or data platforms.
- Experience working in regulated industries such as Healthcare or FinTech.
- Knowledge of FinOps and cloud cost optimization.
Skills: google cloud platform (gcp),infrastructure as code (iac),gke,gcp,kubernetes (gke),terraform,python
Click on Apply to know more.