Cittabase Solutions
Website:
cittabase.com
Job details:
GCP / SRE / DevOps Engineer is a senior technical leadership role responsible for the reliability, scalability, and automation of a cloud-native GCP ecosystem. You will operate complex distributed systems built on GKE and SpringBoot microservices, drive toil elimination through advanced automation, and bridge application development with platform stability — ensuring a resilient, secure, and fully pipeline-driven environment.
Key Responsibilities
- Leadership & Programme Management
- Act as the primary liaison between the client and Cittabase engineering teams, driving alignment on priorities, SLAs, and delivery commitments.
- Allocate and prioritize Jira bugs, incidents, and stories across support engineers; track team deliverables and enforce SLA adherence.
- Own and present weekly, monthly, and quarterly operational metrics reports to client stakeholders.
- Provide technical mentorship to the team; unblock engineers through hands-on debugging, architecture guidance, and escalation support.
- Incident Management & Site Reliability
- Serve as Incident Commander for high-priority outages — coordinate cross-functional response, lead blameless post-mortems, and drive systemic prevention measures.
- Drive incident management activities ensuring timely resolution; participate in on-call rotations via PagerDuty.
- Define, monitor, and enforce SLOs, SLIs, and error budgets across streaming and batch data platforms.
- Implement distributed tracing (OpenTelemetry) to diagnose latency and message loss; establish DORA metrics and Four Golden Signals (Latency, Traffic, Errors, Saturation) as operational baselines.
- Automate operational runbooks to reduce manual toil and mean time to recovery.
- Observability & Cost Optimisation
- Design and maintain Grafana and Cloud Monitoring dashboards tracking the Four Golden Signals and platform health across all workstreams.
- Monitor and optimise GCP cloud costs using billing insights and FinOps best practices.
- Platform Engineering & Infrastructure
- Design and manage production-grade GKE clusters ensuring high availability for SpringBoot microservices; implement and optimise cloud-native persistence using AlloyDB.
- Configure and maintain Apigee gateways for secure, low-latency API management.
- Manage the health and scaling of Kafka brokers and Pub/Sub topics to ensure zero message loss in streaming pipelines.
- Support operational health of large-scale data processing tooling: BigQuery, Dataflow, and Cloud Composer orchestration.
- Own end-to-end infrastructure lifecycle via Terraform — establish reusable module standards and state management for consistent environments across the GCP ecosystem.
- CI/CD & DevOps Automation
- Architect and maintain robust CI/CD pipelines using GitHub Actions; transition manual deployment processes into fully automated, gated workflows for Cloud Functions, Dataflow, and Cloud Composer.
Qualifications
- 7+ years of experience in SRE or DevOps, with at least 4 years of hands-on leadership within the Google Cloud Platform (GCP).
- Proven experience in a leadership capacity during critical outages (Incident Commander) and a strong background in post-mortem documentation and remediation.
- Proven track record of building unified dashboards in Grafana and managing complex alerting rotations in PagerDuty.
- Expert-level knowledge of Kubernetes (GKE), including service mesh, ingress controllers, and cluster security.
- Advanced mastery of Terraform, GitHub Actions, and Kafka infrastructure management.
- Deep familiarity with GCP’s data and serverless offerings, including AlloyDB, Dataflow, Cloud Functions, and BigQuery.
- GCP Professional Cloud DevOps Engineer or Professional Cloud Architect certification is highly desirable.
Click on Apply to know more.