Job Summary
We are looking for a Senior Platform Engineer to design, implement, and operate cloud-native platforms on Microsoft Azure. The role focuses on Kubernetes, cloud infrastructure, automation, observability, security, and Disaster Recovery, ensuring highly available, secure, and resilient production environments.
Key Responsibilities
● Design, deploy, and manage Azure infrastructure and AKS clusters.
● Implement and maintain Disaster Recovery, backup, and business continuity solutions.
● Automate infrastructure and deployments using Terraform, Helm, and Azure DevOps.
● Manage Kubernetes networking, ingress, storage, and security.
● Deploy and maintain observability platforms including Prometheus, Grafana, and Loki.
● Manage TLS certificates, secrets, and platform security.
● Support production environments, troubleshoot critical incidents, and drive root cause analysis.
● Plan and execute platform migrations and infrastructure upgrades.
● Create and maintain technical documentation, architecture diagrams, HLD/LLD, SOPs, operational runbooks, troubleshooting guides, and Disaster Recovery procedures.
● Define and maintain platform engineering standards, policies, governance frameworks, and best practices across Azure and Kubernetes environments.
● Establish cloud and Kubernetes governance covering RBAC, naming and tagging standards, resource organisation, security controls, namespaces, resource limits, ingress, secrets, and storage.
● Define Infrastructure-as-Code and CI/CD governance, including reusable Terraform modules, code review
standards, state management, pipeline controls, approval gates, environment promotion, and artifact/version management.
● Participate in architecture and technical design reviews for new platforms, applications, integrations, and infrastructure changes.
● Drive platform security and compliance readiness through security baselines, vulnerability remediation, access reviews, secrets/certificate management, audit controls, and policy enforcement.
● Define and track platform availability, SLIs/SLOs, capacity, performance, and operational health.
● Drive incident and problem management practices, including root cause analysis, corrective actions, and prevention of recurring incidents.
● Perform capacity planning, performance optimisation, and cloud cost optimization across platform infrastructure.
● Own DR testing, RTO/RPO validation, backup/restore standards, recovery procedures, and evidence from periodic recovery exercises.
● Evaluate new platform technologies, conduct POCs, and establish approved patterns before production adoption.
● Provide technical leadership, knowledge sharing, and mentoring to engineers on Azure, Kubernetes, Terraform, CI/CD, security, and platform operations.
Required Skills
● Microsoft Azure
● Kubernetes (AKS/OpenShift)
● Docker & Helm
● Terraform
● Azure DevOps / CI/CD
● Prometheus, Grafana, Loki
● Azure Networking (VNets, NSGs, Private Endpoints, Firewall)
● Linux & Bash scripting
● Disaster Recovery, Backup & Restore strategies
● Experience with PostgreSQL, MongoDB, MySQL, or Azure SQL
● Cloud & Platform Governance
● Kubernetes Security, Governance & Policy Enforcement
● Azure Policy, RBAC & Security Controls
● Infrastructure-as-Code standards and reusable Terraform patterns
● CI/CD governance, release controls, and environment promotion strategies
● Technical documentation (HLD, LLD, SOPs, runbooks, and architecture diagrams)
● Architecture design and technical design reviews
● Incident, Problem & Root Cause Analysis management
● Capacity planning, performance optimization & FinOps / cloud cost optimization
● Security, compliance, audit controls & operational governance
● RTO/RPO planning, DR testing, backup and recovery governance
● Git / GitOps practices and source control standards
● Technical leadership, mentoring, and cross-functional collaboration
Preferred
● Banking or regulated industry experience
● Strong understanding of High Availability, Disaster Recovery, and production operations
● Experience defining enterprise platform standards, policies, and governance frameworks.
● Experience working in security- and compliance-controlled environments.
● Experience leading technical design reviews, platform modernisation, and cloud transformation initiatives