Website:
synthlane.com
Job details:
Company Description
Synthlane is a leading IT services and consulting firm that helps businesses, large enterprises, and government institutions solve complex technology challenges with strategic, scalable solutions. The company specializes in enterprise IT consulting, digital transformation, cybersecurity and risk management, cloud and infrastructure services, and custom software development. Synthlane focuses on aligning technology initiatives with business objectives, enabling clients to modernize operations, strengthen security, and improve productivity. Its experienced team of consultants, engineers, and strategists works closely with clients to design future-ready solutions that deliver measurable results. Organizations seeking a trusted partner for reliable, innovation-driven IT services rely on Synthlane to support long-term success.
Role Description
We are seeking an accomplished Senior Site Reliability Engineer (SRE) to lead the
design, implementation, and evolution of highly available, scalable, and resilient systems
across our multi-cloud infrastructure. In this senior role, you will drive architectural
decisions, establish reliability standards, and mentor teams while ensuring operational
excellence across complex distributed systems. You will partner with engineering
leadership, development teams, and product stakeholders to shape infrastructure
strategy, implement sophisticated automation, and champion a culture of reliability
engineering.
As a Senior SRE, you'll tackle sophisticated, large-scale challenges using cutting-edge
technologies across AWS and Azure platforms. You will lead critical initiatives that
impact system reliability at scale, architect solutions for complex infrastructure problems,
and guide teams in adopting industry-leading practices that drive meaningful
improvements across our entire technology ecosystem.
Requirements
● Architect and implement highly reliable, scalable, and cost-effective infrastructure
solutions for mission-critical applications across multi-cloud environments (AWS and
Azure).
● Lead the definition and refinement of service level objectives (SLOs), service level
indicators (SLIs), and error budgets, establishing reliability standards across the
organization.
● Design and implement sophisticated Infrastructure as Code (IaC) solutions using
Terraform, Ansible, and Azure Resource Manager (ARM) templates or Bicep.
● Drive automation strategies to eliminate toil, improve operational efficiency, and enable
self-service capabilities for development teams.
● Lead incident response efforts, conduct thorough post-incident reviews, and implement
systemic improvements to prevent recurrence.
● Champion cloud-native architectures and modern reliability practices, serving as a
technical advisor for infrastructure and platform decisions.
● Participate in and help optimize the on-call rotation, ensuring sustainable practices and
effective escalation procedures.
● Establish and maintain comprehensive documentation standards, runbooks, and
knowledge repositories that enable team autonomy and effective incident response.
● Design and implement advanced monitoring, logging, and alerting strategies using
observability platforms to enable proactive issue detection and resolution.
● Lead container orchestration initiatives using Kubernetes (AKS, EKS) and implement
sophisticated deployment strategies including blue-green, canary, and progressive
delivery patterns.
● Ensure security, compliance, and governance standards are embedded throughout the
infrastructure lifecycle, implementing security-as-code practices.
● Drive capacity planning, performance optimization, and cost management initiatives
across cloud platforms.
● Architect and implement highly reliable, scalable, and cost-effective infrastructure
solutions for mission-critical applications across multi-cloud environments (AWS and
Azure).
● Lead the definition and refinement of service level objectives (SLOs), service level
indicators (SLIs), and error budgets, establishing reliability standards across the
organization.
● Design and implement sophisticated Infrastructure as Code (IaC) solutions using
Terraform, Ansible, and Azure Resource Manager (ARM) templates or Bicep.
● Drive automation strategies to eliminate toil, improve operational efficiency, and enable
self-service capabilities for development teams.
● Lead incident response efforts, conduct thorough post-incident reviews, and implement
systemic improvements to prevent recurrence.
● Champion cloud-native architectures and modern reliability practices, serving as a
technical advisor for infrastructure and platform decisions.
● Participate in and help optimize the on-call rotation, ensuring sustainable practices and
effective escalation procedures.
● Establish and maintain comprehensive documentation standards, runbooks, and
knowledge repositories that enable team autonomy and effective incident response.
● Design and implement advanced monitoring, logging, and alerting strategies using
observability platforms to enable proactive issue detection and resolution.
● Lead container orchestration initiatives using Kubernetes (AKS, EKS) and implement
sophisticated deployment strategies including blue-green, canary, and progressive
delivery patterns.
● Ensure security, compliance, and governance standards are embedded throughout the
infrastructure lifecycle, implementing security-as-code practices.
● Drive capacity planning, performance optimization, and cost management initiatives
across cloud platforms.
● Collaborate with architecture and security teams to establish platform standards,
reference architectures, and best practices.
Click on Apply to know more.