SkillsCapital
Website:
skillscapital.io
Job details:
We are hiring Senior Site Reliability Engineers to join a leading global technology organization building and operating a growing portfolio of enterprise SaaS platforms and cloud-native products.
This is an opportunity to work on large-scale production infrastructure while collaborating with software engineers, platform architects, cloud teams, and security specialists to build highly reliable, scalable, and automated platforms supporting millions to billions of real-time transactions across multiple enterprise applications.
Location - 100% Remote
Start Date - Immediate to 4 Weeks
What You'll Do
Own the reliability, observability, and operational excellence of multiple enterprise SaaS platforms.
Build and maintain production-grade observability platforms using Grafana, Prometheus, and Loki.
Define and implement Service Level Objectives (SLOs), Service Level Indicators (SLIs), and Error Budgets across products.
Design and improve incident response processes, on-call practices, escalation procedures, and post-incident reviews.
Build intelligent alerting systems that reduce noise and provide actionable diagnostic context.
Automate operational tasks including safe remediation, infrastructure provisioning, deployment workflows, and platform operations.
Design, build, and maintain CI/CD pipelines using GitHub Actions.
Manage Infrastructure as Code using Terraform/OpenTofu across multi-account AWS environments.
Design and operate large-scale event streaming infrastructure supporting real-time and batch processing workloads.
Improve platform reliability, scalability, resilience, and operational efficiency.
Optimise cloud infrastructure performance and cost across AWS services.
Collaborate with engineering teams to improve platform standards, deployment reliability, and developer experience.
Support platform integration and modernization initiatives while maintaining production stability.
Document operational procedures, platform architecture, and engineering standards.
Required Skills
8–12 years of experience in Site Reliability Engineering (SRE), Platform Engineering, DevOps, or Cloud Infrastructure Engineering.
Strong hands-on experience with AWS including ECS, EKS, IAM, VPC, RDS, CloudWatch, and Cost Explorer.
Deep expertise with Terraform/OpenTofu across multi-account and multi-environment AWS deployments.
Strong experience managing enterprise CI/CD pipelines using GitHub Actions.
Experience operating production-scale event streaming infrastructure, including reliability, operations, and performance optimisation.
Hands-on experience with Grafana, Prometheus, and Loki in production environments.
Strong understanding of SRE principles, including SLOs, Error Budgets, incident management, and operational excellence.
Experience improving reliability for large-scale distributed systems.
Excellent troubleshooting, root cause analysis, and production support skills.
Strong communication and collaboration skills within distributed engineering teams.
Preferred Skills
Experience building internal platform engineering capabilities.
Experience integrating or modernising enterprise platforms.
Strong understanding of automation and AI-assisted operational workflows.
Experience implementing automated remediation for known production failure scenarios.
Knowledge of cloud cost optimisation strategies.
Experience with secrets management solutions such as AWS Secrets Manager or HashiCorp Vault.
Familiarity with Agile software delivery and modern engineering practices.
AWS certifications are an advantage.
Why Consider This Opportunity
Opportunity to build and operate enterprise-scale production platforms supporting multiple SaaS products.
Work on modern Site Reliability Engineering practices across cloud-native environments.
Exposure to AWS, Terraform, GitHub Actions, Observability, Event Streaming, Infrastructure as Code, and Platform Engineering.
Collaborate with highly experienced engineering and platform teams.
Own production reliability, scalability, and operational excellence across mission-critical enterprise systems.
Fully remote work environment.
Long-term project opportunities with extension potential.
Fast interview process and onboarding.
If you are passionate about Site Reliability Engineering, cloud infrastructure, observability, automation, and building highly reliable distributed systems, we would love to hear from you.
Click on Apply to know more.