EPAM Systems
Website:
epam.com
Job details:
We are in search of an experienced
Lead/Senior SRE Engineer with strong Dynatrace expertise to join our team.
In this role, you will drive the implementation of DevOps and SRE practices, shape the technology roadmap, and ensure the reliability, performance, and observability of our production systems. You will collaborate closely with application teams and product owners to foster a culture of operational excellence and continuous improvement.
Responsibilities
- Deploy and standardize DevOps & SRE practices to improve how teams run production
- Steer conversations on the SRE technology roadmap and align stakeholders on priorities
- Establish and maintain SLIs and SLOs, plus MTTR, Lead time for change, Deployment Frequency, and Change Failure Rate measurements
- Implement and manage monitoring, alerting, operability, and observability for applications using Dynatrace, Splunk, and Grafana
- Evaluate performance through assessments and monitoring, and propose targeted performance enhancements
- Hold application teams accountable for performance and availability SLAs
- Work with product owners to manage error budget, prioritize toil backlog, and validate progress against team, application, and incident metrics
- Join an on-call rotation to handle production incidents and outages
- Improve the continuous integration & continuous deployment (CI/CD Pipeline) through ongoing optimization
- Use structured troubleshooting, incident management, and root cause analysis to resolve issues
- Advocate for automation and build automated processes wherever possible
- Execute cybersecurity measures via continuous vulnerability assessment and risk management
- Produce periodic reporting on progress for management and the customer
- Coordinate with application teams to simplify platform adoption and manage communication within the team and with customers
- Review the current system and create plans for enhancements and improvements
Requirements
- Bachelor's degree in Computer Science or a related discipline, or comparable hands-on experience
- 5+ years of overall IT experience, with 5+ years working in DevOps or SRE teams
- Practical experience supporting and operating production infrastructure
- Strong understanding of CI/CD and delivery workflows
- Clear knowledge of observability foundations: monitoring, logging, and tracing
- Advanced expertise with Dynatrace and Splunk
- Working knowledge of a leading cloud provider, including AWS, Azure, or GCP
- Demonstrated ability to run high-availability, fault-tolerant, scalable, distributed software in production environments
- Ability to work autonomously and collaboratively, with strong organizational and interpersonal skills and experience growing operational maturity
- Excellent analytical and problem-solving skills with strategic thinking and calm troubleshooting under pressure
- Flexibility to quickly adapt to new technologies and approaches
- English level B2 (Upper-Intermediate) or higher
Click on Apply to know more.