Website:
buildandhire.com
Job details:
SRE Observability Engineer
We are currently hiring for an
SRE Observability Engineer for a fast-growing software company in Pune, India
Key Responsibilities
- Design and implement end-to-end observability solutions using modern monitoring and logging platforms.
- Define and monitor Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Key Performance Indicators (KPIs) to ensure system reliability.
- Perform Root Cause Analysis (RCA), trend analysis, and support Problem Management initiatives.
- Build dashboards, alerting rules, capacity reports, and anomaly detection mechanisms.
- Analyze logs, metrics, and distributed traces to identify performance bottlenecks and improve system reliability.
- Identify monitoring gaps and enhance observability coverage across cloud-native applications.
- Support load and performance testing activities.
- Participate in Major Incident Management and ensure timely service recovery.
- Collaborate with DevOps and Engineering teams to improve system resilience, scalability, and operational excellence.
- Develop and validate recovery plans while driving continuous reliability improvements.
Required Skills / Primary Skills
- Strong experience with Dynatrace, Grafana, Kibana, and the ELK Stack.
- Hands-on experience with Splunk and OpenTelemetry.
- Experience with cloud monitoring on AWS and/or Azure.
- Strong understanding of distributed tracing and observability best practices.
- Experience monitoring Kubernetes-based applications and infrastructure.
- Knowledge of Site Reliability Engineering (SRE) principles.
- Experience with monitoring, alerting, incident management, and performance optimization.
- Strong analytical, troubleshooting, and Root Cause Analysis (RCA) skills.
Additional Skills
- SRE or DevOps certifications are preferred.
- Experience supporting high-volume, real-time distributed systems.
- Knowledge of cloud-native architectures and microservices.
- Excellent problem-solving and communication skills.
- Ability to collaborate effectively with cross-functional engineering and operations teams.
- Self-driven with a proactive approach to improving system reliability and performance.
Notice Period
Immediate Joiners Preferred
If you're passionate about building highly reliable, scalable, and observable cloud platforms, we'd love to hear from you. Apply now and become part of an exciting technology team.
Click on Apply to know more.