Job Description
SN Required Information Details
1 Role Senior Observability Engineer
2 Required Technical Skill Set · Designing, deploying, and scaling solutions that provide deep insights into our complex systems, directly influencing system reliability and strategic technology decisions · Hands-on experience in at least one of OpenTelemetry, Grafana, ITRS Geneos, Google Cloud Observability · OpenShift or Kubernetes administration · Prometheus and PromQL · Helm charts · Grafana
Designing, deploying, and scaling solutions that provide deep insights into our complex systems, directly influencing system reliability and strategic technology decisions. · Core Observability Skills - Candidates must have hands-on experience in at least one of the following key areas: 1. OpenTelemetry: Implementing and managing the OpenTelemetry framework. 2. Grafana Enterprise Stack: Deep knowledge of Mimir, Loki, and Tempo. 3. ITRS Geneos: Advanced administration and scaling in an enterprise environment. 4. Google Cloud Observability: Expertise with Google's cloud-native monitoring suite. · Essential Platform & Tooling Skills 1. Container Orchestration: Proficient in OpenShift or Kubernetes administration. 2. Metrics & Querying: Strong experience with Prometheus and PromQL. 3. Deployment: Expertise in creating and managing Helm charts for application deployment. 4. Grafana: Skilled in dashboard creation, data source management, and alerting.
Good-to-Have · Scripting: Automation experience with Python or Bash. · Familiarity with enterprise deployment pipelines (e.g., Lightspeed).
Soft Skills Practical problem solving and strategic thinking skills Demonstrated leadership, interpersonal skills and relationship building skills
Service oriented attitude Ability to work in a fast-paced environment Experience working or leading requirement gathering efforts for multiple large development projects at one-time Proficient using basic technical tools and systems Good interpersonal and communication skills
SN Responsibility of / Expectations from the Role
1 Design, build, and manage end-to-end observability solutions (metrics, logs, traces) for enterprise-wide deployment.
2 Drive the strategic transition from legacy monitoring tools (Geneos ITRS) to a modern, unified observability stack.
3 Administer and scale a containerized observability platform on OpenShift/Kubernetes.
4 Automate operational tasks, deployments, and configurations using Python, Bash, and Helm.
5 Collaborate with application teams to define standards and promote observability best practices.
6 Providing in-depth analysis with interpretive thinking to define problems and develop innovative solutions.
7 Directly impacting the business by influencing strategic functional decisions through advice, counsel, or provided services.
8 Persuading and influencing others through strong and comprehensive communication and diplomacy skills.
9 Performing other duties and functions as assigned