- Location
- Mumbai Metropolitan Region
- Job type
- Full-time
Required skills
- Ansible
- configuration management
- DevOps
- Git
- infrastructure management
- Jenkins
- Kubernetes
- production support
- SRE
- ServiceNow
About the role
Infrasoft Technologies Ltd
Website:
kiya.ai
Job details:
Job Description
Role/Principal Accountabilities :
- Build the plan based on the in-house Observability stack.
- Deliver a plan to meet Operational Resilience monitoring requirements based on IBS/ CIF business functions: Settlement/Collateral etc.
- Produce a business dashboard to give visibility of business-critical services and show real-time service status.
- Align with the inhouse Observability team to ensure best practices are achieved.
- Align with ServiceNow Configuration Management to ensure any missing assets are added to ServiceNow.
- Work with Operational Resilience and infrastructure management to gain buy-in and visibility of the project.
- Lead a team of monitoring experts.
- Use SRE monitoring principles.
- Align with SRE management.
- Align with Infra management.
- Contribute to Observability standards and procedures.
- Work with the SRE team to optimize the operation model.
Secondary
- Align with Application teams to ensure their systems are integrated into the overall service dashboard.
Skills & Experience Required
- Strong communication presenting and documentation.
- Team management and stakeholder skills.
- Experience of driving change in a large organisation.
- Demonstrate understanding of large enterprise systems and the technologies on which they are built.
- Strong analytical skills and a solid understanding from a Production Support monitoring point of view.
- In-depth understanding of the technical aspects of delivering monitoring integration.
Knowledge Required
- Background in Observability Engineering covering Metrics, Logs and Traces and the interaction between telemetry.
- Strong Experience with Observability tools such as Open Telemetry, Grafana UI, Mimir, Loki, Tempo, Grafana Agent, Prometheus.
- Appreciation of SLI/SLOs for measuring and reporting on service levels.
- Awareness of notification frameworks for escalation automation.
- Experience of DevOps tools i.e. Jenkins, Ansible Tower, GIT and build & deploy pipelines.
- Exposure to Kubernetes, containers & micro-service concepts (not essential, but beneficial).
(ref:hirist.tech)
Click on Apply to know more.
This page is fully interactive when JavaScript is enabled. Please enable JavaScript to apply or browse related roles.