Quess Corp Limited
Website:
quesscorp.com
Job details:
Position Title: EMT Lead – Event Management Technology | Experience Required: 7–8 Years
Function: IT Operations / Technical Operations Center (TOC)| Employment Type: Full-Time
Role Summary
We are seeking an experienced Event Management Technology (EMT) Lead to own the enterprise event management platform and strategy — correlating, filtering, and prioritizing events generated across monitoring tools such as Dynatrace, SolarWinds, and VeloCloud into meaningful, actionable alerts. This role sits at the intersection of monitoring, automation, and incident response, and is responsible for reducing alert noise, improving signal-to-noise ratio, and ensuring the right events reach the right teams at the right time within a 24x7 TOC environment.
Key Responsibilities
Event Management Platform Ownership
- Own the enterprise event management/correlation platform, integrating event feeds from Dynatrace, SolarWinds, VeloCloud, and other monitoring/observability tools.
- Design and maintain event correlation rules, deduplication logic, suppression policies, and event enrichment to reduce alert fatigue and noise.
- Define event severity/priority frameworks aligned with business impact and service criticality.
- Continuously tune event thresholds and correlation logic based on incident trends and false-positive analysis.
Team Leadership & TOC Operations
- Lead and mentor a team of event/monitoring engineers operating in a 24x7 TOC, ensuring consistent shift coverage and escalation discipline.
- Own the event-to-incident lifecycle — ensuring critical events are correctly triaged, escalated, and routed to the right resolver groups.
- Drive shift handovers, ticket hygiene, and documentation standards across the team.
- Conduct performance reviews, skills assessments, and training for team members on event management tools and processes.
Integration & Automation
- Integrate the event management platform with ITSM tools (ServiceNow, Remedy, Jira Service Management) for automated incident/ticket creation.
- Drive auto-remediation and self-healing workflows for known, recurring event patterns.
- Partner with monitoring tool owners (Dynatrace, SolarWinds, VeloCloud admins) to ensure clean, well-structured event data feeds into the platform.
- Support AIOps initiatives — anomaly detection, event clustering, and predictive alerting — to mature the event management capability.
Incident & Problem Management
- Act as an escalation point during major incidents (P1/P2), ensuring rapid event-to-response handoff to Incident Management.
- Drive RCA for event storms, missed escalations, or correlation failures, and implement preventive improvements.
- Partner with Incident/Problem Management to reduce MTTD (Mean Time to Detect) and MTTA (Mean Time to Acknowledge).
Reporting & Stakeholder Management
- Build dashboards and reports on event volumes, noise reduction metrics, correlation effectiveness, and escalation SLAs.
- Present regular updates to IT leadership on event management maturity, trends, and improvement roadmap.
- Collaborate with Network, Infrastructure, Application, and Security teams to align event priorities with business-critical services.
Required Skills & Experience
- 7–8 years of overall IT experience, with at least 3–4 years focused on event management, monitoring, or TOC leadership.
- Strong hands-on experience with event/alert data from Dynatrace, SolarWinds, and VeloCloud.
- Experience with event management/correlation platforms (e.g., ServiceNow Event Management, BMC TrueSight, Moogsoft, PagerDuty, xMatters, or similar).
- Solid understanding of ITSM processes (Incident, Problem, Change Management) and event-to-ticket workflows.
- Experience working in a 24x7 TOC/NOC environment, including shift management and major incident escalation.
- Proven ability to design correlation rules, suppression logic, and severity frameworks to reduce alert noise.
- Team leadership experience (managing 5+ engineers) in an operations environment.
- Strong stakeholder communication skills, with experience presenting to senior IT leadership.
Good to Have
- Exposure to AIOps platforms and machine-learning-based event correlation.
- Scripting/automation experience (Python, PowerShell) for event pipeline customization.
- Familiarity with cloud-native monitoring/event sources (Azure Monitor, AWS CloudWatch).
- Relevant certifications (ITIL Foundation, Dynatrace, ServiceNow Event Management).
Education
- Bachelor's degree in Computer Science, Information Technology, or related field (or equivalent practical experience).
Click on Apply to know more.