Phykon
Website:
phykon.com
Job details:
Location: Thiruvananthapuram, Kerala, India (Onsite)
Experience: 8+ Years
Employment Type: Full-Time
Department: Network Operations Center (NOC)
Reports To: [Head of Operations / Engineering – to be confirmed]
About The Role
We are seeking an experienced NOC / SRE Lead to own the health, observability, and incident response for a connected fleet of field-deployed systems and the data center infrastructure that supports it.
This is the keystone L3 role in our Network Operations Center (NOC). You will own deep diagnosis and root-cause analysis on the running system, command the hardest escalations, and build the standards, runbooks, observability architecture, and operator-certification program that the rest of the NOC inherits.
As the L3 Lead, you will act as the escalation point above L2 support teams, taking ownership of the most complex incidents while keeping escalations to L4 product engineering limited to genuine code and firmware issues — protecting both fleet uptime and engineering velocity.
This is a greenfield opportunity. In the near term, the role combines a hands-on L3 technical function with NOC leadership. As the NOC scales, the role is designed to grow into a Principal SRE or NOC Manager track.
Key Responsibilities
- Fleet Health & Observability
- Architect the monitoring stack the NOC runs on — scrape architecture, alert rules, dashboards, and SLOs — rather than only consuming it.
- Own end-to-end visibility of the connected fleet across data center, edge Kubernetes, Linux, and network layers.
- Continuously tune alerting to reduce noise and catch degradation early.
- Data Center & Infrastructure Operations
- Oversee the health and availability of data center and colocation infrastructure — compute, storage, network, power, and cooling dependencies — supporting the fleet.
- Coordinate with data center operators, colocation providers, and remote-hands teams for physical interventions, maintenance windows, and capacity changes.
- Manage hardware fault handling, RMA workflows, and site-level incident response across distributed data center and edge sites.
- Coordinate with telecommunications providers, colocation partners, cloud vendors, and infrastructure service providers to resolve service-impacting issues and maintain operational continuity.
- Incident Management & Response
- Serve as incident commander on the hardest escalations, including Sev1/P1 incidents, and drive them to resolution within SLA.
- Perform deep diagnosis and root-cause analysis on the running system using logs, metrics, and telemetry.
- Diagnose complex issues across:
- Network connectivity, routing, NAT/CGNAT, and tunnels
- Degraded or intermittent links to field-deployed hardware
- Production Kubernetes and Linux platform dependencies at the edge
- Assume end-to-end ownership of major incidents, service disruptions, and customer escalations until resolution and formal closure.
- Act as the highest operational escalation point within the NOC for complex technical and service- impacting incidents.
- Provide timely updates to internal stakeholders, leadership teams, and customers during major incidents and service outages.
- Standards, Runbooks & Operator Certification
- Codify diagnostic procedures, escalation paths, and operational standards that make the NOC repeatable and scalable.
- Own the operator-certification layer that qualifies operators to run and support the fleet.
- Develop and continuously improve runbooks and operational procedures inherited by L1/L2 teams.
- Escalation & Coordination
- Maintain escalation hygiene: reserve L4 product engineering for genuinely complex issues requiring code or firmware changes.
- Clearly define problem scope, business impact, and affected systems when escalating, with supporting logs and evidence.
- Facilitate technical communication across engineering and operations teams during major incidents, providing timely stakeholder updates.
- Continuous Service Improvement
- Lead post-incident reviews and drive corrective actions to closure.
- Identify recurring incidents and contribute preventive improvements and monitoring optimization.
- Client & Stakeholder Engagement
- Act as a primary technical point of contact for customers, partners, vendors, and internal stakeholders on operational and service-related matters.
- Participate in customer meetings, operational reviews, service reviews, and technical discussions as a subject matter expert.
- Provide clear, professional, and executive-level communication during major incidents, service disruptions, and operational escalations.
- Build and maintain strong working relationships with customer technical teams while driving accountability and service excellence.
- Leadership & Team Development
- Provide technical leadership, mentoring, and guidance to L1 and L2 engineers.
- Drive knowledge sharing, operational maturity, and continuous improvement initiatives across the NOC function.
- Support the development and maintenance of operator certification programs and competency frameworks.
- Participate in recruitment, onboarding, training, and capability development activities as the NOC organization scales.
Required Skills & Qualifications
- 8+ years in SRE, production engineering, or senior NOC roles with strong platform depth (not generalist IT-NOC).
- Production Kubernetes and Linux experience, ideally in edge or distributed environments rather than pure cloud.
- Ownership-level experience with Prometheus and Grafana — designing scrape architectures and alert rules, not just reading dashboards.
- Hands-on experience with data center operations — running production infrastructure across data center, colocation, or distributed edge sites.
- Strong practical networking: TCP/IP, NAT/CGNAT, tunnels, and degraded-link diagnosis.
- Demonstrated on-call leadership and Sev1 incident command in production environments.
- Proven ability to codify standards, runbooks, and certification programs for operations teams.
- Strong analytical, troubleshooting, and documentation skills, with excellent communication across engineering and business stakeholders.
- Strong ownership mindset with the ability to independently drive incidents, projects, and operational improvements to completion.
- Ability to remain calm, structured, and decisive during high-severity incidents and customer escalations.
- Excellent stakeholder management and communication skills, including the ability to communicate complex technical concepts to both technical and non-technical audiences.
- Experience interacting directly with enterprise customers, vendors, and senior stakeholders in operational environments.
Tools & Technologies
Observability & Telemetry:
Prometheus, Grafana, and related monitoring/alerting platforms
Platform:
Kubernetes, Linux (edge/distributed deployments)
Automation:
Infrastructure-as-Code tooling (advantage)
Incident & Collaboration:
Incident management, ticketing, and on-call platforms
Preferred Qualifications
- Greenfield NOC build experience — standing up a NOC, runbooks, and on-call from scratch (strongest positive signal).
- Prior NOC/SRE experience supporting a hardware or field-deployed product.
- Exposure to OT/BMS environments and comfort bridging IT and industrial systems.
- Relevant certifications (Kubernetes, Linux, networking, or cloud) are an advantage.
- Highly desired: Proven experience building or scaling a Network Operations Center (NOC) from the ground up, including monitoring strategy, runbooks, escalation models, operational processes, and team development.
Work Environment
- Operates within a 24×7 Network Operations Center supporting a live fleet of field-deployed systems and distributed data center infrastructure.
- May require participation in rotational shifts, weekend coverage, planned maintenance activities, and on-call support arrangements.
- Serves as a key escalation point for critical incidents and major customer-impacting events.
- Acts as a primary technical point of contact for customers, partners, and internal stakeholders during operational reviews and service incidents.
- Expected to perform effectively in high-pressure operational environments while maintaining clear communication, leadership, and decision-making capabilities.
- Provides technical leadership, mentoring, and operational guidance to L1 and L2 support teams.
Click on Apply to know more.