VergeCloud
Website:
vergecloud.com
Job details:
About The Role
We're looking for a Senior SRE with 5+ years of experience to help operate and evolve
the infrastructure behind our CDN — from the baremetal edge nodes serving traffic, to
the Kubernetes-based control plane, to the observability and data pipelines that tell us
what's actually happening on the network. You'll work across the full stack: provisioning
and configuration management, CI/CD, Kubernetes, and the logging/monitoring systems
that let us catch problems before customers do.
This is a hands-on, senior individual-contributor role for someone who's comfortable
moving between low-level networking debugging and higher-level platform/infra
automation, and who can drive decisions with less oversight.
What You'll Do
Operate and maintain our fleet of self-managed baremetal edge servers, using
SaltStack (or Ansible) for configuration management and automation
Manage and improve our managed Kubernetes cluster, deploying and maintaining
services via Helm
Build and maintain CI/CD pipelines (GitHub Actions) for infrastructure and service
deployments
Own and extend observability tooling — Prometheus, Grafana, Alertmanager, and
distributed tracing with Grafana Tempo
Maintain and query ClickHouse at scale — millions of rows ingested daily, with a
focus on schema design and low-latency query tuning
Manage cloud service configuration as code using Terraform
Diagnose and resolve issues across the stack — from DNS resolution and
BGP/routing anomalies to TCP/IP-level performance regressions and application What
You'll Need
Requirements
What You'll Need
5+ years of experience in an SRE, DevOps, or infrastructure/platform engineering
role
Deep understanding of networking fundamentals — DNS, TCP/IP, and CDN concepts
(caching, routing, anycast, edge delivery) — you should be comfortable reading a
packet capture or debugging a DNS resolution chain
Solid experience operating Linux systems in production, including self-managed
baremetal infrastructure
Hands-on experience with configuration management tools — SaltStack or Ansible
Experience with Kubernetes in production, including deploying and managing
services via Helm
Experience building CI/CD pipelines, ideally with GitHub Actions
Working knowledge of Terraform or similar IaC tools
Practical experience with Prometheus, Grafana, and Alertmanager for monitoring and
alerting; familiarity with distributed tracing (Grafana Tempo or similar)
Ability to write code in Go or Python for automation, tooling, or internal services
Strong debugging skills across layers — from kernel/network to application to
infrastructure automation
Good communication skills and comfort working in a small, high-ownership team
Nice to Have
Experience with BGP / anycast routing in a production CDN or network operator
context
Experience With ClickHouse Or Another Columnar/analytical Database At Scale
Experience with Kafka/Redpanda or similar streaming systems for log/data pipelines
Prior experience at a CDN, ISP, hosting provider, or similar network-heavy operator
Experience with GitOps workflows (ArgoCD or similar)
Click on Apply to know more.