- Location
- Pune District, Maharashtra, India
- Job type
- Full-time
Required skills
- Python
- Ansible
- BGP
- Datadog
- DevOps
- distributed storage
- DNS
- DPDK
- end-to-end
- incident response
- IP
- kernel
- Kubernetes
- Linux
- load balancing
- OpenStack
- SRE
- Swift
- TCP
- Terraform
About the role
Kaseya
Website:
kaseya.com
Job details:
Position Summary
- Looking for a Senior DevOps Engineer with strong hands-on Linux Administration, Systems, Networking and Infrastructure expertise.
- Own and optimize highly available, large-scale infrastructure supporting Kaseya’s Backup & Data Protection platform.
- Troubleshoot complex issues across Linux OS, storage, networking, Kubernetes and distributed systems.
- Automate infrastructure operations, improve reliability and drive operational excellence at scale.
- Work closely with Engineering, SRE and Architecture teams to build scalable and resilient infrastructure.
Required Skills
- 8–12 years of DevOps/SRE/Linux Infrastructure experience with strong Linux administration fundamentals.
- Advanced Linux troubleshooting: kernel/sysctl tuning, CPU/memory, processes, filesystems, I/O and performance analysis.
- Strong Networking: TCP/IP, DNS, routing, load balancing, VLANs, iptables/nftables, BGP/ECMP.
- Production experience with Kubernetes, Terraform/Pulumi, Ansible and CI/CD.
- Strong scripting skills in Shell, Python or Go for automation and operational tooling.
- Experience with Prometheus/Grafana/Datadog, incident management and production troubleshooting.
- Strong understanding of security, RBAC, secrets management and Linux hardening.
Desired Skills
- Experience with OpenStack, Ceph, Swift, Cinder, Neutron or distributed storage.
- Large-scale infrastructure experience across 1,000+ / 5,000+ servers or nodes.
- Bare-metal provisioning using PXE, Ironic, Foreman, IPMI/Redfish.
- Advanced networking: VXLAN/EVPN, OVS/OVN, SR-IOV, DPDK.
- Storage technologies such as NVMe/NVMe-oF, iSCSI, XFS, Ceph.
- Experience with backup/DR, RPO/RTO, replication and data protection platforms.
- Multi-cloud/private-cloud and infrastructure cost optimization experience.
Role Expectations
- Deep Linux expertise with hands-on troubleshooting of kernel, CPU, memory, filesystem and storage performance.
- Strong understanding of Linux networking and distributed infrastructure at scale.
- Ability to troubleshoot issues end-to-end across OS → Network → Storage → Kubernetes → Application.
- Build automation for provisioning, configuration, patching, monitoring and infrastructure operations.
- Drive SRE practices, observability, SLOs, incident response and root-cause analysis.
- Participate in designing highly available, secure and scalable infrastructure for a petabyte-scale backup platform.
- Ability to work independently on complex infrastructure problems and provide technical leadership/mentorship.
Click on Apply to know more.
This page is fully interactive when JavaScript is enabled. Please enable JavaScript to apply or browse related roles.