Office Beacon ASPL
Website:
officebeacon.com
Job details:
Kindly find below the Job Description for the post of Senior AI Infrastructure Engineer.
Location: Remote
Shift: US Shift/ Night Shift
Position Overview
We are seeking a highly experienced AI Infrastructure Engineer to support and optimize advanced GPU-based server environments designed for AI/ML workloads and high-performance computing (HPC). The ideal candidate will possess deep expertise in NVIDIA GPU infrastructure, Linux systems, distributed compute environments, Mellanox InfiniBand networking, and AI cluster operations. This role involves Level 3 (L3) support, troubleshooting, optimization, and maintenance of enterprise AI compute environments containing high-density GPU servers.
Key Responsibilities
- Manage and support high-performance AI server infrastructure with:
- 192+ CPU cores
- 2TB RAM configurations
- 8x NVIDIA GPU server environments
- Troubleshoot complex GPU infrastructure issues involving:
- NVIDIA GPU hardware
- CUDA environments
- GPU drivers and firmware
- NVLink and GPU fabric architecture
- PCIe performance issues
- Support and optimize Mellanox InfiniBand networking environments for low-latency AI workloads and distributed compute operations.
- Monitor and maintain AI/HPC clusters used for machine learning training and inference workloads.
- Diagnose and resolve Linux system-level issues across GPU compute nodes.
- Work with containerized and orchestration technologies such as:
- Docker
- Kubernetes
- Slurm
- NVIDIA Container Toolkit
- Support distributed AI workloads and GPU resource allocation across multiple systems.
- Perform system tuning, patching, firmware updates, and performance optimization.
- Collaborate with internal engineering teams to improve infrastructure scalability, reliability, and uptime.
- Develop documentation, SOPs, and escalation procedures for infrastructure operations.
Required Skills & Experience
- 5+ years of experience supporting Linux-based enterprise infrastructure
- Strong experience with NVIDIA GPU systems in AI/HPC environments
- Hands-on experience with CUDA and GPU driver management
- Experience supporting GPU clusters or AI training infrastructure
- Knowledge of Mellanox InfiniBand networking
- Strong understanding of distributed compute environments
- Experience with Docker and Kubernetes
- Experience troubleshooting performance bottlenecks in GPU environments
- Familiarity with HPC or AI datacenter infrastructure
Preferred Qualifications
- Experience with: NVIDIA B200/B300, NVLink, RDMA, Slurm workload manager, OpenStack, AI model training infrastructure
- NVIDIA certifications preferred
- RHCE or Kubernetes certifications are a plus
Soft Skills
- Strong analytical and troubleshooting abilities
- Ability to work independently in complex technical environments
- Excellent communication and documentation skills
- Comfortable working in fast-paced AI infrastructure environments
Support Level
Level 3 (L3) Infrastructure Support
Environment
AI Infrastructure / HPC / GPU Datacenter Operations
Click on Apply to know more.