Xloud Technologies
Website:
xloud.tech
Job details:
We’re looking for an engineer who can take ownership of designing, deploying and supporting GPU infrastructure for AI laboratories and research computing.
This role involves hands-on work across NVIDIA GPU servers, Lustre storage, high-speed networking and shared computing platforms, from initial architecture through deployment and ongoing support.
What you’ll work on:
• Design and deploy Linux-based GPU and HPC clusters.
• Configure NVIDIA drivers, CUDA, MIG and GPU workload scheduling.
• Deploy and support Lustre, including high availability, performance tuning and recovery.
• Manage containerised AI workloads using Kubernetes and HPC scheduling with Slurm.
• Configure and troubleshoot high-speed Ethernet, RDMA/RoCE and storage connectivity.
• Automate provisioning and maintenance using Ansible, Python and Bash.
• Run benchmarks, diagnose performance bottlenecks and validate failover.
• Lead technical discussions with customers and OEMs, document deployments and support laboratory administrators.
What we’re looking for:
• Proven experience deploying and supporting production HPC or GPU environments.
• Strong Linux administration and troubleshooting skills.
• Hands-on experience with NVIDIA GPUs and parallel storage, particularly Lustre.
• Understanding of compute, storage and networking dependencies.
• Ability to own technical delivery and resolve complex infrastructure incidents.
Additional experience with InfiniBand, BeeGFS, OpenPBS, NVIDIA AI Enterprise, Prometheus/Grafana, VMware or Proxmox would be valuable.
You’ll help build the infrastructure that researchers, faculty and students use for AI development and experimentation.
Click on Apply to know more.