Website:
intellifyai.ai
Job details:
Role Description
We are looking for a highly skilled DevOps Engineer to join our engineering team and take ownership of our cloud infrastructure, deployment pipelines, and production environments.
The ideal candidate will have experience building and managing cloud-native infrastructure for generative AI products, taking solutions from MVP to production at scale. You will be responsible for designing secure, highly available, and scalable infrastructure across AWS and Azure, automating deployments, improving system reliability, and enabling rapid software delivery.
This is a hands-on Individual Contributor (IC) role requiring strong expertise in infrastructure automation, container orchestration, CI/CD, observability, and security. You will work closely with engineering, AI, and product teams to architect robust deployment strategies for enterprise-grade AI applications.
Experience with LiveKit, Pipecat, real-time media infrastructure, and self-hosted AI platforms is highly desirable.
Key Responsibility
- Design, implement, and manage cloud infrastructure across AWS and Azure.
- Build and maintain scalable CI/CD pipelines for automated application deployments.
- Architect highly available, secure, and fault-tolerant production environments.
- Manage Kubernetes clusters and containerised workloads using Docker.
- Automate infrastructure provisioning using Infrastructure as Code (Terraform, CloudFormation, or equivalent).
- Implement monitoring, logging, alerting, and observability using tools such as Prometheus, Grafana, ELK, OpenTelemetry, or similar.
- Optimise infrastructure performance, scalability, availability, and cost.
- Deploy and manage AI applications and supporting infrastructure from MVP through enterprise-scale production.
- Deploy and maintain self-hosted AI workloads, inference servers, and GPU-enabled environments.
- Configure networking, DNS, SSL, load balancing, firewalls, VPNs, and cloud security best practices.
- Collaborate with engineering teams to improve release processes, developer productivity, and platform reliability.
- Implement disaster recovery, backup strategies, and business continuity planning.
- Ensure infrastructure complies with security, governance, and operational best practices.
Qualifications
- 7+ years of hands-on DevOps or Cloud Infrastructure experience.
- Strong experience with AWS and Azure cloud platforms.
- Experience building and scaling AI products from MVP to enterprise production.
- Strong understanding of cloud architecture, solution design, and infrastructure planning.
- Hands-on experience with Kubernetes, Docker, and container orchestration.
- Strong knowledge of Infrastructure as Code using Terraform, CloudFormation, or equivalent.
- Experience with CI/CD tools such as GitHub Actions, Azure DevOps, Jenkins, GitLab CI, or similar.
- Experience with monitoring and observability tools, including Prometheus, Grafana, ELK, OpenTelemetry, or New Relic.
- Configure and maintain TURN/STUN servers (Coturn) for WebRTC connectivity.
- Experience with LiveKit, Pipecat, WebRTC, or real-time voice/video infrastructure is highly preferred.
- Experience deploying and managing Linux-based production servers.
- Strong scripting skills in Bash, Python, or PowerShell.
- Good understanding of networking, security, IAM, SSL/TLS, load balancing, and cloud-native architectures.
- Excellent troubleshooting and incident management skills.
- Strong communication and collaboration skills with cross-functional engineering teams.
- Bachelor's/Master's degree in Computer Science, Information Technology, or a related field.
- AWS, Azure, Kubernetes, or Terraform certifications are a plus.
- Experience working in AI, SaaS, or enterprise technology startups is highly preferred.
Nice to Have
- Experience hosting LLM inference servers
- Experience deploying Whisper, Deepgram, or speech processing pipelines
- Experience with AI agent frameworks
- Experience building production RAG infrastructure
- Knowledge of security compliance (ISO 27001, SOC 2)
- Cost optimisation across AWS and Azure
What we are looking for
- Strong troubleshooting and production support skills
- Ability to manage mission-critical production environments
- Ownership mindset with attention to reliability and scalability
- Experience working in fast-paced startup environments
- Excellent documentation and communication skills
- Ability to independently architect and deploy infrastructure
Click on Apply to know more.