TELUS Digital
Website:
telusdigital.com
Job details:
Job Description
Telus International AI Data Solutions, or TI AI for short, has been a global leader in AI data services since 2005. Listed in Feb 2021 on the NYSE, TI has a market cap of USD 3 billion. With the acquisition of Playment in July 2021, TI now offers end-to-end AI training data solutions for video, images, audio, and text, serving the world’s top ML teams. Our platform supports growing AI and ML use cases of customers across industries - including automotive, agritech, insurance, retail, aerial imaging, Extended Reality (AR/VR), and more.
Why is it important?
High-quality training data is the most crucial building block of ML, AI, and GenAI algorithms that are shaping tech advancements today. Enterprises adopting AI spend a considerable amount of their time in pre- and post-processing this data and require a combination of sophisticated tools
supporting multiple data types, as well as a skilled and diverse workforce to ensure high-quality
labelled data.
Since quality training data is the first step in quality AI solutions, our AI Data platforms and solutions are designed to cover end-to-end requirements for customers - from collection, selection, annotation, and even advanced GenAI tasks. Our global AI community of 1M+ contributors also helps companies test and improve AI and ML models, with wide subject matter expertise across data types, languages, geographies, and cultural contexts - enabling smarter products, better customer solutions, and so much more. This helps our customers focus on business outcomes instead of advanced AIOps.
Responsibilities:
● Own the reliability, scalability, and performance of production systems and services.
● Design and implement highly available, fault-tolerant, and distributed infrastructure.
● Define and drive observability strategy, including monitoring, logging, and alerting.
● Build and maintain scalable CI/CD pipelines to enable fast and reliable deployments.
● Automate infrastructure provisioning and operational workflows using IaC tools.
● Lead incident management, root cause analysis (RCA), and implement preventive measures.
● Define and track SLIs, SLOs, and SLAs aligned with business and product requirements.
● Collaborate closely with engineering teams to improve system design, deployment processes,
and operational excellence.
● Optimise cloud infrastructure for cost, performance, and efficiency.
● Own and improve on-call processes; mentor engineers in handling production incidents.
● Drive best practices for security, networking, and infrastructure reliability.
● Participate in architecture and design reviews to ensure system resilience and scalability.
● Document system architecture, runbooks, and operational processes.
Requirements:
● 6+ years of experience in DevOps, SRE, or Cloud Engineering roles.
● Proven experience building software using Python, Go, Rust, or JavaScript, with strong scripting
capabilities.
● Hands-on experience with cloud platforms such as AWS, GCP, or Azure.
● Expertise in Infrastructure as Code tools like Terraform, Ansible, or CloudFormation.
● Strong experience with containerization (Docker) and orchestration (Kubernetes).
● Solid understanding of Linux systems, networking concepts, and security best practices (IAM,
VPNs, firewalls).
● Experience with monitoring and observability tools like Prometheus, Grafana, ELK Stack, New
Relic, or Datadog.
● Proven experience in building and maintaining CI/CD pipelines (CircleCI, ArgoCD, Jenkins,
GitHub Actions, GitLab CI, etc.).
● Deep hands-on experience owning cloud security end-to-end.
● Strong debugging and troubleshooting skills for complex production systems.
● Comfortable operating in fast-moving engineering environments, balancing long-term
infrastructure investment with immediate operational needs.
● Experience owning and continuously improving on-call processes - including rotation design,
escalation policies, runbook culture, and post-incident review cadence.
Nice to Have:
● Experience with large-scale distributed systems and microservices architecture.
● Experience building and running operators on Kubernetes / knowledge of internals.
● Hands-on experience with cost optimisation and capacity planning in cloud environments.
● Experience mentoring engineers and driving engineering best practices.
● Prior involvement in architectural reviews and cross-team technical decision-making.
● Strong understanding of compliance, security standards, and DevSecOps practices.
● Experience building internal developer platforms or self-service infrastructure tools.
● Development exposure – experience contributing to backend services or application code (e.g., APIs, microservices), with strong software engineering fundamentals.
● MLOps experience – familiarity with deploying, monitoring, and managing ML models in
production, including tools like ML pipelines, model versioning, and data workflows.
Click on Apply to know more.