TELUS Digital
Website:
telusdigital.com
Job details:
SDE 3 - DevOps Engineer
Bengaluru
5 Days Working - Hybrid
Telus Digital
*Responsibilities:*
* Own the reliability, scalability, and performance of production systems and services
* Design and implement highly available, fault-tolerant, and distributed infrastructure
* Define and drive observability strategy, including monitoring, logging, and alerting
* Build and maintain scalable CI/CD pipelines to enable fast and reliable deployment
* Automate infrastructure provisioning and operational workflows using IaC tools
* Lead incident management and root cause analysis (RCA) and implement preventive measures
* Define and track SLIs, SLOs, and SLAs aligned with business and product requirements
* Collaborate closely with engineering teams to improve system design, deployment processes, and operational excellence
* Optimize cloud infrastructure for cost, performance, and efficiency
* Own and improve on-call processes; mentor engineers in handling production incidents
* Drive best practices for security, networking, and infrastructure reliability
* Participate in architecture and design reviews to ensure system resilience and scalability
* Document system architecture, runbooks, and operational processes
*Requirements:*
* 8+ years of experience in a DevOps, SRE, or cloud engineering role
* Proven experience building software using Python, Go, Rust, or JavaScript, with strong scripting capabilities.
Leading Experience must.
*Hands-on experience with cloud platforms such as AWS, GCP, or Azure.
* Expertise in Infrastructure as Code tools like Terraform, Ansible, or CloudFormation.
* Strong experience with containerization (Docker) and orchestration (Kubernetes)
* Solid understanding of Linux systems, networking concepts, and security best practices (IAM, VPNs, firewalls)
* Experience with monitoring and observability tools like Prometheus, Grafana, ELK Stack, New Relic, or Datadog.
* Proven experience in building and maintaining CI/CD pipelines (CircleCI, ArgoCD, Jenkins, GitHub Actions, GitLab CI, etc
* Deep hands-on experience owning cloud security end-to-end.
* Strong debugging and troubleshooting skills for complex production systems.
* Comfortable operating in fast-moving engineering environments, balancing long-term infrastructure investment with immediate operational needs.
* Experience owning and continuously improving on-call processes - including rotation design, escalation policies, runbook culture, and post-incident review cadence.
*Nice to Have*
* Experience with large-scale distributed systems and microservices architecture.
* Experience building and running operators on Kubernetes / knowledge of internals.
* Hands-on experience with cost optimization and capacity planning in cloud environments.
* Experience mentoring engineers and driving engineering best practices.
* Prior involvement in architectural reviews and cross-team technical decision-making.
* Strong understanding of compliance, security standards, and DevSecOps practices.
* Experience building internal developer platforms or self-service infrastructure tools.
* Development exposure: Experience contributing to backend services or application code (e. g., APIs, microservices) with strong software engineering fundamentals.
* MLOps experience: Familiarity with deploying, monitoring, and managing ML models in production, including tools like ML pipelines, model versioning, and data workflows
Click on Apply to know more.