Website:
talent500.com
Job details:
About T-Mobile:
T-Mobile US, Inc. (NASDAQ: TMUS), headquartered in Bellevue, Washington, is America’s supercharged Un-carrier, connecting millions through its strong nationwide network and flagship brands, T-Mobile and Metro by T-Mobile. Customers benefit from an unmatched combination of value, quality, and exceptional service experience.
TMUS Global Solutions:
TMUS Global Solutions is a world-class technology powerhouse accelerating the company’s global digital transformation. With a culture built on growth, inclusivity, and global collaboration, the teams here drive innovation at scale, powered by bold thinking.
Site Reliability Engineer
Job Overview:
- At T-Mobile, we don’t just build technology — we empower people. We believe in investing in YOU — your growth, your leadership, and your long-term impact. We’re unstoppable when driven individuals come together to solve bold challenges, inspire innovation, and create platforms that power the future.
- This role ensures the reliability and resilience of digital infrastructure, enabling efficient software development and deployment. It focuses on automating processes and reducing manual effort to prevent operational incidents, improve system performance, and enhance overall operational efficiency.
- The role requires expertise in programming, scripting, incident response management, and a variety of technical tools to maintain system robustness, improve deployment quality, and drive continuous operational improvements. Success is measured by system stability, incident reduction, and ongoing gains in operational efficiency.
- As part of Enterprise Procurement Engineering organization, this role supports a large-scale enterprise device management procurement and application delivery ecosystem that enables critical operations.
- The Site Reliability Engineer leverages automation, CI/CD practices, scripting, observability, and incident management expertise to improve reliability, scalability, and operational efficiency across a complex technology environment.
- This work directly impacts organizational stability and customer experience by ensuring the availability, performance, and reliability of critical systems.
Key Responsibilities:
- Automate processes to accelerate software development and deployment while minimizing manual interventions, including CI/CD pipelines, deployment automation, and endpoint management workflows where appropriate.
- Design, build, and enhance automation solutions that improve operational efficiency, deployment consistency, and service reliability across complex enterprise environments.
- Enhance system reliability and resilience by identifying issues and implementing preventive measures to reduce downtime and improve operational stability.
- Conduct root cause analysis and collaborate with problem management teams to prevent incident recurrence and improve system operations.
- Leverage programming, scripting, and incident response expertise to improve system robustness, deployment quality, and operational efficiency.
- Implement and support secure automation practices, including secrets management, credential lifecycle automation, and integration with approved enterprise security platforms.
- Partner with product, engineering, cybersecurity, and operations teams to design and implement scalable deployment and automation solutions.
- Support modernization initiatives involving mobile device management platforms, application deployment automation, platform migrations, and operational process improvements.
- Improve deployment visibility, monitoring, operational reporting, and observability capabilities to enhance traceability and operational awareness.
- Apply problem-solving and analytical skills to prevent operational incidents and maintain system stability.
- Continuously learn new skills and technologies to adapt to changing environments and drive innovation.
- Perform other duties and projects as assigned.
Qualifications:
- Bachelor's degree in Computer Science, Engineering, or a related technical field and 5+ years of relevant experience; or an advanced degree with 1+ year of relevant experience; or an equivalent combination of education and experience.
- 5–8 years of experience in Site Reliability Engineering (SRE), DevOps, platform engineering, operations, or software development environments.
- 5–8 years of experience troubleshooting production and customer-related issues while supporting business-critical systems and managing customer relationships.
- 5–8 years of experience developing automation and software solutions using Python, Bash, or similar programming languages.
- Experience designing, implementing, and maintaining CI/CD pipelines using GitLab or comparable platforms.
- Experience automating deployments, configuration management, infrastructure provisioning, and operational workflows.
- Experience integrating and automating enterprise secrets management solutions such as CyberArk, HashiCorp Vault, or similar platforms.
- Experience applying cybersecurity best practices and secure engineering principles across software delivery and operational environments.
- Experience with cloud platforms, infrastructure automation, containerized environments, and Kubernetes.
- Experience with observability, monitoring, deployment visibility, and operational reporting solutions.
- Strong understanding of SRE principles, DevOps practices, system reliability, resiliency, operational excellence, incident management, root cause analysis, and continuous service improvement.
- Proven ability to design, implement, and maintain reliable, scalable automation solutions for complex enterprise environments.
- Strong analytical and problem-solving skills with the ability to diagnose and resolve complex operational and production issues.
- Ability to collaborate effectively with product, engineering, cybersecurity, operations, and support teams while contributing to technical design decisions and engineering best practices.
- Excellent verbal and written communication skills, with the ability to explain technical concepts to both technical and non-technical audiences.
- Demonstrated ability to balance operational support, project delivery, modernization initiatives, and continuous improvement efforts.
- Experience mentoring peers, sharing technical knowledge, and adapting quickly to evolving technologies, tools, and business requirements.
- Commitment to automation, operational excellence, service reliability, and continuous improvement.
Licenses and Certifications:
Preferred:
- AWS Certified DevOps Engineer
- Certified Kubernetes Administrator (CKA)
- Google Cloud Certified – Professional DevOps Engineer
Click on Apply to know more.