Website:
talent500.com
Job details:
About T-Mobile:
T-Mobile US, Inc. (NASDAQ: TMUS) is America’s supercharged Un-carrier, powered by an award-winning 5G network that connects more people in more places than ever before. With its unique value proposition of best network, best value, and best experiences, T-Mobile is redefining connectivity, fueling competition, and driving the next wave of innovation in wireless and beyond. Headquartered in Bellevue, Washington, T-Mobile provides services through its subsidiaries and operates its flagship brands, T-Mobile, Metro by T-Mobile, and Mint Mobile.
TMUS Global Solutions:
TMUS Global Solutions is a world-class technology organization accelerating T-Mobile’s global digital transformation. Our teams combine engineering talent, technology expertise, and collaborative ways of working to build secure, scalable solutions that improve customer and employee experiences. We foster innovation, agility, transparency, and strong enterprise partnerships to deliver measurable business outcomes.
About the Role:
- The Manager, Software Engineering – AI Platform Engineering leads T-Mobile's offshore engineering team responsible for delivering AI Platform & Agent Systems Engineering and Platform & Reliability Engineering capabilities. This team includes AI Platform Engineers, Platform & Reliability Engineers, and a Senior Product Owner working closely with onshore engineering, architecture, product, AI/ML, and platform teams.
- This is a hands-on engineering leadership role responsible for engineering execution, production reliability, delivery predictability, talent development, and operational excellence across AI platform services supporting conversational AI, agentic AI, enterprise integrations, and cloud-native platform capabilities.
- The manager partners with onshore engineering leaders to execute platform roadmaps, improve engineering practices, and ensure reliable operation of production AI services. Success is measured through engineering quality, platform reliability, predictable delivery, operational excellence, engineering capability growth, and effective collaboration across globally distributed teams.
What You'll Do:
- Lead, coach, and develop an offshore engineering team consisting of AI Platform Engineers, Platform & Reliability Engineers, and a Senior Product Owner.
- Drive predictable delivery of AI platform capabilities, platform engineering initiatives, and reliability improvements in partnership with onshore engineering and product teams.
- Partner with onshore architects, engineering managers, and product managers to execute roadmap priorities and ensure alignment across distributed teams.
- Foster engineering excellence through software development best practices, code quality, automated testing, CI/CD, and operational discipline.
- Ensure reliability, availability, observability, and operational readiness of production AI platform services.
- Support delivery of AI platform capabilities including agent orchestration, conversational platforms, SDKs, enterprise integrations, and Model Context Protocol (MCP)-enabled services.
- Lead hiring, performance management, mentoring, and career development while building a collaborative, high-performing engineering culture.
- Drive incident management, root cause analysis, and continuous improvement initiatives to enhance platform reliability and operational excellence.
- Ensure compliance with security, governance, and software engineering standards while optimizing cloud resources and platform costs.
- Communicate delivery progress, risks, dependencies, and engineering priorities to stakeholders across the organization.
What You'll Bring:
- Bachelor’s degree in Computer Science, Software Engineering, Information Systems, or a related field, or equivalent practical experience.
- 10+ years of software engineering, platform engineering, reliability engineering, or cloud-native systems experience.
- 5+ years leading engineering teams, including hiring, performance management, delivery accountability, and talent development.
- Demonstrated ownership of production reliability, incident management, SLOs, operational excellence, and service lifecycle management.
- Hands-on technical background sufficient to evaluate architecture, engineering quality, and technical tradeoffs.
- Experience with cloud-native technologies, including Kubernetes, CI/CD platforms, Infrastructure-as-Code, and modern observability tooling.
- Experience leading distributed or offshore engineering teams.
- Experience partnering with Product Owners or Product Managers in Agile delivery environments.
- Strong communication skills and the ability to influence engineering and business stakeholders.
Must Have Skills:
- Software engineering leadership and talent development
- Site Reliability Engineering, production operations, and incident management
- Kubernetes and cloud-native platform engineering
- CI/CD, Infrastructure-as-Code, and observability
- Distributed and offshore engineering team leadership
Nice-to-Have:
- Experience operating AI platforms, agentic AI systems, LLM-powered services, or AI infrastructure in production.
- Familiarity with LLM gateways, conversational AI platforms, AI observability, and AI governance frameworks.
- Experience leading multi-disciplinary teams spanning software engineering, platform engineering, reliability engineering, and product delivery functions.
- Experience building engineering organizations that support high-scale, customer-facing platforms.
- Track record of establishing reliability engineering and operational excellence practices for emerging platforms.
Click on Apply to know more.