Website:
talent500.com
Job details:
About T-Mobile:
T-Mobile US, Inc. (NASDAQ: TMUS), headquartered in Bellevue, Washington, is America’s supercharged Un-carrier, connecting millions through its strong nationwide network and flagship brands, T-Mobile and Metro by T-Mobile. Customers benefit from an unmatched combination of value, quality, and exceptional service experience.
TMUS Global Solutions:
TMUS Global Solutions is a world-class technology powerhouse accelerating the company’s global digital transformation. With a culture built on growth, inclusivity, and global collaboration, the teams here drive innovation at scale, powered by bold thinking.
Job Overview:
At T-Mobile, we don’t just build technology — we empower people. We believe in investing in YOU — your growth, your impact, and your future. We’re unstoppable when individuals like you come together to solve bold challenges, inspire innovation, and build platforms that serve millions.
As a Senior Site Reliability Engineer, you’ll join a world-class engineering team focused on building and scaling intelligent infrastructure for LLM-based applications, AI services, and enterprise-scale backend systems. You’ll contribute to the design and implementation of observability, automation, and incident response strategies that ensure our platforms are high-performing, reliable, and cost-effective. You’ll play a key role in driving operational excellence, supporting platform scalability, and collaborating across engineering and architecture teams. This role provides growth opportunities to influence large-scale architecture and AI/ML reliability.
What You’ll Do:
- Implement and maintain observability, monitoring, and alerting systems for AI platforms and mission-critical backend services serving the T-Mobile Supply Chain domain.
- Collaborate in designing telemetry pipelines, logging infrastructure, and metrics dashboards using tools such as Splunk, Prometheus, Grafana, and Open Telemetry.
- Contribute to the development of SLOs, SLIs, and real-time health indicators across platform services and APIs.
- Participate in on-call rotations and lead resolution of high-impact incidents, including root cause analysis and postmortem reporting.
- Work closely with platform engineering teams to enforce governance, compliance, and security standards in production environments.
- Improve deployment pipelines, CI/CD workflows, and infrastructure automation (e.g. GitLab).
- Tune and scale infrastructure components such as Kafka, HAP Roxy, RMQ, databases, and distributed APIs.
- Support capacity planning, cost analysis, and system tuning to optimize platform performance.
- Advocate for automation-first operations, reducing toil through scripting and system reliability tooling.
- Contribute to documentation, runbooks, and knowledge sharing across SRE and engineering teams.
- Mentor junior engineers and participate in the culture of technical rigor and continuous improvement.
What You’ll Bring:
- Bachelor’s degree in Computer Science, Engineering, or a related field (Master’s preferred).
- 7+ years of experience in SRE, DevOps, or operations engineering in cloud-based environments.
Engineering & Operations:
- CI/CD for integrations and infrastructure
- Automated testing (integration/regression)
- Performance & scalability engineering
- SRE fundamentals (SLIs, SLOs, RCA)
- Platform cost visibility & optimization
- Security-first design & compliance readiness
- Hands-on experience with monitoring, alerting, and incident response in distributed systems.
- Strong coding/scripting ability in Python, Java, or shell scripting languages like Bash or PowerShell.
- Proficiency in CI/CD pipelines, GitLab workflows
- Strong working knowledge of SQL and NoSQL databases, including Oracle DB and MongoDB.
- Working knowledge of AI/ML systems, APIs, and modern LLM tooling is a strong plus.
- Familiarity with observability tools such as Splunk, Grafana, Prometheus.
- Experience with Kubernetes, container orchestration, and hybrid/multi-cloud deployments (Azure, AWS, GCP, OCI).
- Demonstrated ability to work in fast-paced, incident-driven environments with high stakes and uptime requirements.
Preferred Qualifications:
- Experience supporting AI workloads, model inference systems, or LLM-enabled platforms.
- Exposure to AIOps or related ML platform observability and reliability practices.
- Experience in highly regulated finance industry with understanding of compliance and audit controls.
- Familiarity with platforms such as Open AI or AI Gateway patterns.
- Background in building secure, zero-downtime platforms with enterprise-scale SLAs.
Must Have Skills:
- Strong grasp of SRE best practices including SLOs, SLIs, postmortems, and chaos engineering.
- Proficiency in diagnosing system bottlenecks across infrastructure, application, and network layers.
- Experience driving automation across observability, configuration, and deployment domains.
- Experience integrating with OFSLL or 3rd party finance systems
- Hands-on experience of OCI, OFSLL maintenance and support
- Strong communicator and collaborator in cross-functional technical teams.
- Curiosity-driven mindset with a passion for learning emerging AI technologies and systems reliability.
Why Join T-Mobile India:
At T-Mobile India, you won’t just contribute to world-class technology—you’ll help build it. You’ll work with global leaders, solve complex system challenges, and build platforms that redefine how technology powers customer experience.
We’re more than just a telecom company—we’re a technology powerhouse leading the way in AI, data, and digital innovation. And we do it all with heart, grit, and a passion for empowering people.
Join us and shape the future of intelligent platforms that serve millions — at the scale and speed of T-Mobile.
Click on Apply to know more.