Website:
7-elevengsc.com
Job details:
Role Description
Why Join 7-Eleven Global Solution Center?
When you join us, you'll embrace ownership as teams within specific product areas take responsibility for end-to-end solution delivery, supporting local teams and integrating new digital assets. Challenge yourself by contributing to products deployed across our extensive network of convenience stores, processing over a billion transactions annually. Build solutions for scale, addressing the diverse needs of our 84,000+ stores in 19 countries. Experience growth through cross-functional learning, encouraged and applauded at 7-Eleven GSC. With our size, stability, and resources, you can navigate a rewarding career. Embody leadership and service as 7-Eleven GSC remains dedicated to meeting the needs of customers and communities.
Why We Exist, Our Purpose and Our Transformation?
7-Eleven is dedicated to being a customer-centric, digitally empowered organization that seamlessly integrates our physical stores with digital offerings. Our goal is to redefine convenience by consistently providing top-notch customer experiences and solutions in a rapidly evolving consumer landscape. Anticipating customer preferences, we create and implement platforms that empower customers to shop, pay, and access products and services according to their preferences. To achieve success, we are driving a cultural shift anchored in leadership principles, supported by the realignment of organizational resources and processes.
At 7-Eleven we are guided by our
Leadership Principles. Each principle has a defined set of behaviours which help guide the 7-Eleven GSC team to Serve Customers and Support Stores.
- Be Customer Obsessed
- Be Courageous with Your Point of View
- Challenge the Status Quo
- Act Like an Entrepreneur
- Have an “It Can Be Done” Attitude
- Do the Right Thing
- Be Accountable
About This Opportunity
As a Senior AI Platform RunOps Engineer, you will own day-2 operations for the AI platform. This role is focused on availability, incident response, platform support, runbooks, operational governance, release readiness, service health, auditability, and continuous improvement across models, agents, gateways, MCP services, and supporting platform infrastructure.
Responsibilities
- Own the operational health of AI platform services, including availability, support readiness, monitoring, and operational escalations.
- Build and maintain production runbooks, support procedures, incident response workflows, and service recovery patterns for AI platform components.
- Monitor and improve platform performance across latency, throughput, error rates, token usage, and operational cost signals.
- Drive release readiness, environment promotion, rollback safety, and operational acceptance criteria for AI platform changes.
- Support the production operations of gateways, model-serving layers, agent runtimes, MCP services, and platform integrations.
- Implement and maintain operational controls around secrets, audit logging, identity, service access, and compliance evidence.
- Partner with platform and security teams on DR, failover, backup, resilience, and infrastructure recovery planning.
- Operate enterprise platform capabilities including AKS, managed identities, Key Vault integrations, SSO mappings, and certificate management where required.
- Provide production support inputs for AI platform governance, change management, onboarding, and service lifecycle decisions.
- Support Apigee X operational integration for AI-facing APIs, including traffic controls, monitoring, onboarding, and secure exposure patterns.
Required Qualifications
- 7+ years of experience in production operations, SRE, platform support, DevOps, or cloud infrastructure engineering.
- Strong hands-on experience supporting distributed systems in production, including Linux systems, containers, CI/CD, and cloud-native services.
- Strong experience in monitoring, logging, ing, incident response, and service health management.
- Experience writing and maintaining operational runbooks, troubleshooting guides, and recovery procedures.
- Experience with identity, access control, audit logging, secrets handling, and security-aware platform operations.
- Experience supporting AI, data, or analytics platforms in production environments is strongly preferred.
- Strong scripting or development capability in Python and other automation-friendly languages.
- Ability to collaborate effectively across engineering, security, support, and business teams during operational events.
Preferred Qualifications
- Experience with enterprise AI platform operations, including model-serving services, agent runtimes, prompt governance, or evaluation systems.
- Experience with Databricks, Dataiku, MLflow, Airflow, or similar platforms requiring production-grade operational governance.
- Experience with AKS or Kubernetes-backed platform operations, including autoscaling, certificate management, identity delegation, and cluster troubleshooting.
- Experience with Apigee X or similar API management platforms for production operations and secure service exposure.
Click on Apply to know more.