Redplum AI
Website:
redplum.com
Company:
https://www.linkedin.com/company/redplumai
Industries: Technology, Information and Internet
Job details:
About the Role
We're building the future of autonomous infrastructure operations.
Imagine a world where production incidents are detected, investigated, and resolved by AI before humans even know something is wrong. That's the future we're creating.
We're looking for an exceptional Platform Engineer who can build distributed systems from scratch, design world-class infrastructure, and help define the next generation of AI-powered SRE Agents—similar to products like Resolve.ai, but built for the future.
This is not a traditional DevOps or Platform Engineering role. You'll own the entire platform, from cloud infrastructure and Kubernetes to agent orchestration, observability, security, and production reliability. You'll work at the intersection of infrastructure, distributed systems, and AI.
If you enjoy solving problems that have never been solved before, you'll fit right in.
What You'll Build
- AI-powered SRE agents capable of monitoring, investigating, diagnosing, and remediating production incidents.
- A scalable platform capable of running thousands of autonomous agents simultaneously.
- Infrastructure that is reliable, secure, and designed for enterprise-scale workloads.
- Internal developer platform that enables rapid product development.
- Event-driven systems connecting observability, cloud infrastructure, Kubernetes, CI/CD, security, and AI models.
- Autonomous workflows for deployments, rollbacks, debugging, and production operations.
Responsibilities
Platform Engineering
- Design and build highly scalable cloud infrastructure from scratch.
- Build Kubernetes-native platforms supporting multi-tenant workloads.
- Create developer platforms that abstract infrastructure complexity.
- Own networking, storage, compute, IAM, secrets, and service architecture.
- Design systems for high availability, disaster recovery, and global scale.
- Build internal tooling that accelerates engineering velocity.
Infrastructure & Reliability
- Design production-grade observability using logs, metrics, traces, and events.
- Build incident detection and automated remediation pipelines.
- Improve reliability, latency, performance, and cost efficiency.
- Create scalable deployment strategies including canary, blue-green, and progressive delivery.
- Build resilient distributed systems capable of self-healing.
AI Agent Platform
- Build infrastructure powering autonomous SRE agents.
- Integrate LLMs with infrastructure APIs and production systems.
- Design agent execution frameworks, memory systems, planning engines, and tool orchestration.
- Build secure execution environments for autonomous actions.
- Enable AI agents to reason over logs, metrics, traces, dashboards, Kubernetes resources, cloud APIs, Git repositories, CI/CD pipelines, and incident history.
Engineering Excellence
- Own architecture decisions from prototype to production.
- Build reusable platform components and developer tooling.
- Define engineering standards and infrastructure best practices.
- Continuously improve platform scalability, reliability, and security.
What We're Looking For
Strong Fundamentals
- Expert understanding of Linux, networking, operating systems, and distributed systems.
- Strong computer science fundamentals.
- Excellent debugging and systems thinking skills.
- Ability to build production systems from first principles.
Infrastructure Expertise
Experience with several of the following:
- Kubernetes
- Docker
- AWS, GCP, or Azure
- Terraform/OpenTofu
- Helm
- Service Mesh (Istio, Linkerd)
- Kafka, NATS, RabbitMQ
- PostgreSQL, Redis
- Elasticsearch/OpenSearch
- Prometheus
- Grafana
- OpenTelemetry
- ArgoCD
- GitHub Actions
- Vault
- Cloudflare
- CDN architectures
AI Experience (Bonus)
- Experience building AI agents.
- Familiarity with LLM APIs.
- MCP (Model Context Protocol)
- LangGraph
- AutoGen
- CrewAI
- Agent orchestration frameworks
- RAG architectures
- Function/tool calling
- AI workflow automation
You Might Be a Great Fit If You
- Have built infrastructure from zero to production.
- Have designed systems serving millions of requests.
- Love debugging distributed systems.
- Automate everything.
- Think reliability should be engineered—not monitored.
- Care deeply about developer experience.
- Can move from architecture to implementation without handoffs.
- Want to build category-defining AI infrastructure.
Nice to Have
- Experience at fast-growing startups.
- Contributions to open-source infrastructure projects.
- Experience building developer platforms.
- Experience operating production Kubernetes clusters.
- Experience with observability platforms.
- Security engineering knowledge.
- Experience with incident management systems.
- Experience building internal engineering tools.
What Success Looks Like
Within your first year, you'll have:
- Built the foundational platform powering autonomous SRE agents.
- Created infrastructure capable of scaling to enterprise workloads.
- Reduced production operational overhead through intelligent automation.
- Enabled engineering teams to ship faster with a world-class developer platform.
- Helped define the architecture for one of the most advanced AI-powered reliability platforms in the industry.
Why Join Us?
This is a rare opportunity to build a category-defining AI infrastructure company from day one.
You'll work on some of the hardest engineering problems in modern software—building autonomous AI agents that can understand, investigate, and resolve production issues without human intervention. Every architectural decision you make will shape the future of the platform.
What You'll Get
- Work with one of India's strongest engineering and product teams, alongside people who have built products used by hundreds of millions of users globally.
- Build from day one—you'll have significant ownership over architecture, technology choices, and engineering culture.
- Solve world-class engineering challenges across distributed systems, Kubernetes, cloud infrastructure, AI agents, observability, and reliability.
- Competitive compensation with one of the best salary packages in the market for exceptional talent.
- Meaningful ESOPs, giving you real ownership in the company and the opportunity to share in its long-term success.
- High autonomy, low bureaucracy—move fast, ship often, and see your work directly impact customers.
- Learn every day by working with ambitious engineers and product leaders who love solving difficult problems.
- Help define an entirely new category at the intersection of AI, infrastructure, and autonomous operations.
If you're excited by building products and want to work with exceptionally talented people while owning meaningful parts of the technology stack, we'd love to meet you.
Click on Apply to know more.