Website:
getswym.com
Job details:
Staff Platform Engineer — Infra & Common Layer — SwymRemote-first · Full-time · Bengaluru-based team. Sits within Platform Engineering, working closely with swym-integration, SRE, and the AI/automation function.
What Swym is and why this role exists
Every day, millions of shoppers save products they want but are not ready to buy. Most brands never know it happened. Swym fixes that. We capture what shoppers want at the earliest and most specific point in their journey, and we make that signal usable across every channel a brand has, until the purchase completes. 48,000+ brands use Swym today.
Building the feature is rarely the hard part anymore. With AI doing first-draft engineering, most features move from idea to working prototype fast. What doesn't move on its own is adoption: whether a merchant who has access to a capability actually turns it on, uses it, and keeps using it. That gap, between shipped and adopted, is what this role exists to close.
Every restock or price-drop event on a merchant's store fans out to thousands of alerts—email, SMS, push—across Wishlist Plus, Back in Stock, and Price Drop. As our footprint has scaled to tens of thousands of Shopify merchants, our high-fan-out pipelines—originally built on Redis—are reaching the limits of their original architecture. To support our next phase of enterprise growth, we are modernizing this foundation: introducing strict multi-tenant isolation, guaranteed message SLAs, and high-signal, actionable observability that eliminates operational noise.
Building a new queue architecture is straightforward. The real engineering craft lies in executing a zero-downtime migration of a live, high-throughput system without merchant disruption, while establishing the SLA discipline and observability that makes the next five years of this pipeline predictable and resilient. Moving our infrastructure from "handles current volume" to "provably reliable per message, per merchant" is what this role exists to achieve.
How we actually work
We are a lean team that runs on AI and GitHub. No wiki. If it matters, it is in the repo or it does not exist.
When a merchant workflow needs instrumenting, AI drafts the event schema. When a channel needs a first-pass message, AI writes it. When we need to know which merchants are underusing a feature, AI pulls the query. Humans decide what the data actually means, which channel earns the spend or the send, and what gets shipped to a merchant's inbox or in-app surface. That model runs across this role specifically: positioning is not this role's job, activation is.
What you will own
In your first 90 days — Master the pipeline and map the architecture- Get fluent in the swym-integration layer end-to-end: how a webhook becomes an alert, where Redis operates in that path, and where structural bottlenecks occur under peak load.
- Identify the highest-impact queue to migrate first and defend that strategy with real traffic data.
- Ship the first feature-toggled queue migration off Redis—live in production, with a thoroughly tested, zero-risk rollback path.
At six months — Scalable event architecture in motion- At least 5 core queues are live on the new pipeline behind feature toggles, with automated DLQ and replay mechanisms fully active in production.
- Achieve rock-solid queue stability across migrated pathways, maintaining 99.99%+ uptime even during high-volume peak sale events.
- Ship OpenTelemetry tracing from enqueue to completion for migrated paths, backed by per-merchant SLA dashboards that flag potential bottlenecks before they impact delivery.
At twelve months — Enterprise-grade reliability and legibility- Queue migration off Redis is complete or strictly scoped to deliberate edge-case workloads.
- A structured P1–P4 severity framework is live, complete with clear ownership and actionable runbooks that allow engineers across the organization to act with total confidence.
- Telemetry and trace data are exposed as queryable interfaces for both engineers and LLM agents.
- Implement service-to-service SLAs, noisy-neighbor isolation, and dedicated priority lanes for enterprise merchants.
- The success metric at twelve months isn't just migrated queues—it's an operating environment so resilient and legible that on-call duty becomes completely calm.
Who you are
- Proven Migration Experience: You have deep distributed systems experience and have led at least one zero-downtime architecture migration at real production scale. You know how to safely evolve live systems without dropping events or interrupting service.
- SLA & System Reliability Mindset: You think in failure modes, multi-tenant isolation, and explicit performance guarantees. When reviewing observability, your goal is high signal-to-noise ratio and immediate actionability.
- AI-First Workflow: You leverage AI to draft instrumentation, scaffold migration tools, and query traces—while keeping critical judgment on cutover risks, system boundaries, and architecture squarely in your hands.
- Background: 7+ years in Platform, SRE, or Distributed Infrastructure roles (as a guideline). Clojure experience and familiarity with high-fan-out notification systems are a strong plus, but engineering judgment and migration experience matter most.
What is in it for you
Most infrastructure modernizations are scoped as background "maintenance" bolted onto someone's existing workload without the authority to change core architecture. Here, you own the entire path: from evolving legacy queueing patterns to establishing the SLA and telemetry standards for our core messaging engine.
If your idea of a great career win is: "I took a foundational high-throughput system, modernized it live without dropping a single merchant event, and turned a complex pipeline into a bulletproof enterprise platform," you'll thrive here.
LogisticsCompensation: Competitive cash and meaningful equity. We will talk specifics early.
How to apply
Please apply directly through this LinkedIn job post or send your resume across to rishin.babu@swymcorp.com or sakshi.gupta@swymcorp.com
Click on Apply to know more.