Website:
sceneplay.com
Job details:
We are hiring a senior scalability engineer to help prepare a modern microservices backend for high-concurrency production usage. This role is focused on backend performance, reliability, capacity planning, async workflow design, database scalability, and observability.
This is not a generic DevOps-only role. We are looking for someone who can reason deeply about how backend systems behave under load: database pressure, queue depth, retries, rate limits, stuck jobs, async workflows, API latency, and failure recovery.
Responsibilities
- Analyze backend architecture and identify scalability bottlenecks across APIs, services, database, cache, queues, storage, and third-party integrations.
- Design and execute load tests for high-concurrency user flows.
- Tune Postgres queries, indexes, connection pools, pagination, and transaction patterns.
- Design reliable queue/background-job systems with retries, idempotency, dead-letter queues, and reconciliation jobs.
- Define observability for scale: metrics, dashboards, request IDs, error tracking, alerts, and stuck-state monitoring.
- Improve reliability of webhook-driven and async processing workflows.
- Partner with backend engineers to implement code-level and architecture-level scalability improvements.
- Recommend infrastructure and autoscaling changes based on measured bottlenecks.
Required Skills
- 4+ years of backend, platform, SRE, performance engineering, or scalability engineering experience.
- In-depth system design knowledge is of paramount importance.
- Strong understanding of high-concurrency backend systems and distributed systems failure modes.
- Deep Postgres experience: query plans, indexing, connection pools, slow queries, transaction design, and pagination.
- Experience with Kafka, Redis, queues, background workers, retry/backoff, DLQs, and idempotent processing.
- Hands-on load testing experience with k6, Artillery, Locust, JMeter, or similar.
- Strong observability experience with logs, metrics, dashboards, alerting, and error tracking.
- Cloud/container familiarity with AWS, Docker, ECS/Fargate, Kubernetes, or similar.
- Ability to work with backend application code and collaborate with engineering teams.
Nice To Have
- Production Node.js or NestJS experience.
- TypeORM or ORM performance tuning experience.
- DevOps/platform experience deploying and operating applications at scale.
- Experience with object storage, upload-heavy systems, media processing, or webhook-heavy workflows.
- OpenTelemetry or distributed tracing experience.
- Terraform, Pulumi, AWS CDK, or infrastructure-as-code experience.
What Success Looks Like
- We know what breaks first at 10x and 100x traffic.
- Critical backend flows have load-test baselines and dashboards.
- Database and queue bottlenecks are identified and mitigated.
- Async workflows are retryable, idempotent, and observable.
- The team has clear production scalability priorities before launch.
Employment Details
- Role: Senior Scalability / Platform Engineer
- Location: Remote only
- Seniority: Senior / Lead-level preferred
Click on Apply to know more.