Source description
About the role
We are hiring a senior scalability engineer to help prepare a modern microservices backend for high-concurrency production usage. This role is focused on backend performance, reliability, capacity planning, async workflow design, database scalability, and observability. This is not a generic Dev Ops-only role. We are looking for someone who can reason deeply about how backend systems behave under load: database pressure, queue depth, retries, rate limits, stuck jobs, async workflows, API latency, and failure recovery. Responsibilities - Analyze backend architecture and identify scalability bottlenecks across APIs, services, database, cache, queues, storage, and third-party integrations. - Design and execute load tests for high-concurrency user flows. - Tune Postgres queries, indexes, connection pools, pagination, and transaction patterns. - Design reliable queue/background-job systems with retries, idempotency, dead-letter queues, and reconciliation jobs. - Define observability for scale: metrics, dashboards, request IDs, error tracking, alerts, and stuck-state monitoring. - Improve reliability of webhook-driven and async processing workflows. - Partner with backend engineers to implement code-level and architecture-level scalability improvements. - Recommend infrastructure and autoscaling changes based on measured bottlenecks. Required Skills - 4+ years of backend, platform, SRE, performance engineering, or scalability engineering experience. - In-depth system design knowledge is of paramount importance. - Solid understanding of high-concurrency backend systems and distributed systems failure modes. - Deep Postgres experience: query plans, indexing, connection pools, slow queries, transaction design, and pagination. - Experience with Kafka, Redis, queues, background workers, retry/backoff, DLQs, and idempotent processing. - Hands-on load testing experience with k6, Artillery, Locust, JMeter, or similar. - Strong observability experience with logs, metrics, dashboards, alerting, and error tracking. - Cloud/container familiarity with AWS, Docker, ECS/Fargate, Kubernetes, or similar. - Ability to work with backend application code and collaborate with engineering teams. Nice To Have - Production Node.js or NestJS experience. - TypeORM or ORM performance tuning experience. - Dev Ops/platform experience deploying and operating applications at scale. - Experience with object storage, upload-heavy systems, media processing, or webhook-heavy workflows. - Open Telemetry or distributed tracing experience. - Terraform, Pulumi, AWS CDK, or infrastructure-as-code experience. What Success Looks Like - We know what breaks first at 10x and 100x traffic. - Critical backend flows have load-test baselines and dashboards. - Database and queue bottlenecks are identified and mitigated. - Async workflows are retryable, idempotent, and observable. - The team has clear production scalability priorities before launch. Employment Details - Role: Senior Scalability / Platform Engineer - Location: Remote only - Seniority: Senior / Lead-level preferred We are hiring a senior scalability engineer to help prepare a modern microservices backend for high-concurrency production usage. This role is focused on backend performance, reliability, capacity planning, async workflow design, database scalability, and observability. This is not a generic Dev Ops-only role. We are looking for someone who can reason deeply about how backend systems behave under load: database pressure, queue depth, retries, rate limits, stuck jobs, async workflows, API latency, and failure recovery. Responsibilities - Analyze backend architecture and identify scalability bottlenecks across APIs, services, database, cache, queues, storage, and third-party integrations. - Design and execute load tests for high-concurrency user flows. - Tune Postgres queries, indexes, connection pools, pagination, and transaction patterns. - Design reliable queue/background-job systems with retries, idempotency, dead-letter queues, and reconciliation jobs. - Define observability for scale: metrics, dashboards, request IDs, error tracking, alerts, and stuck-state monitoring. - Improve reliability of webhook-driven and async processing workflows. - Partner with backend engineers to implement code-level and architecture-level scalability improvements. - Recommend infrastructure and autoscaling changes based on measured bottlenecks. Required Skills - 4+ years of backend, platform, SRE, performance engineering, or scalability engineering experience. - In-depth system design knowledge is of paramount importance. - Solid understanding of high-concurrency backend systems and distributed systems failure modes. - Deep Postgres experience: query plans, indexing, connection pools, slow queries, transaction design, and pagination. - Experience with Kafka, Redis, queues, background workers, retry/backoff, DLQs, and idempotent processing. - Hands-on load testing experience
More at Sceneplay