Role Description
As a Site Reliability Engineer at Runware, you will help ensure these systems remain reliable, performant and resilient as we scale. This is a highly technical, hands-on role working across software, infrastructure and production operations to improve observability, reduce incidents, eliminate operational toil and build lasting improvements across complex distributed systems.
-
Own and improve the reliability, availability and performance of critical production services across the Runware platform
-
Define and evolve our reliability practices, including SLIs, SLOs, alerting, observability and production-readiness standards
-
Investigate complex production issues across distributed systems, APIs, networking, queues, databases and GPU-backed workloads, participating in our engineering on-call rotation
-
Lead and contribute to incident reviews and RCAs, turning recurring failure modes into lasting engineering improvements
-
Reduce operational toil through automation, automated remediation and improvements to deployment safety, recovery and system resilience
-
Work closely with Engineering and DevOps teams on capacity planning, performance, scaling and architectural improvements as the platform grows
Qualifications
-
Strong experience operating and troubleshooting production systems at scale in an SRE, Production Engineering, Platform Engineering or similar role
-
Strong understanding of distributed systems and comfortable debugging across applications, databases, queues, containers, networking and infrastructure
-
Experience designing and operating observability systems using metrics, logs and distributed tracing
-
Understanding of SRE principles including SLIs, SLOs, error budgets, capacity planning, incident management and reducing operational toil
-
Experience with Kubernetes, containers, IaC and automated deployment practices
-
Ability to write software and automation using languages such as Python, Go or PHP
-
Take strong ownership of production problems and comfortable participating in an engineering on-call rotation, taking issues from initial investigation through to long-term remediation
Requirements
-
Experience operating high-throughput or low-latency APIs and distributed systems
-
Experience with bare-metal infrastructure, GPU environments or AI and ML workloads
-
Experience with RabbitMQ or other distributed messaging and queueing systems
-
Experience operating MySQL, Redis, ClickHouse or similar production data systems
-
Experience with global traffic management, load balancing, CDN platforms and hybrid infrastructure environments
-
Experience building automated scaling, capacity management or self-healing systems
Benefits
-
Generous paid time off β vacation, sick days, public holidays
-
Meaningful stock options β share in the upside you create
-
Remote-first setup β work from home anywhere we can employ you
-
Flexible hours β own your schedule outside core collaboration blocks
-
Family leave β paid maternity, paternity, and caregiver time
-
Company retreats β twice-yearly gatherings in inspiring locations