Senior Site Reliability Engineer @Stratus
All Others
Salary unspecified
Remote Location
πŸ‡ΊπŸ‡Έ USA Only
Employment Type contract
Posted YDay

[Hiring] Senior Site Reliability Engineer @Stratus

YDay - Stratus is hiring a remote Senior Site Reliability Engineer. πŸ’Έ Salary: unspecified πŸ“Location: USA

Role Description

The Senior Site Reliability Engineer is accountable for how Stratus behaves in production. Reporting to the Senior Director, Platform Engineering, this role brings genuine SRE discipline to a platform that MEP contractors run their fabrication shops on β€” where downtime does not mean a degraded experience, it means work stops on a job site. Stratus is a ~50-person, primarily remote, Series B software company growing quickly.

Unrelenting Reliability is one of our company values, and this is the role that operationalizes it. You will define what reliable means in numbers, instrument the system so we can see it, and close the loop from production signal back into engineering priority. This is an engineering role with a production mandate: you will write code, tune queries, build alerting, run load tests, and lead incidents β€” and you will be measured on customer-visible availability and latency, not on tickets closed.

The load-bearing problems on your plate are:

  • Establishing service level objectives and an error budget the whole engineering organization operates against, with the measurement infrastructure to back them.
  • Building the detection and response capability β€” high-signal alerting, clear on-call and escalation paths, per-system incident ownership, and well-exercised recovery paths β€” so that production problems are caught and resolved fast.
  • Building the performance and capacity engineering practice for our data and messaging layers, with headroom measured rather than assumed.

You will work across every engineering pod, with our platform and security functions, and with the customer-facing teams who see reliability problems first. The right candidate is comfortable being the person who says a number out loud and defends it, and is drawn to a place where the reliability practice is being built rather than maintained.

Qualifications

  • 6+ years of professional engineering experience, with at least 3+ years in a dedicated SRE or production engineering role at a B2B SaaS company.
  • Demonstrated ownership of SLOs and error budgets in production β€” you have defined them, measured them, argued about them with product leadership, and changed engineering behavior with them.
  • Deep, practical observability skills with Prometheus, Loki, Tempo, and Grafana or comparable: you have instrumented real systems and built alerting that pages on customer impact rather than on CPU.
  • Hands-on incident command experience at meaningful severity, including building on-call and escalation practice from the ground up.
  • Strong database performance skills β€” query profiling, index design, connection pooling, and diagnosing saturation under load. MongoDB experience strongly preferred; comparable document or relational depth acceptable.
  • Production Kubernetes experience (AKS preferred) sufficient to debug a live problem β€” pod scheduling, resource limits, networking, and service mesh behavior (Istio preferred).
  • Solid coding ability in at least one general-purpose language (Go, Python, C#, or TypeScript); willingness to work in a C#/.NET codebase.
  • Experience with load and performance testing tooling (k6, JMeter, Gatling, or comparable) and with turning results into engineering priority.
  • Production experience on Azure or AWS, with real understanding of the failure modes of managed services.
  • Fluency with AI-assisted engineering tooling and a track record of designing AI-leveraged workflows for your team β€” this is a graded expectation at every level at Stratus.
  • Excellent written communication: you write post-mortems, runbooks, and reliability reports that executives and engineers both read and act on.
  • Judgment and steadiness under pressure, and the credibility to tell engineering and product leadership something they do not want to hear.

Requirements

  • Experience establishing or maturing an SRE practice β€” defining the discipline, not inheriting it.
  • Experience with Sentry or comparable application error-monitoring platforms.
  • Experience operating event-driven and real-time systems β€” message brokers (Azure Service Bus, Kafka) and websocket or push layers (SignalR or comparable).
  • Experience operating MongoDB Atlas at production scale, including replica set topology and Atlas performance tooling.
  • Experience with durable workflow orchestration (Temporal or comparable).
  • Background in multi-region or multi-zone architecture and DR design against stated recovery objectives.
  • Experience with incident.io, PagerDuty, or comparable incident management platforms.
  • Familiarity with DORA metrics and with reliability work inside SOC 2 or NIST 800-171 scope.
  • Experience with legacy monolith reliability β€” improving the operational behavior of an ASP.NET or comparable application you cannot rewrite.
  • Domain interest in MEP, BIM, AEC, or construction technology.
  • Prior experience in a Series B / growth-stage company navigating the transition from product-market fit to scale.

Benefits

  • Comprehensive and competitive health benefits plan
  • Matching 401k contributions
  • 20 days annual PTO
  • Primarily remote work with occasional annual team onsites.

Company Description

Stratus participates in E-Verify. After you join the team, we'll verify your eligibility to work in the U.S. by submitting information from your Form I-9 to the Social Security Administration and, if needed, the Department of Homeland Security. This process happens post-hire only - we never use E-Verify to pre-screen applicants.

Before You Apply
️
πŸ‡ΊπŸ‡Έ Be aware of the location restriction for this remote position: USA Only
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Senior Site Reliability Engineer @Stratus
All Others
Salary unspecified
Remote Location
πŸ‡ΊπŸ‡Έ USA Only
Employment Type contract
Posted YDay
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
️
πŸ‡ΊπŸ‡Έ Be aware of the location restriction for this remote position: USA Only
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
Γ—
Apply to the best remote jobs
before everyone else

Access 125,000+ vetted remote jobs and get daily alerts.

4.9 β˜…β˜…β˜…β˜…β˜… from 500+ reviews

⚑ 126,674+ remote jobs, refreshed hourly

πŸ”” Real-time alerts: Apply first, direct to employer

πŸ›‘οΈ Vetted companies, no scams, true remote only

Unlock All Jobs Now

Maybe later