Staff Site Reliability Engineer @Horizon3
Software Development
Salary usd 199,750 - 2..
Remote Location
πŸ‡ΊπŸ‡Έ USA Only
Employment Type full-time
Posted 1wk ago

[Hiring] Staff Site Reliability Engineer @Horizon3

1wk ago - Horizon3 is hiring a remote Staff Site Reliability Engineer. πŸ’Έ Salary: usd 199,750 - 270,000 per year πŸ“Location: USA

Role Description

We are seeking a hands-on Staff Site Reliability Engineer to own and evolve the reliability strategy, operating model, and engineering-wide standards supporting our platform. This is a foundational role for an experienced engineer who will set technical direction across teams, lead the highest-impact reliability initiatives, and establish the practices and systems that enable engineering to operate production services safely.

  • Own and evolve the engineering-wide SRE strategy, operating model, and reliability standards, aligning them to customer impact, business priorities, and risk.
  • Lead cross-functional alignment across Infrastructure, product, service, security, and business stakeholders to improve reliability, observability, incident response, and operational readiness across multiple teams.
  • Establish an organization-wide approach to service ownership, meaningful SLIs and SLOs, and error budgets for critical customer paths and services.
  • Define and drive adoption of observability standards across pipelines and platform components that report on service health, performance and operational risk.
  • Set the standard for dashboards, actionable alerting, runbooks, and escalation paths.
  • Drive end-to-end complex cross-functional reliability initiatives.
  • Set and raise the engineering-wide bar for incident management, incident command, on-call health, post-incident learning, and recovery readiness.
  • Shape the technical direction, operational model, and growth path of the SRE function.
  • Participate in a 24/7 on-call rotation and help design an on-call model that is sustainable, appropriately staffed, and continuously improved.

Qualifications

  • Experience designing, operating, and troubleshooting large scale distributed systems in production environments.
  • Deep knowledge of reliability engineering, observability, incident management, and production operations, with demonstrated ability to turn that knowledge into standards and practices adopted by others.
  • Experience in establishing SLIs, SLOs, actionable alerts, observability, and service ownership.
  • Backend experience building backend systems and automation that reduce operational toil, strengthen safeguards, and improve operational efficiency.
  • Experience in leading high severity incidents and improving incident response programs.
  • Excellent written and verbal communication skills including technical designs, runbooks, postmortems, and operational documentation.

Requirements

  • Python and Terraform (Infrastructure as Code), or equivalent automation and infrastructure-as-code tools.
  • Experience with Observability tools such as Datadog, New Relic, Grafana, or equivalent platforms.
  • Experience operating production services in AWS and Kubernetes.
  • Experience with CI/CD pipelines such as Gitlab CI, ArgoCD, or GitOps workflows.

Benefits

  • Inclusive Team: We value diversity and promote an inclusive culture where everyone can thrive.
  • Growth Opportunities: Be part of a dynamic and growing team with numerous career development opportunities.
  • Innovative Culture: Work in a collaborative environment that encourages creativity and out-of-the-box thinking.
  • Hybrid & Remote Work: We embrace a mix of remote and hybrid work models depending on role and location, including our Chicago office, where some roles require regular in-office presence.
  • Competitive Compensation: We offer competitive salary, equity and benefits. Our benefits include health, vision & dental insurance for you and your family, a flexible vacation policy, and generous parental leave.
Before You Apply
️
πŸ‡ΊπŸ‡Έ Be aware of the location restriction for this remote position: USA Only
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Staff Site Reliability Engineer @Horizon3
Software Development
Salary usd 199,750 - 2..
Remote Location
πŸ‡ΊπŸ‡Έ USA Only
Employment Type full-time
Posted 1wk ago
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
️
πŸ‡ΊπŸ‡Έ Be aware of the location restriction for this remote position: USA Only
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
Γ—
Apply to the best remote jobs
before everyone else

Access 125,000+ vetted remote jobs and get daily alerts.

4.9 β˜…β˜…β˜…β˜…β˜… from 500+ reviews

⚑ 126,991+ remote jobs, refreshed hourly

πŸ”” Real-time alerts: Apply first, direct to employer

πŸ›‘οΈ Vetted companies, no scams, true remote only

Unlock All Jobs Now

Maybe later