Site Reliability Engineer @Runpod
All Others
Salary usd 150,000 - 2..
Remote Location
๐Ÿ‡บ๐Ÿ‡ธ USA Only
Employment Type full-time
Posted 2mths ago

[Hiring] Site Reliability Engineer @Runpod

2mths ago - Runpod is hiring a remote Site Reliability Engineer. ๐Ÿ’ธ Salary: usd 150,000 - 200,000 per year ๐Ÿ“Location: USA

Role Description

The Reliability team owns the availability, performance, and operational excellence of Runpodโ€™s global platform. While infrastructure teams build the systems, the Reliability team ensures those systems remain resilient, observable, and scalable under real-world production conditions.

This team is responsible for:

  • Defining and enforcing reliability standards across engineering
  • Designing incident response processes and improving recovery times
  • Building observability systems and reliability tooling
  • Driving SLO adoption and production readiness reviews
  • Reducing operational toil through automation

As a Site Reliability Engineer on the Reliability team, you will focus on ensuring the stability and resilience of Runpodโ€™s distributed platform. You will partner with engineering teams to improve system design, strengthen observability, and prevent incidents before they happen.

This role blends software engineering with production operations. Youโ€™ll work on reliability frameworks, SLO design, automation, and production hardening, reducing errors and improving performance across different services and infrastructure.

This is a high-impact role central to maintaining trust with developers running critical AI workloads on Runpod.

Your Impact

  • Increase platform uptime and reduce incident frequency and duration
  • Establish and operationalize SLIs/SLOs across services
  • Improve MTTR through better tooling, automation, and runbooks
  • Strengthen production readiness standards
  • Drive long-term systemic reliability improvements

You will influence how reliability is defined and measured across Runpod and help build the operational backbone of the company.

Responsibilities

  • Reliability Engineering
    • Define and implement SLIs/SLOs for critical services
    • Lead incident response and coordinate cross-team mitigation efforts
    • Conduct blameless postmortems and ensure corrective actions are completed
    • Perform production readiness reviews for new services and features
    • Identify systemic risks and drive preventative improvements
  • Observability & Monitoring
    • Design and improve monitoring, alerting, and dashboards (Prometheus, Grafana, etc.)
    • Improve signal-to-noise ratio in alerts and reduce alert fatigue
    • Build internal tooling for reliability tracking and reporting
    • Improve visibility into GPU performance and distributed systems health
  • Automation & Toil Reduction
    • Automate recurring operational workflows
    • Build tools and scripts (Python, Go, Bash) to eliminate manual processes
    • Improve deployment safety through automation and guardrails
    • Strengthen CI/CD reliability and release processes
  • Cross-Functional Reliability Advocacy
    • Partner with engineering teams to improve system resilience
    • Provide guidance on fault tolerance, scalability, and failure handling
    • Contribute to architectural discussions with a reliability-first mindset

Qualifications

  • 5+ years of experience in SRE, Reliability Engineering, or Production Engineering
  • Strong Linux systems and Networking expertise
  • Experience managing containerized production systems
  • Strong understanding of distributed systems and failure modes
  • Experience defining and managing SLIs/SLOs
  • Proven incident response and postmortem leadership experience
  • Strong scripting or programming skills
  • Experience with monitoring and alerting systems
  • Excellent written communication skills
  • Successful completion of a background check

Preferred

  • Experience with GPU infrastructure or AI/ML platforms
  • Experience improving reliability in high-growth or large scale environments
  • Familiarity with GPU observability tooling
  • Experience with Infrastructure as Code
  • Experience working in startup environments
  • Experience building internal reliability platforms or frameworks

Benefits

  • The competitive base pay for this position ranges from $150,000- $200,000 USD.
  • Meaningful equity in a fast-growing company - everyone on the team receives stock options.
  • Generous medical, dental & vision plans.
  • Flexible PTO - take the time you need to recharge.
  • Most roles are remote work first with an inclusive, collaborative teams utilizing Slack as the main form of internal communication.
  • Join a passionate team on the cutting edge of AI infrastructure.
Before You Apply
๏ธ
๐Ÿ‡บ๐Ÿ‡ธ Be aware of the location restriction for this remote position: USA Only
โ€ผ Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Site Reliability Engineer @Runpod
All Others
Salary usd 150,000 - 2..
Remote Location
๐Ÿ‡บ๐Ÿ‡ธ USA Only
Employment Type full-time
Posted 2mths ago
Apply for this position
Did not apply โœ“
Applied โœ“
Sent Follow-Up โœ“
Interview Scheduled โœ“
Interview Completed โœ“
Offer Accepted โœ“
Offer Declined โœ“
Application Denied โœ“
Unlock 125,000+ Remote Jobs
๏ธ
๐Ÿ‡บ๐Ÿ‡ธ Be aware of the location restriction for this remote position: USA Only
โ€ผ Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply โœ“
Applied โœ“
Sent Follow-Up โœ“
Interview Scheduled โœ“
Interview Completed โœ“
Offer Accepted โœ“
Offer Declined โœ“
Application Denied โœ“
Unlock 125,000+ Remote Jobs
ร—
Apply to the best remote jobs
before everyone else

Access 125,000+ vetted remote jobs and get daily alerts.

4.9 โ˜…โ˜…โ˜…โ˜…โ˜… from 500+ reviews

โšก 128,758+ remote jobs, refreshed hourly

๐Ÿ”” Real-time alerts: Apply first

๐Ÿ›ก๏ธ Vetted companies, no scams, true remote only, direct access to employer

Unlock All Jobs Now

Maybe later