Site Reliability Engineer @Sphere Partners
Devops
Salary unspecified
Remote Location
Employment Type full-time
Posted 2mths ago

[Hiring] Site Reliability Engineer @Sphere Partners

2mths ago - Sphere Partners is hiring a remote Site Reliability Engineer. πŸ’Έ Salary: unspecified πŸ“Location: Worldwide

Role Description

As a Site Reliability Engineer, you will own and improve the reliability, availability, and operational health of a large-scale cloud platform. You'll collaborate closely with Software Engineers, Infrastructure, Customer Operations, and Product teams while helping evolve production support processes and operational standards.

The role combines traditional SRE responsibilities with modern AI-assisted engineering practices, leveraging AI tools to improve incident response, documentation, operational workflows, and engineering productivity.

This position is ideal for someone who enjoys solving complex production challenges, improving observability, automating repetitive operational work, and building scalable reliability practices.

Responsibilities

  • Act as the primary technical escalation point for critical production incidents, providing hands-on support during high-severity outages.
  • Improve platform reliability by reviewing new product launches, infrastructure changes, and production readiness before release.
  • Design, implement, and optimize monitoring, alerting, and observability solutions across cloud infrastructure and applications.
  • Analyze operational metrics, recurring alerts, and incident trends to reduce alert fatigue and improve overall system health.
  • Lead incident investigations and post-mortems, ensuring root causes are identified and preventative actions are implemented.
  • Collaborate with Engineering, Infrastructure, Customer Operations, and external support teams to coordinate incident response and customer communications.
  • Participate in capacity planning, peak traffic readiness, disaster recovery exercises, and system performance reviews.
  • Develop and maintain operational runbooks, documentation, and incident response procedures.
  • Improve internal reliability tooling and automate operational workflows using modern AI-assisted development tools.
  • Drive continuous improvements in operational excellence through automation, standardization, and proactive reliability initiatives.

Qualifications

  • 4+ years of experience as a Site Reliability Engineer, DevOps Engineer, Production Engineer, or a similar infrastructure-focused role.
  • Strong experience supporting production systems running on AWS.
  • Hands-on experience with monitoring and observability platforms such as Datadog, AWS CloudWatch, New Relic, or similar.
  • Experience with incident management platforms such as PagerDuty.
  • Strong understanding of production incident management, root cause analysis, and post-incident review processes.
  • Experience working with ticketing and documentation platforms such as Jira and Confluence.
  • Familiarity with operational dashboards and reporting tools (Looker or similar BI platforms).
  • Experience building operational documentation, runbooks, and support processes.
  • Comfortable working outside regular business hours when critical production incidents require senior engineering support.
  • Experience using AI-assisted engineering tools (Claude, GitHub Copilot, Cursor, or similar) to improve engineering workflows, automate documentation, incident triage, reporting, or operational tasks.
  • Strong scripting or automation skills (Python, Bash, or similar) are considered a plus.

Nice to Have

  • Experience working in fintech, payments, financial services, or other high-availability environments.
  • Experience with Infrastructure as Code (Terraform, CloudFormation, or similar).
  • Familiarity with Kubernetes and containerized environments.
Before You Apply
️
worldwide Be aware of the location restriction for this remote position: Worldwide
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Site Reliability Engineer @Sphere Partners
Devops
Salary unspecified
Remote Location
Employment Type full-time
Posted 2mths ago
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
️
worldwide Be aware of the location restriction for this remote position: Worldwide
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
Γ—
Apply to the best remote jobs
before everyone else

Access 125,000+ vetted remote jobs and get daily alerts.

4.9 β˜…β˜…β˜…β˜…β˜… from 500+ reviews

⚑ 127,038+ remote jobs, refreshed hourly

πŸ”” Real-time alerts: Apply first, direct to employer

πŸ›‘οΈ Vetted companies, no scams, true remote only

Unlock All Jobs Now

Maybe later