Senior Site Reliability Engineer @Banyan Software
Software Development
Salary usd 145,000 - 1..
Remote Location
Employment Type full-time
Posted 2wks ago

[Hiring] Senior Site Reliability Engineer @Banyan Software

2wks ago - Banyan Software is hiring a remote Senior Site Reliability Engineer. πŸ’Έ Salary: usd 145,000 - 170,000 per year πŸ“Location: USA, Canada

Role Description

We are seeking a highly experienced and hands-on SRE to own the operational excellence of the modernized SaaS applications produced by the Banyan AI Factory. This is not a role focused on building the factory itself; instead, you will run the reliability of the modernized applications the factory delivers to our Operating Companies (OpCos).

You will join a team that provides 24x7 coverage with rotating on-call responsibilities, serving as Tier 1 Site Reliability Engineering (SRE) for our OpCos’ distributed applications. Day to day this will include:

  • Automated deployments
  • Cloud service integration
  • Application performance and availability monitoring/observability
  • Security incident response across our two target clouds β€” Amazon Web Services (AWS) and Microsoft Azure

Key Responsibilities

  • 24x7 Operations & On-Call: Operate as part of a team providing round-the-clock coverage of OpCo containerized applications, participating in a rotating on-call schedule to ensure continuous availability and rapid response.
  • Tier 1 SRE & Operations: Serve as Tier 1 SRE for the modernized applications, managing day-to-day cloud integrations across our two target clouds β€” AWS and Azure β€” to keep production systems healthy, performant, and secure.
  • Performance & Availability Monitoring/Observability: Implement and maintain robust application observability tooling (monitoring, logging, tracing) to track performance and availability, proactively detect degradation, and drive down mean-time-to-detect and mean-time-to-resolve.
  • Disaster Recovery and Service Restoration: Develop, maintain, test, and execute disaster recovery and business continuity procedures. Ensure the timely recovery and restoration of services following geographic disruptions, cyber incidents, infrastructure failures, or other disaster events.
  • Security Incident Response: Respond to security incidents and operational events affecting OpCo SaaS platforms, executing established runbooks, coordinating remediation.
  • Automation & Infrastructure-as-Code: Use Infrastructure-as-Code (Terraform) and CI/CD pipelines (e.g., GitHub Actions, GitLab CI) to manage, deploy, and automate the operational environments of modernized applications, reducing toil and improving consistency.
  • AI Agents & DevSecOps Scale: Build scale in our DevSecOps practice by designing, building, and operating AI agents that automate SRE tasks and incident response, reducing toil and accelerating detection, triage, and remediation.
  • Hands-on Problem Solving: Serve as a technical escalation point for operational challenges, applying strong analytical skills to resolve infrastructure, network, and automation issues across distributed, multi-tenant SaaS environments while navigating technical ambiguity.

Qualifications

  • Experience: 5–7 years of progressive experience in Software Engineering, and/or Site Reliability Engineering, with a focus on operating distributed systems.
  • Automation Coding Experience: Deep expertise in Python, Javascript, or Go. Building automation and integrations between tools.
  • Containerization: Deep expertise in container technologies (Docker/Kubernetes) supporting highly scalable and resilient distributed systems.
  • Infrastructure-as-Code with Terraform: Have experience working with modules at scale.
  • Cloud Native Services: Hands-on experience operating production workloads on Amazon Web Services (AWS) and/or Microsoft Azure.
  • CI/CD & Automation: Deep history of hands-on work with CI/CD platforms (GitHub Actions, GitLab CI) and embedding DevSecOps practices directly into operational workflows.
  • Operations, Monitoring & Observability: Experience with application level logging, troubleshooting, and tracing tools.
  • AI-Fluent Engineering: Experience with AI-assisted engineering tools such as Claude Code or similar.
  • Application Performance Management (APM): Familiarity with APM tooling and practices.
  • Incident & Security Response: Demonstrated experience participating in on-call rotations, responding to production and security incidents, and executing disaster recovery procedures.
  • Communication & Collaboration: Exceptional communication, presentation, and collaboration skills.
  • Education: Bachelor’s degree in Computer Science or a related technical field.

Preferred Skills (A Plus)

  • Familiarity with advanced cloud security tools like Wiz, Prisma Cloud, and Checkov.

Salary Information

The expected base salary for this position is approximately USD $145,000 - $170,000 for US-based candidates and CAD $120,000 - $145,000 for Canada-based candidates, excluding annual bonus and equity (when applicable).

Diversity, Equity, Inclusion & Equal Employment Opportunity at Banyan

Banyan affirms that inequality is detrimental to our Global Teams, associates, our Operating Companies, and the communities we serve. As a collective, our goal is to impact lasting change through our actions. Together, we unite for equality and equity.

Please Note

Banyan Software does not accept unsolicited resumes or applications submitted via email, LinkedIn, or other direct channels. All candidates must apply through our official Careers site to be considered for employment.

Recruitment Notice

Banyan Software may use artificial intelligence (AI) tools to assist in screening and/or assessing applicants during the recruitment process.

Beware of Recruitment Scams

We have been made aware of individuals fraudulently posing as members of our Talent Acquisition team and extending fake job offers. Protect yourself by verifying communications and reporting suspicious messages.

Before You Apply
️
remote Be aware of the location restriction for this remote position: USA, Canada
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Senior Site Reliability Engineer @Banyan Software
Software Development
Salary usd 145,000 - 1..
Remote Location
Employment Type full-time
Posted 2wks ago
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
️
remote Be aware of the location restriction for this remote position: USA, Canada
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
Γ—
Apply to the best remote jobs
before everyone else

Access 125,000+ vetted remote jobs and get daily alerts.

4.9 β˜…β˜…β˜…β˜…β˜… from 500+ reviews

⚑ 126,845+ remote jobs, refreshed hourly

πŸ”” Real-time alerts: Apply first, direct to employer

πŸ›‘οΈ Vetted companies, no scams, true remote only

Unlock All Jobs Now

Maybe later