Senior Site Reliability Engineer @Cross River
Devops
Salary $160,000 - $200..
Remote Location
πŸ‡ΊπŸ‡Έ USA Only
Employment Type full-time
Posted 1mth ago

[Hiring] Senior Site Reliability Engineer @Cross River

1mth ago - Cross River is hiring a remote Senior Site Reliability Engineer. πŸ’Έ Salary: $160,000 - $200,000 per year πŸ“Location: USA

Role Description

We are seeking a highly skilled and motivated Senior Site Reliability Engineer with 8+ years of hands-on experience ensuring the reliability, scalability, and performance of mission-critical systems. The ideal candidate brings deep expertise in building and maintaining production infrastructure, establishing DevOps best practices, and driving operational excellence across engineering teams. We're looking for someone who takes ownership of system reliability, thrives in a collaborative and fast-paced environment, and is passionate about building resilient financial infrastructure.

  • Define and enforce DevOps guardrails, standards, and best practices to ensure consistency, security, and compliance across engineering teams.
  • Enable Engineering teams to design, implement, and maintain CI/CD pipelines best to enable fast, safe, and repeatable deployments across all environments.
  • Co-Develop and maintain Infrastructure as Code (IaC) with Application teams using tools such as Terraform.
  • Establish and govern deployment strategies including blue/green, canary, and rolling deployments.
  • Build and maintain developer self-service tooling and internal platforms that accelerate delivery while maintaining governance.
  • Champion a "shift-left" culture by embedding reliability, security, and observability practices early in the software development lifecycle.
  • Help define, implement, and monitor Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets for critical services.
  • Build and maintain comprehensive observability stacks including centralized logging, metrics, distributed tracing, and alerting using tools such as New Relic, ELK, Prometheus, and Grafana.
  • Lead incident response and management, including on-call rotations, root cause analysis (RCA), and blameless post-mortems.
  • Perform capacity planning and performance engineering to ensure systems scale efficiently with business growth.
  • Identify and eliminate toil through automation, reducing manual operational overhead.
  • Conduct reliability reviews and chaos engineering exercises to proactively identify and mitigate failure modes.
  • Manage and optimize cloud infrastructure to balance reliability, cost, and performance.
  • Collaborate with software engineering teams to improve system architecture, resiliency patterns, and fault tolerance.
  • Support and maintain .NET-based services running across Windows and Linux environments.

Qualifications

  • 8+ years in SRE, DevOps, or Infrastructure Engineering roles.
  • 5+ years with AWS (preferred); experience with multi-cloud is a plus.
  • Strong proficiency with Terraform.
  • Deep experience designing, building, and maintaining CI/CD pipelines and automation workflows.
  • Strong experience with Docker and container orchestration (ECS preferred).
  • Proficiency with tools such as New Relic, ELK Stack, CloudWatch, Prometheus, Datadog, or Grafana.
  • Proficiency in .NET, PowerShell, Python, Go, or Bash.
  • Strong Linux and Windows systems administration skills.
  • Solid understanding of DNS, load balancing, CDNs, and network security.
  • Strong written and verbal communication skills.

Requirements

  • Experience implementing and managing service mesh technologies (e.g., Istio, Linkerd, AWS App Mesh).
  • Familiarity with SRE frameworks as outlined in Google's SRE handbook.
  • Experience with secrets management (e.g., HashiCorp Vault, AWS Secrets Manager).
  • Understanding of compliance and regulatory requirements in financial services (SOC 2, PCI-DSS, etc.).
  • Experience with chaos engineering tools (e.g., Gremlin, Litmus, AWS Fault Injection Simulator).
  • Experience supporting .NET applications in production environments.
  • Financial industry / banking infrastructure experience is helpful, but not required.
  • Crypto / blockchain infrastructure experience is helpful, but not required.
  • Experience with GitOps workflows and patterns.
  • Familiarity with cost optimization and FinOps practices in cloud environments.

Key Metrics of Success

  • System uptime and availability targets consistently met or exceeded.
  • Reduction in mean time to detect (MTTD) and mean time to resolve (MTTR).
  • Adoption and adherence to DevOps guardrails across engineering teams.
  • Measurable reduction in operational toil through automation.
  • Healthy error budget management across critical services.
Before You Apply
️
πŸ‡ΊπŸ‡Έ Be aware of the location restriction for this remote position: USA Only
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Senior Site Reliability Engineer @Cross River
Devops
Salary $160,000 - $200..
Remote Location
πŸ‡ΊπŸ‡Έ USA Only
Employment Type full-time
Posted 1mth ago
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
️
πŸ‡ΊπŸ‡Έ Be aware of the location restriction for this remote position: USA Only
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
Γ—
Apply to the best remote jobs
before everyone else

Access 125,000+ vetted remote jobs and get daily alerts.

4.9 β˜…β˜…β˜…β˜…β˜… from 500+ reviews

⚑ 126,783+ remote jobs, refreshed hourly

πŸ”” Real-time alerts: Apply first, direct to employer

πŸ›‘οΈ Vetted companies, no scams, true remote only

Unlock All Jobs Now

Maybe later