Senior Site Reliability Engineer @Fingerprint
Software Development
Salary usd 152,000 - 2..
Remote Location
πŸ‡ΊπŸ‡Έ USA Only
Employment Type full-time
Posted 2wks ago

[Hiring] Senior Site Reliability Engineer @Fingerprint

2wks ago - Fingerprint is hiring a remote Senior Site Reliability Engineer. πŸ’Έ Salary: usd 152,000 - 205,000 per year πŸ“Location: USA

Role Description

Are you a systems-minded engineer who is happiest when production tells you something you didn't expect? Do you care less about how a system looks on a diagram than about how it behaves at 3am under load it wasn't designed for? Do you want to own reliability for a platform that answers millions of identification requests a day, where being wrong or being slow is a customer-visible event? If so, we have the perfect opportunity for you.

We're looking for a Senior Site Reliability Engineer to join our Infrastructure team and take ownership of how our platform behaves in production. This is a hands-on engineering role, not an oversight one β€” you'll write code and infrastructure, own systems end to end, and be measured by whether the things you own stay fast, available, and predictable as we grow.

You'll work across observability, incident response, capacity and performance, change safety, and the tooling that makes all of it routine. You'll define what "reliable" means for the critical paths you own, instrument them so we know before customers do, and partner with product engineering teams to make their services operable by design rather than by heroics.

Responsibilities

  • Own the reliability of core production systems end to end β€” you instrument them, set targets for them, operate them, and are accountable for how they behave under real traffic.
  • Define and maintain SLIs and SLOs for the critical paths you own, wire them into dashboards and alerts, and use error budget burn as the evidence base for what gets fixed next.
  • Drive alert quality: raise signal, kill noise, and close the gap where customers notice a problem before our monitoring does.
  • Take a lead role in incident response β€” investigate systematically across service boundaries, restore service, and write postmortems that produce follow-ups people actually complete.
  • Build secure, resilient, and cost-efficient infrastructure, with explicit attention to failure modes: timeouts and retries, backpressure and load shedding, graceful degradation, and blast radius containment.
  • Do capacity and performance work with real data β€” load testing, profiling, saturation analysis, and headroom planning ahead of growth rather than after an incident.
  • Improve change safety: progressive delivery, automated rollback, meaningful pre-production signal, and deployment practices that make shipping boring.
  • Manage infrastructure through code and configuration (we primarily use Terraform), consistently applying patterns that align with our overall service architecture.
  • Design, write, and ship software and developer-facing tooling that reduces toil and makes operating services straightforward for the engineers who own them.
  • Run deliberate failure testing β€” game days and chaos exercises, staging first β€” to find the gaps and safe limits before customers do.
  • Partner with product engineering teams on production readiness for new and high-risk services: capacity, failure modes, rollback plans, runbooks, and on-call handoff.
  • Participate in the on-call rotation, and improve it: better runbooks, clearer escalation, less pager fatigue for everyone in it.
  • Approach all engineering work with a security lens β€” actively looking for vulnerabilities in your own work and in peer reviews.
  • Act as the go-to person for hard production problems in your area, and mentor engineers through code review, pairing, and design feedback so operational knowledge doesn't silo.

Qualifications

  • 6–10 years of experience in SRE, production engineering, infrastructure, or backend engineering within primarily cloud-based environments (AWS preferred).
  • A track record of owning a system end to end β€” you've designed something significant, shipped it, operated it, and lived with the consequences when it misbehaved.
  • Hands-on experience defining and operating against SLIs, SLOs, and error budgets.
  • Strong incident skills: you've led or been a primary responder on high-severity, customer-facing incidents.
  • Depth in distributed systems failure modes in high-throughput, low-latency environments.
  • Depth in cloud infrastructure fundamentals: networking, load balancing, containerization (EKS/Kubernetes), and distributed systems.
  • Strong hands-on experience managing infrastructure through code and configuration (Terraform or equivalent).
  • Solid programming skills in Go, Python, or a comparable language.
  • Fluency with observability tooling (Datadog, Prometheus, Grafana, OpenTelemetry, or similar).
  • Hands-on experience operating Redis/ElastiCache in production.
  • Fluency with software engineering best practices: source control, code review, comprehensive test coverage across edge cases and errors, and safe deployment.
  • A high level of personal ownership and autonomy, with real experience working without clearly defined requirements.
  • Pragmatism over purity β€” you know reliability competes with delivery.
  • Strong written and verbal communication in English.
  • AI-native by default β€” you use AI tools as a normal part of how you investigate incidents, analyze telemetry, write runbooks, and build tooling.

Compensation & Transparency

For US-based employees, the cash compensation range for this role is $152,000 – $205,000. We set standard ranges for all US roles based on function, level, and geographic location, benchmarked against similar stage growth companies. To comply with local legislation and provide greater transparency, we share salary ranges on all job postings.

Due to regulatory and security reasons, there’s a small number of countries where we cannot have Fingerprint teammates based. Additionally, because Fingerprint is an all-remote company and people can join our workforce from almost any country, we do not sponsor visas. Fingerprint teammates need to be authorized to work from their home location.

We are dedicated to creating an inclusive work environment for everyone. We embrace and celebrate the unique experiences, perspectives and cultural backgrounds that each employee brings to our workplace.

Before You Apply
️
πŸ‡ΊπŸ‡Έ Be aware of the location restriction for this remote position: USA Only
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Senior Site Reliability Engineer @Fingerprint
Software Development
Salary usd 152,000 - 2..
Remote Location
πŸ‡ΊπŸ‡Έ USA Only
Employment Type full-time
Posted 2wks ago
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
️
πŸ‡ΊπŸ‡Έ Be aware of the location restriction for this remote position: USA Only
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
Γ—
Apply to the best remote jobs
before everyone else

Access 125,000+ vetted remote jobs and get daily alerts.

4.9 β˜…β˜…β˜…β˜…β˜… from 500+ reviews

⚑ 126,845+ remote jobs, refreshed hourly

πŸ”” Real-time alerts: Apply first, direct to employer

πŸ›‘οΈ Vetted companies, no scams, true remote only

Unlock All Jobs Now

Maybe later