Role Description
As a Staff Site Reliability Engineer, you will join our Government SRE team and own both the technical reliability of government environments and the coordination of compliant, efficient deployments into them. This team is at the critical intersection of the unique compliance requirements of a regulated environment and the need to establish a consistent software experience for users and developers in commercial environments. You will work closely with cross-functional teams, including security, compliance, operations, validation, and engineering, to lead best practices for cloud infrastructure, continuous delivery, and government release processes.
What Will You Do?
-
Drive continuous software delivery, resolve incidents, run post mortems, and create automation strategies for deployment, self-testing, and alerting.
-
Lead and execute incident management for production issues, ensuring rapid recovery, root cause analysis, and preventative follow-up actions.
-
Improve and optimize the observability strategy by collaborating with application engineering teams to design monitoring solutions that enhance alerting capabilities and reduce noise.
-
Define, implement, and monitor SLOs, SLIs, and SLAs in collaboration with product and engineering teams to align with business objectives.
-
Design, develop, and maintain software solutions that address operational, compliance, and pipeline challenges.
-
Own and coordinate all government environment releases, driving process improvements to enhance the release pipeline's efficiency, reliability, and visibility.
-
Understand product architecture and service dependencies to manage risk and implement effective testing strategies.
-
Partner cross-functionally with engineering, product, SecOps, compliance, and leadership teams to align priorities, define testing strategies, and resolve challenges.
-
Ensure all infrastructure and deployments meet FedRAMP, government regulations, and industry standards, while maintaining required release documentation and risk assessments.
Qualifications
-
8+ years of experience in SRE, DevOps, or Infrastructure Engineering for SaaS products, with 4+ years running operations at a large scale.
-
2+ years of production experience with a container orchestration system (Kubernetes preferred) and Continuous Delivery.
-
Strong understanding of compliance frameworks relevant to government deployments (e.g., FedRAMP, DoD, NIST 800 53, NIST 800 137).
-
Multi cloud experience in AWS/GCP (expertise within AWS preferred).
-
Demonstrated experience with at least one main programming language (Python, Go, Ruby, etc.) and proficiency in bash scripting to improve operational workflows.
-
Familiarity with GitOps frameworks, IaC tooling (Terraform or Pulumi), and deployment strategies (blue green, rolling deploys, canary deploys).
-
Experience with industry standard observability stacks (Prometheus, Grafana, ELK, OpenTelemetry, etc.) and incident management processes.
-
Proven background implementing and supporting FedRAMP, security, risk management, and compliance processes for software releases.
-
Experience working directly with government agencies or in highly regulated industries.
-
Familiarity with testing strategies and automation in large-scale environments.
Benefits
-
Equity & Rewards
-
Restricted Stock Units (RSUs)
-
Employee Stock Purchase Plan (ESPP)
-
Flexible time off
-
Paid company holidays and paid sick time
-
Gender-neutral parental leave
-
Grandparent leave
-
Medical, dental, and vision coverage
-
401(k) retirement plan with company match
-
Life and disability insurance
-
Health and dependent care FSA
-
Voluntary benefits (hospital, accident, critical illness)
-
Employee Assistance Program (EAP)
-
ARAG pre-paid legal
-
Nationwide pet insurance
-
Cancer Care program
-
Global business travel medical insurance
-
Home office allowance
-
Mobile phone reimbursement
-
Wellness coach
-
Wellness/gym reimbursement
-
Fertility coverage
-
Adoption & surrogacy reimbursement