Role Description
Weβre seeking a Sr. Site Reliability Engineer (SRE) to help lead the transformation of our infrastructure to support a more modern, containerized, and highly scalable architecture. This role will play a critical part in evolving our platform from legacy Azure-based services toward a Kubernetes-driven, microservices-oriented environment.
-
Take ownership of complex, ambiguous infrastructure challenges and drive them through to practical, scalable solutions.
-
Partner closely with engineering and data teams to ensure reliability, performance, and scalability are built into our systems from the ground up.
Qualifications
-
5+ years of experience in Site Reliability Engineering, DevOps, or Infrastructure Engineering.
-
Hands-on experience working in cloud environments (Azure preferred, AWS or GCP acceptable).
-
Strong experience with Kubernetes in production environments.
-
Experience deploying and managing applications on Kubernetes using Helm.
-
Proficiency with Infrastructure-as-Code tools such as Terraform.
-
Strong scripting skills (Python, Bash, or similar) used to automate and solve infrastructure challenges.
-
Experience with observability, monitoring, and incident response in production environments.
-
Experience building or supporting CI/CD pipelines (GitHub Actions and/or GitLab CI/CD; experience with CI/CD migrations is a plus).
-
Solid understanding of networking fundamentals, system design, and cloud infrastructure components.
-
Familiarity with Azure Entra ID, app registrations, federated identity credentials, and workload identity.
-
Proven ability to take loosely defined problems and drive them to practical, scalable solutions.
-
Strong communication skills and ability to collaborate across engineering and non-technical stakeholders.
-
Comfort operating in fast-paced environments with evolving priorities.
-
Understanding of PHI handling requirements, access control patterns, and audit controls in a healthcare environment.
Requirements
-
Contribute to the migration from legacy Azure services and function-based architectures to containerized, microservices-based systems.
-
Help design, build, and scale Kubernetes-based infrastructure and supporting tooling.
-
Partner with engineering teams to ensure new systems are designed for reliability, scalability, and operational efficiency from day one.
-
Drive standardization across infrastructure to reduce silos and enable broader team ownership.
-
Monitor cloud usage and spending, identify inefficiencies, and recommend and implement cost optimization strategies.
-
Design, deploy, and manage secure networking, including public and private endpoints, environment segmentation, site-to-site and point-to-site VPNs, and inter-environment connectivity.
-
Design and maintain highly available, resilient systems in a cloud-native environment.
-
Implement and evolve observability practices including monitoring, alerting, and logging (e.g., Datadog, Prometheus, Grafana).
-
Define and manage SLIs, SLOs, and SLAs aligned to system performance and user experience.
-
Lead incident response efforts and drive root cause analysis and long-term improvements.
-
Build and optimize CI/CD pipelines to support fast, safe, and repeatable deployments.
-
Champion Infrastructure-as-Code practices using tools such as Terraform to eliminate manual processes.
-
Leverage scripting (Python, Bash, or similar) to solve problems, automate workflows, and reduce operational toil.
-
Drive capacity planning, performance tuning, and infrastructure improvements to support rapid growth.
-
Proactively identify system risks and scalability bottlenecks before they impact customers.
-
Contribute to infrastructure strategy and help shape how the platform evolves as the business scales.
-
Document systems, processes, and best practices to improve team-wide reliability and reduce single points of failure.
-
Contribute to cross-training efforts as the team moves toward broader ownership and standardization.
-
Share knowledge and elevate the team through mentorship and collaboration.
Benefits
-
Professional growth opportunities with compelling career paths.
-
Healthy work-life balance supported by flexible paid time off (PTO).
-
Comprehensive benefits package, including medical, dental, vision, STD & LTD insurance for full-time team members.
-
401(k) savings plan with employer matching contributions.