Role Description
We are looking for a skilled DevOps Engineer to help build, maintain, and scale our cloud infrastructure and deployment pipelines. In this role, you will be the backbone of our engineering delivery, ensuring that our applications are highly available, secure, and deployable at a moment's notice. You will take ownership of our containerized environments using Docker and Kubernetes (EKS), manage our cloud resources in AWS as code with Terraform, and streamline our CI/CD workflows using GitHub Actions.
Our platform supports healthcare operations, so protecting patient data is part of everything we build. You will work independently while partnering closely with our engineers, supporting both the applications we build and a growing set of self-hosted open-source tools. If you are passionate about automating manual processes, eliminating downtime, and building resilient systems that never depend on a single person, this is the role for you.
Qualifications
-
5+ years of hands-on experience in a DevOps, Site Reliability (SRE) or Cloud Engineering role, including owning production infrastructure.
-
Production experience managing infrastructure with Terraform (Terragrunt, Pulumi or CloudFormation experience also counts).
-
Strong production experience running infrastructure in AWS, including IAM, networking and managed databases.
-
Deep understanding of Docker, and production experience running Kubernetes (EKS preferred), including Helm, networking, upgrades and scaling.
-
A track record of building complex, automated pipelines in GitHub Actions, including secure cloud authentication (OIDC).
-
Solid understanding of cloud networking: DNS, load balancing, VPCs, subnets, security groups and private connectivity.
-
Strong Python or Bash skills for automating operational work.
-
A history of being the primary owner of infrastructure while working closely with application engineers.
Requirements
-
AWS Management: Provision, configure, and maintain scalable cloud infrastructure across various AWS services (e.g., EC2, RDS, S3, VPC, IAM).
-
Infrastructure as Code: Own our Terraform and Terragrunt codebase across AWS, Cloudflare and GitHub.
-
Account and environment structure: Plan and carry out changes to our AWS account structure.
-
Security & Compliance: Enforce security best practices across our cloud environments.
-
Kubernetes Administration: Deploy, manage, upgrade and scale our EKS clusters.
-
Dockerization: Work closely with software engineers to containerize applications.
-
Pipeline Development: Design, build, and maintain robust Continuous Integration and Continuous Deployment (CI/CD) pipelines using GitHub Actions.
-
Release Engineering: Automate testing, staging, and production deployments.
-
Process Automation: Identify bottlenecks in the development lifecycle and write scripts to automate repetitive operational tasks.
-
Observability: Run our monitoring, alerting, logging and tracing.
-
Backups and recovery: Own database backups and disaster recovery.
-
Incident Response: Participate in an on-call rotation to troubleshoot and resolve production issues.
-
Partner with engineering: Work independently day to day, in close collaboration with our engineers.
-
Avoid single points of failure: Keep runbooks and documentation current.
-
Enable self-service: Give engineers safe, scoped ways to deploy and operate their own services.
Preferred Skills
-
Experience running infrastructure that handles PHI in a HIPAA-regulated environment.
-
Experience running third-party open-source apps in production.
-
Experience with AWS Organizations, or with moving workloads and data between AWS accounts.
Bonus Skills
-
Experience preparing for or supporting a SOC 2 audit.
-
Experience with the Grafana stack (Loki, Tempo, Mimir), Prometheus, OpenTelemetry or similar.