Role Description
We're seeking an experienced Senior DevOps Engineer to architect, maintain, and scale our critical infrastructure. You'll be responsible for ensuring system reliability, optimizing performance, and implementing modern deployment strategies that enable our engineering teams to ship with confidence. This role requires deep technical expertise across AWS, Kubernetes, and modern DevOps tooling, combined with strategic thinking about capacity planning and system evolution.
-
Architect and scale our AWS and Kubernetes infrastructure as usage and team size grow
-
Design and implement zero-downtime deployment strategies with automated rollback
-
Build and maintain the Terraform modules the rest of engineering relies on for infrastructure changes
-
Lead capacity planning and cost management across production systems
-
Build observability into our services with Honeycomb, Sentry, and CloudWatch so issues surface before they become incidents
-
Manage CI/CD pipelines in GitHub Actions that let engineers ship with confidence
-
Own networking and traffic routing decisions, including load balancing and DNS
Qualifications
-
7+ years in DevOps, Site Reliability Engineering, or Infrastructure Engineering, with deep AWS expertise (EKS, VPC, S3, RDS/Aurora) and fluency in the Well-Architected framework
-
Has run Kubernetes in production, including cluster administration and multi-environment management, and writes Terraform at scale (module development, state management)
-
Has deployed CNCF tooling in production, like Helm and Karpenter, and built CI/CD pipelines in GitHub Actions
-
Has used observability tools (Honeycomb, Sentry, CloudWatch) to diagnose production issues, and implemented zero-downtime deployments with automated rollbacks
-
Strong networking fundamentals (TCP/IP, DNS, load balancing, traffic routing) and experience managing async job processing and message queues at scale
-
Scripts and automates in Python, Go, Bash, or similar, with a track record of capacity planning, performance tuning, and cost management
-
Familiarity authoring AGENTS.md files and DNA scaffolding that encode our conventions and steer AI review bots
Requirements
-
Has designed alerting strategies and on-call runbooks for production services, and built dashboards for real-time service health
-
Background in database monitoring for RDS Postgres, plus experience building custom metrics and instrumentation for application-specific insight
-
Experience profiling applications to find performance bottlenecks, and working with Cloudflare WAF (origin shielding, firewall rules)
-
Familiarity with Okta SSO integration and experience designing RBAC across AWS and Kubernetes
Benefits
-
Healthcare Coverage β Comprehensive health, dental, and vision plans, including two $0/month medical plan options for employees
-
Equity β Ownership in what we're building
-
401(k) plan through Vestwell
-
Flex Benefit β $500/year for home office equipment, productivity tools, learning & development or fitness/wellness
-
Commuter Benefits β $100/month for SF based employees
-
PTO β Flexible paid time off and company holidays
-
Parental Leave β Paid leave to support growing families
-
Wellness & Family Support β Free Talkspace membership, One Medical access (location-dependent), and Kindbody discounts for family planning