Role Description
As our Senior/Staff DevOps Engineer, you'll decide where our infrastructure goes next and take the rest of engineering there. This is a hands-on role in the truest sense: you set the technical direction, ship it with your own hands, and own the outcomes for reliability, cost, security, and developer experience across the company. You'll partner directly with our product engineering teams, embedded in what they're building, removing infrastructure friction, and shaping infrastructure around real product needs.
Why this role:
-
Real ownership, end to end.
-
Technical leadership without the management overhead.
-
AI-native by default.
Responsibilities
-
Set and drive the technical vision and quarterly roadmap for our infrastructure proactively, with clear trade-offs and measurable goals.
-
Run and evolve our AWS + Kubernetes (EKS) infrastructure: cluster management, autoscaling (Karpenter), policy enforcement (Kyverno), and zero-downtime operations.
-
Own Infrastructure as Code end to end (Terraform, AWS CDK in TypeScript) and our GitLab CI/CD (reusable/shared templates, OIDC, self-managed GitLab).
-
Build and own observability that teams actually use (Prometheus, Grafana, OpenTelemetry, OpenSearch, CloudWatch; log pipelines, APM) a shared view of system health, not a dashboard nobody opens.
-
Keep blue-green deployments and health-gated automated rollback fast and boring; own zero-downtime PostgreSQL schema migrations (expand/contract) and CI migration gating.
-
Own security engineering in a HIPAA environment: secrets hygiene (rotation, short-lived credentials, leak scanning), PHI-aware handling of logs and data, and Vault managed as code.
-
Partner directly with product teams to remove infrastructure friction and improve developer experience by design.
-
Use agentic AI as a core part of your workflow, integrating autonomous-agent output into production.
Qualifications
-
6+ years in DevOps/infrastructure engineering, with strong systems fundamentals and solid Linux administration and troubleshooting (performance analysis, resource management, process debugging).
-
Hands-on AWS (EC2, EKS, RDS, ElastiCache, Lambda, SQS, EventBridge, API Gateway, ALB, S3) and production Kubernetes/EKS (cluster management, node scaling, policy enforcement; Karpenter, Kyverno, or similar).
-
Strong Infrastructure as Code (Terraform and AWS CDK in TypeScript) and CI/CD ownership (GitLab CI/CD: reusable/shared templates, OIDC id_tokens, self-managed GitLab).
-
Monitoring and observability in practice (Prometheus, Grafana, OpenTelemetry, OpenSearch, CloudWatch; log-shipping and error tracking/APM).
-
Practical security engineering (secrets rotation, short-lived credentials, leak scanning, PHI-aware logging) and HashiCorp Vault as code (KV, JWT/OIDC auth for CI, policy design).
-
Blue-green deployments with automated, health-gated rollback; PostgreSQL zero-downtime schema migrations (expand/contract) and migration gating in CI.
-
Containers (Docker, ECR, immutable tags, image lifecycle) and network/protocol fundamentals (load balancing, TLS, DNS).
-
Hands-on agentic AI workflows (Claude Code or similar): delegating to autonomous agents and integrating their output into production.
-
A developer-focused mindset, strong problem-solving for complex system issues, and strong technical writing (docs-as-code, ADRs, design docs via MRs).
-
Fluent Russian and English (B1).
-
Experience working effectively in remote, distributed teams.
Requirements
-
Experience in a regulated/compliance-heavy environment (HIPAA, SOC 2, or similar).
-
Configuration management (Ansible) for VM fleet management.
-
Node.js application operations (pm2, npm), our stack is Node.js + TypeScript.
-
GitOps tooling (ArgoCD, Flux) and deeper PostgreSQL database administration.
-
AWS certifications.
Benefits
-
High-impact environment: A chance to contribute to a product-driven company in the medical tech space.
-
Growth & compensation: Clear growth opportunities and a competitive compensation package.
-
Remote-first: Fully remote long-term collaboration under a B2B model.
-
Health & wellness: Health insurance after the probation period, plus sports & wellness compensation.
-
Career development: Clear growth opportunities, including personalized English lessons via Preply.
-
Time off: 19 paid vacation days annually, 4 additional wellness days each year, and paid sick leave for the first 5 working days.
-
Culture: Thoughtful gifts for key life events and offline corporate events.