Role Description
We run a distributed, camera-based video monitoring and AI alerting platform. The system spans the full spectrum of modern and legacy infrastructure: an AWS-hosted fleet of Java microservices and workers, a GPU-backed computer-vision inference pipeline, a real-time streaming and presence layer, and thousands of on-premise edge "media boxes" deployed in the field that ingest camera feeds, serve HLS video, and stream events back to the cloud.
This is a reliability-first role. Your primary job is to keep a large, mixed operational estate healthy at scale:
-
Meaningful service objectives
-
Trustworthy alerting
-
Sound capacity
-
Tested disaster recovery
-
Fast, calm incident response
If you think in SLOs, error budgets, and blameless postmortems, and you are happiest when a noisy, fragile system becomes quiet and predictable on your watch, this is your role.
What you'll keep reliable
-
Computer-vision / AI models: frame-based inference services running on GPU EC2 (g4dn-class, AWS Deep Learning AMIs), fed by camera frames from S3 and a Kafka (Amazon MSK) event bus, with Redis (ElastiCache) for state.
-
Python services: the AI/alerts inference tier and supporting tooling.
-
Legacy Java services: ~140 Java 8 services and libraries (REST APIs, SQS/SNS workers, Lambda functions) running on Jetty 9.4, deployed to Elastic Beanstalk, ECS, and Lambda.
-
Edge appliances ("media boxes"): Ubuntu 22.04 / Docker Compose appliances managed remotely over AWS IoT Core secure tunnelling, with Cloudflare tunnels for egress.
-
Data and messaging backbone: PostgreSQL (RDS) across many schemas, TimescaleDB for analytics, Redis, DynamoDB, Amazon MSK (Kafka), SQS/SNS, and Kinesis.
What you'll do Reliability and operations (the core of the role)
-
Own service objectives. Define SLIs and SLOs for the services that matter, manage error budgets, and use them to drive prioritization and change-rate decisions.
-
Make observability trustworthy. Own alert quality end-to-end; build the dashboards and the custom metrics, exporters, and instrumentation (CloudWatch, OpenTelemetry) needed to see the system clearly.
-
Lead incident response. Run incidents calmly, drive mean-time-to-recovery down, and produce blameless postmortems with action items that actually get closed.
-
Plan capacity and performance. Forecast and right-size compute (especially GPU), Kafka/MSK throughput and partitioning, RDS/TimescaleDB load, and Redis.
-
Own business continuity and disaster recovery. Backups, replication, failover, and recovery for RDS, MSK, Redis, and the edge fleet.
-
Keep the edge fleet healthy. Remote diagnosis and recovery over AWS IoT, container auto-update over systemd timers.
-
Engineer away toil. Write real software (Python, Golang, Bash) to automate operational work.
-
Govern production change safely. Enforce collaborative, reviewed change management.
Delivery and platform (in support of reliability)
-
Keep CI/CD healthy and safe: CircleCI with Bazel/Gradle builds, OIDC-based AWS auth, container builds to ECR, and EB/ECS/Lambda deploys.
-
Maintain Terraform for the AWS estate (compute, networking, IAM, databases, messaging, monitoring).
-
Harden security and compliance: IAM least-privilege, Secrets Manager/KMS, TLS and certificate management.
Qualifications
-
AWS certification is mandatory. A current AWS Certified DevOps Engineer β Professional or AWS Certified Solutions Architect β Professional is strongly preferred.
-
10+ years in Site Reliability Engineering or production operations at scale.
-
Demonstrated SLO/error-budget practice.
-
Strong production observability skills.
-
Proven incident command.
-
Capacity planning and performance experience across compute, databases, and a messaging or streaming system (Kafka/MSK ideal).
-
Disaster recovery ownership.
-
Software engineering ability for automation.
-
Expert with Terraform (or equivalent IaC) and strong Linux administration.
-
Database operations experience with PostgreSQL.
-
A reliability mindset.
Requirements
-
Operating GPU workloads and serving computer-vision or ML models in production.
-
Apache MSK / Kafka and streaming-data operations.
-
AWS IoT Core at scale.
-
Managing a fleet of edge / on-premise devices.
-
Operating and modernizing legacy systems.
-
Chaos engineering / game-day practice, and capacity modeling.
-
Familiarity with Bazel in a monorepo; Cloudflare, Cognito/Auth0, API Gateway.
How we work, and what we expect from this hire:
-
Observability must be trustworthy.
-
Change management is collaborative, never unilateral.
-
Follow-through over activity.
-
Prefer the right fix to a quick patch.