Senior Site Reliability Engineer @COGNATIV
Devops
Salary unspecified
Remote Location
Employment Type full-time
Posted 2mths ago

[Hiring] Senior Site Reliability Engineer @COGNATIV

2mths ago - COGNATIV is hiring a remote Senior Site Reliability Engineer. πŸ’Έ Salary: unspecified πŸ“Location: Worldwide

Role Description

We run a distributed, camera-based video monitoring and AI alerting platform. The system spans the full spectrum of modern and legacy infrastructure: an AWS-hosted fleet of Java microservices and workers, a GPU-backed computer-vision inference pipeline, a real-time streaming and presence layer, and thousands of on-premise edge "media boxes" deployed in the field that ingest camera feeds, serve HLS video, and stream events back to the cloud.

This is a reliability-first role. Your primary job is to keep a large, mixed operational estate healthy at scale:

  • Meaningful service objectives
  • Trustworthy alerting
  • Sound capacity
  • Tested disaster recovery
  • Fast, calm incident response

If you think in SLOs, error budgets, and blameless postmortems, and you are happiest when a noisy, fragile system becomes quiet and predictable on your watch, this is your role.

What you'll keep reliable

  • Computer-vision / AI models: frame-based inference services running on GPU EC2 (g4dn-class, AWS Deep Learning AMIs), fed by camera frames from S3 and a Kafka (Amazon MSK) event bus, with Redis (ElastiCache) for state.
  • Python services: the AI/alerts inference tier and supporting tooling.
  • Legacy Java services: ~140 Java 8 services and libraries (REST APIs, SQS/SNS workers, Lambda functions) running on Jetty 9.4, deployed to Elastic Beanstalk, ECS, and Lambda.
  • Edge appliances ("media boxes"): Ubuntu 22.04 / Docker Compose appliances managed remotely over AWS IoT Core secure tunnelling, with Cloudflare tunnels for egress.
  • Data and messaging backbone: PostgreSQL (RDS) across many schemas, TimescaleDB for analytics, Redis, DynamoDB, Amazon MSK (Kafka), SQS/SNS, and Kinesis.

What you'll do Reliability and operations (the core of the role)

  • Own service objectives. Define SLIs and SLOs for the services that matter, manage error budgets, and use them to drive prioritization and change-rate decisions.
  • Make observability trustworthy. Own alert quality end-to-end; build the dashboards and the custom metrics, exporters, and instrumentation (CloudWatch, OpenTelemetry) needed to see the system clearly.
  • Lead incident response. Run incidents calmly, drive mean-time-to-recovery down, and produce blameless postmortems with action items that actually get closed.
  • Plan capacity and performance. Forecast and right-size compute (especially GPU), Kafka/MSK throughput and partitioning, RDS/TimescaleDB load, and Redis.
  • Own business continuity and disaster recovery. Backups, replication, failover, and recovery for RDS, MSK, Redis, and the edge fleet.
  • Keep the edge fleet healthy. Remote diagnosis and recovery over AWS IoT, container auto-update over systemd timers.
  • Engineer away toil. Write real software (Python, Golang, Bash) to automate operational work.
  • Govern production change safely. Enforce collaborative, reviewed change management.

Delivery and platform (in support of reliability)

  • Keep CI/CD healthy and safe: CircleCI with Bazel/Gradle builds, OIDC-based AWS auth, container builds to ECR, and EB/ECS/Lambda deploys.
  • Maintain Terraform for the AWS estate (compute, networking, IAM, databases, messaging, monitoring).
  • Harden security and compliance: IAM least-privilege, Secrets Manager/KMS, TLS and certificate management.

Qualifications

  • AWS certification is mandatory. A current AWS Certified DevOps Engineer – Professional or AWS Certified Solutions Architect – Professional is strongly preferred.
  • 10+ years in Site Reliability Engineering or production operations at scale.
  • Demonstrated SLO/error-budget practice.
  • Strong production observability skills.
  • Proven incident command.
  • Capacity planning and performance experience across compute, databases, and a messaging or streaming system (Kafka/MSK ideal).
  • Disaster recovery ownership.
  • Software engineering ability for automation.
  • Expert with Terraform (or equivalent IaC) and strong Linux administration.
  • Database operations experience with PostgreSQL.
  • A reliability mindset.

Requirements

  • Operating GPU workloads and serving computer-vision or ML models in production.
  • Apache MSK / Kafka and streaming-data operations.
  • AWS IoT Core at scale.
  • Managing a fleet of edge / on-premise devices.
  • Operating and modernizing legacy systems.
  • Chaos engineering / game-day practice, and capacity modeling.
  • Familiarity with Bazel in a monorepo; Cloudflare, Cognito/Auth0, API Gateway.

How we work, and what we expect from this hire:

  • Observability must be trustworthy.
  • Change management is collaborative, never unilateral.
  • Follow-through over activity.
  • Prefer the right fix to a quick patch.
Before You Apply
️
worldwide Be aware of the location restriction for this remote position: Worldwide
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Senior Site Reliability Engineer @COGNATIV
Devops
Salary unspecified
Remote Location
Employment Type full-time
Posted 2mths ago
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
️
worldwide Be aware of the location restriction for this remote position: Worldwide
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
Γ—
Apply to the best remote jobs
before everyone else

Access 125,000+ vetted remote jobs and get daily alerts.

4.9 β˜…β˜…β˜…β˜…β˜… from 500+ reviews

⚑ 126,783+ remote jobs, refreshed hourly

πŸ”” Real-time alerts: Apply first, direct to employer

πŸ›‘οΈ Vetted companies, no scams, true remote only

Unlock All Jobs Now

Maybe later