Senior Site Reliability Engineer @Software Mind
Devops
Salary unspecified
Remote Location
Employment Type full-time
Posted 1wk ago

[Hiring] Senior Site Reliability Engineer @Software Mind

1wk ago - Software Mind is hiring a remote Senior Site Reliability Engineer. πŸ’Έ Salary: unspecified πŸ“Location: Worldwide

Role Description

We are the AI Experience Framework team that builds the platform powering ServiceNow's AI-first user interfaces - an SSR runtime (karuna) built on Lit and server-rendered web components, running behind a multi-tier proxy/HTTP2 routing chain with sharded V8 isolate pools, paired with a ServiceNow Glide/Java platform layer (karuna-glide) that supplies metadata, ACLs, and service artifacts. This role owns production reliability for that stack end to end: Kubernetes deployment and operations, observability, and hands-on troubleshooting of both the Node.js and JVM sides of the system - not generalist infrastructure work.

Qualifications

  • Production operations/SRE experience, including hands-on Kubernetes deployment, scaling, and incident response
  • Direct operational experience troubleshooting Node.js in production: reading heap snapshots and CPU profiles, diagnosing event-loop blocking, and understanding process/worker isolation models (V8 isolates or equivalent sandboxing)
  • Direct operational experience troubleshooting JVM-based services in production: GC log analysis, thread dump analysis, and JVM tuning
  • Hands-on experience building and maintaining Prometheus alerting rules and Grafana dashboards from scratch, not just consuming existing ones
  • Strong Linux/networking fundamentals: DNS, load balancing, TCP/HTTP semantics (including HTTP/2), and debugging service-to-service networking inside Kubernetes
  • Experience with CI/CD and infrastructure-as-code for containerized deployments (Helm, GitOps tooling such as ArgoCD/Flux, or equivalent)
  • Track record owning on-call rotations, writing runbooks, and driving postmortems that lead to real reliability improvements
  • Experience with Splunk for log aggregation, search, and production troubleshooting
  • Hands-on experience with in-memory caching systems (Valkey/Redis) in production β€” key design, TTL/eviction tuning, and tenant-scoped cache invalidation
  • Experience managing service-to-service mTLS - certificate issuance, rotation, and format conversion (e.g., PKCS/BCFKS↔PEM) - plus JWT-based service authentication
  • Very good spoken and written English

Requirements

  • Own Kubernetes deployment and operational health for Framework services, including scaling, rollout/rollback strategy, and resource tuning
  • Build and maintain production observability - Grafana dashboards and Prometheus alerting rules - across the SSR runtime and the Glide platform layer
  • Diagnose and resolve Node.js production incidents: event-loop stalls, heap growth, V8 isolate exhaustion, and isolate-pool scheduling issues under concurrent versioned traffic (vN/vN-1)
  • Diagnose and resolve JVM production incidents on the Glide/Java side: GC pressure, thread dumps, and platform-service latency
  • Own incident response for the team: runbooks, on-call rotation, postmortems, and paging hygiene
  • Drive CI/CD and infrastructure-as-code for Kubernetes manifests/Helm and deployment pipelines
  • Partner with the framework engineering team to identify reliability gaps before they become incidents - capacity planning, load testing, chaos/failure-injection where useful
  • Represent production reliability concerns in architecture reviews for new framework capabilities

Benefits

  • Flexible employment and remote work
  • International projects with leading global clients
  • International business trips
  • Non-corporate atmosphere
  • Language classes
  • Internal & external training
  • Private healthcare and insurance
  • Multisport card
  • Well-being initiatives
Before You Apply
️
worldwide Be aware of the location restriction for this remote position: Worldwide
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Senior Site Reliability Engineer @Software Mind
Devops
Salary unspecified
Remote Location
Employment Type full-time
Posted 1wk ago
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
️
worldwide Be aware of the location restriction for this remote position: Worldwide
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
Γ—

Apply to the best remote jobs
before everyone else

Access 125,000+ vetted remote jobs and get daily alerts.

4.9 β˜…β˜…β˜…β˜…β˜… from 500+ reviews
Unlock All Jobs Now

Maybe later