Senior Site Reliability Engineer @Level AI
Devops
Salary unspecified
Remote Location
Employment Type full-time
Posted YDay

[Hiring] Senior Site Reliability Engineer @Level AI

YDay - Level AI is hiring a remote Senior Site Reliability Engineer. πŸ’Έ Salary: unspecified πŸ“Location: India

Role Description

Please note: This role requires working in the EST Time Zone (usually from 8 PM to 4 AM IST) in order to add technical coverage to US customers.

The Senior SRE will be positioned at the intersection of backend engineering, infrastructure operations, and FinOps. The role is explicitly broader than a traditional DevOps engineer and explicitly more hands-on than a pure architect.

  • Infrastructure cost efficiency and FinOps: Own the continued reduction of Kubernetes overprovisioning, drive right-sizing programs, and maintain the cost telemetry that backend teams use to make decisions.
  • GPU throughput optimization: Run a structured experimentation program on on-premise GPU clusters, partnering with AI service owners.
  • Backend enablement, not ownership absorption: Build the tooling, dashboards, and processes that let backend teams from other groups own their own cost and reliability budgets.
  • Reliability instrumentation: Ensure that surface area is captured properly for both cost-at-scale and reliability.
  • Selective security workstreams: Take on a defined slice of the active security work.

Qualifications

  • This role explicitly requires 4-5 years of hands-on systems experience.
  • Backend engineering depth: production experience in Python, Go/Rust, comfortable owning services end to end.
  • Kubernetes at scale: scheduler behaviour, resource requests/limits, HPA/VPA, node pool design, cost-aware autoscaling.
  • Cloud and on-premise infrastructure: GCP fluency, IaC (Terraform), CI/CD, and comfort operating in hybrid setups.
  • GPU workload understanding: familiarity with throughput profiling, batching, KV-cache behavior, inference server tuning, and GPU utilisation metrics.
  • Observability and reliability: metrics, traces, logs, SLOs, and the discipline to instrument systems properly.
  • FinOps mindset: demonstrated history of converting infrastructure choices into measurable cost outcomes.
  • Security baseline: able to take on platform-security workstreams without requiring constant handoff to the DevOps team.
  • Can work in the EST time zone (A must).
Before You Apply
️
remote Be aware of the location restriction for this remote position: India
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Senior Site Reliability Engineer @Level AI
Devops
Salary unspecified
Remote Location
Employment Type full-time
Posted YDay
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 130,000+ Remote Jobs
️
remote Be aware of the location restriction for this remote position: India
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 130,000+ Remote Jobs
Γ—

Apply to the best remote jobs
before everyone else

Access 130,000+ vetted remote jobs and get daily alerts.

4.9 β˜…β˜…β˜…β˜…β˜… from 500+ reviews

⚑ 130,008+ remote jobs, refreshed hourly

πŸ”” Real-time alerts: Apply first

πŸ›‘οΈ Vetted companies, no scams, true remote only

Unlock All Jobs Now

Maybe later