1mth ago - Mirantis is hiring a remote Senior Site Reliability Engineer. πΈ Salary: unspecified πLocation: USA
Role Description
Define what reliability means for a GPU-accelerated AI platform and make it measurable. You will own the service-level indicators and objectives for the K0rdent Observability Framework (KOF) β deriving meaningful SLIs from the signals the platform already emits, and exposing them to Platform Administrators through a clean API. Work spans hybrid, edge, and air-gapped deployments built on the Mirantis K0rdent stack.
We are looking for a Senior SRE who thinks past dashboards to the contract between a platform and its operators. The right candidate can look at raw telemetry from Kubernetes, bare metal, and NVIDIA infrastructure, decide which signals actually predict user-visible reliability, and turn them into SLIs and SLOs that operators can act on. You are equally comfortable writing the service that exposes those SLIs through an API and reasoning about error budgets, alerting quality, and signal-to-noise. You should be self-directed, able to own reliability definitions end to end, and effectively communicate them across teams.
Responsibilities
Qualifications
Requirements
Benefits
| πΊπΈ | Be aware of the location restriction for this remote position: USA Only |
| βΌ | Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more. | οΈ
| πΊπΈ | Be aware of the location restriction for this remote position: USA Only |
| βΌ | Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more. | οΈ
Access 125,000+ vetted remote jobs and get daily alerts.
β‘ 126,845+ remote jobs, refreshed hourly
π Real-time alerts: Apply first, direct to employer
π‘οΈ Vetted companies, no scams, true remote only