Role Description
We're looking for a Site Reliability Engineer to make Inferact's vLLM-powered inference services dependable in production. You'll work alongside the engineers building the platform, bringing a reliability perspective to how systems are designed, shipped, and operated. This is a hands-on engineering role for someone who thinks about failure before launch, simplifies operations, and writes software that makes reliable inference possible at scale.
-
Help own reliability across the production lifecycle, from release safety and observability to incident response and recovery.
-
Define SLOs around availability, latency, and successful inference requests.
-
Build monitoring that helps engineers pinpoint failures.
-
Automate recurring operational work.
-
Drive mitigation, coordinate escalation, and lead post-mortems during incidents.
-
Partner with engineering on capacity planning and operational readiness.
Qualifications
-
Bachelor's degree or equivalent experience in computer science, engineering, systems, infrastructure, or similar.
-
Hands-on experience building or operating production distributed systems, cloud infrastructure, or services where reliability has meaningful user or business impact.
-
Strong programming or scripting ability in Python, Go, Bash, or similar, with experience building tools and automation that solve operational problems.
-
Experience responding to significant production incidents, including mitigation, root cause analysis, escalation, and following corrective actions through to completion.
-
Practical understanding of service-level objectives (SLOs), service-level indicators (SLIs), error budgets, and alerting that distinguishes user-impacting failures from noise.
-
Strong Linux, networking, and systems debugging fundamentals, with the ability to investigate failures across application and infrastructure boundaries.
-
Clear communication, sound judgment under pressure, and the ability to work with engineering teams to identify failure modes and make systems simpler to operate.
Requirements
-
Experience supporting AI inference, model serving, ML infrastructure, GPU workloads, or other latency-sensitive, high-throughput services.
-
Experience deploying, operating, and debugging Kubernetes-based services, with familiarity with Docker, Terraform, or comparable infrastructure tooling.
-
Experience building observability with metrics, logs, traces, dashboards, and actionable alerts across distributed systems.
-
Experience improving deployment safety and recovery through automation, testing, release gates, and rollback procedures.
-
Experience with resource scheduling, capacity planning, or workload isolation in multi-tenant cloud or GPU environments.
Benefits
-
Generous health, dental, and vision benefits.
-
401(k) company match.