Member of Technical Staff @Inferact
All Others
Salary $200,000 - $400..
Remote Location
πŸ‡ΊπŸ‡Έ USA Only
Employment Type full-time
Posted 1mth ago

[Hiring] Member of Technical Staff @Inferact

1mth ago - Inferact is hiring a remote Member of Technical Staff. πŸ’Έ Salary: $200,000 - $400,000 per year πŸ“Location: USA

Role Description

We're looking for a hands-on cluster administration engineer to own and operate the high-performance GPU compute infrastructure that keeps Inferact engineering productive. Inferact runs on expensive, high-performance GPU and HPC clusters across neo-cloud and dedicated compute providers. Your job is to make sure that infrastructure is healthy, available, observable, and usable around the clock.

  • Take ownership of cluster health, GPU availability, monitoring, alerting, scheduling, access, diagnostics, and incident response.
  • Work closely with engineering leadership and infrastructure owners to standardize provisioning, operation, debugging, and scaling of compute across providers.
  • Directly impact how fast Inferact can build, test, and improve the systems powering vLLM.

Qualifications

  • Bachelor's degree or equivalent experience in computer science, engineering, systems administration, or similar.
  • Hands-on experience administering large compute clusters, HPC environments, university or research clusters, supercomputing systems, or production GPU clusters.
  • Strong Linux systems administration fundamentals across networking, processes, storage, package management, shell scripting, logs, access control, and system debugging.
  • Experience operating GPU servers, including driver management, GPU health monitoring, node failures, memory errors, scheduler issues, and hardware diagnostics.
  • Experience with cluster scheduling and resource allocation using SLURM, Kubernetes, or equivalent tooling.
  • Ability to own urgent infrastructure incidents end-to-end when compute issues are blocking engineering teams.
  • Ability to automate operational workflows using Bash, Python, Ansible, Terraform, Helm, or similar tooling.

Requirements

  • Experience operating GPU compute across providers such as Lambda, CoreWeave, Crusoe, Nebius, Together, Fireworks, RunPod, or similar environments.
  • Experience improving cluster utilization, reducing idle or unavailable GPU capacity, and debugging scheduling or resource contention issues.
  • Familiarity with high-performance GPU networking such as InfiniBand, RoCE, NVLink / NVSwitch, RDMA, NCCL, or equivalent systems.
  • Experience with storage for HPC or ML workloads, including NFS, Lustre, Ceph, distributed filesystems, or other high-throughput storage systems.
  • Experience managing secure access, identity, permissions, SSH, VPNs, bastion hosts, secrets, and basic infrastructure security hygiene.
  • Background in research computing, scientific computing, ML infrastructure, SRE, platform engineering, or infrastructure operations for engineering-heavy teams.

Bonus Points

  • Managed GPU or HPC infrastructure in a university lab, national lab, research institution, AI infrastructure company, hedge fund, HFT firm, or large-scale ML platform team.
  • Built monitoring, alerting, runbooks, health checks, or remediation workflows that materially reduced operational toil or incident resolution time.
  • Operated Kubernetes clusters for ML or GPU workloads at meaningful scale.
  • Standardized provisioning, diagnostics, monitoring, and operating patterns across multiple compute providers.
  • Carried real operational responsibility for infrastructure used by many engineers or researchers.

Logistics

  • Location: This role is based in San Francisco, California. Will consider remote in the US for exceptional candidates.
  • Compensation: Depending on background, skills, and experience, the expected annual salary range for this position is $200,000 - $400,000 USD + equity.
  • Visa sponsorship: We sponsor visas on a case-by-case basis.

Benefits

  • Generous health, dental, and vision benefits.
  • 401(k) company match.
Before You Apply
️
πŸ‡ΊπŸ‡Έ Be aware of the location restriction for this remote position: USA Only
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Member of Technical Staff @Inferact
All Others
Salary $200,000 - $400..
Remote Location
πŸ‡ΊπŸ‡Έ USA Only
Employment Type full-time
Posted 1mth ago
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
️
πŸ‡ΊπŸ‡Έ Be aware of the location restriction for this remote position: USA Only
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
Γ—
Apply to the best remote jobs
before everyone else

Access 125,000+ vetted remote jobs and get daily alerts.

4.9 β˜…β˜…β˜…β˜…β˜… from 500+ reviews

⚑ 126,955+ remote jobs, refreshed hourly

πŸ”” Real-time alerts: Apply first, direct to employer

πŸ›‘οΈ Vetted companies, no scams, true remote only

Unlock All Jobs Now

Maybe later