HPC Infrastructure Engineer @ElevenLabs
Engineering
Salary unspecified
Remote Location
Employment Type full-time
Posted 1mth ago

[Hiring] HPC Infrastructure Engineer @ElevenLabs

1mth ago - ElevenLabs is hiring a remote HPC Infrastructure Engineer. πŸ’Έ Salary: unspecified πŸ“Location: Northern America, Europe

Role Description

Every model we train runs on infrastructure this role owns. We operate NVIDIA GPU clusters across bare metal and rented capacity, and we're looking for an engineer to join our small research infrastructure team and make that compute fast, reliable, and boring - in the best sense. When the clusters just work, research moves faster. Your impact is measured directly in training throughput and researcher velocity.

This is a builder-operator role with real breadth:

  • One week you're writing automation that eliminates a whole class of manual work.
  • The next you're benchmarking a new provider's InfiniBand fabric or on-site bringing new hardware online.

You'll have unusual scope and autonomy - we're a lean team where decisions are made by the people closest to the problem.

What you’ll be doing:

  • Operate and improve our GPU fleet end to end: provisioning, scheduling, monitoring, upgrades, capacity planning.
  • Build automation that keeps the fleet healthy without human intervention β€” node health checks, automated draining and remediation, burn-in pipelines for new capacity.
  • Own the stack beneath the training code: OS images, NVIDIA drivers, CUDA, container runtimes, NCCL, high-speed networking (InfiniBand/RoCE).
  • Run and tune job scheduling (Slurm or similar) so researchers get compute fairly and fast.
  • Build and maintain high-performance storage for datasets and checkpoints.
  • Hunt down performance problems: stragglers, degraded links, thermal issues, flaky GPUs β€” and fix the class of problem, not just the instance.
  • Evaluate rented GPU capacity: benchmark it, validate it, hold providers to their SLAs.
  • Hands-on hardware work when it's needed: racking, cabling, diagnostics, coordinating with datacenter staff and vendors.
  • Keep clusters secure by default: access control, network isolation, secrets.

Qualifications

  • Have run large-scale Linux server or GPU environments in production and enjoy both building and operating.
  • Know the NVIDIA stack well β€” drivers, CUDA, NCCL, DCGM β€” or have deep systems experience and learn hardware stacks fast.
  • Are comfortable with bare-metal environments, server hardware, and high-speed networking.
  • Write solid automation in Python and/or Bash, with IaC tools like Ansible or Terraform.
  • Are happy digging into noisy data (metrics, logs, PromQL) to find what's actually wrong.
  • Like owning real scope end to end and being the person others rely on.
  • Don't consider any task above or beneath you β€” datacenter trips included.

Requirements

  • Experience supporting ML training workloads from the infra side (distributed training failure modes, checkpointing patterns).
  • Experience evaluating and working with GPU cloud providers.
  • Parallel filesystems (WEKA, VAST, etc) or large-scale object storage.
  • BMC/IPMI/Redfish automation, PXE provisioning at scale.
  • Power and cooling awareness for dense GPU deployments.

Benefits

  • Innovative culture: You’ll be part of a generational opportunity to define the trajectory of AI, surrounded by a team pushing the boundaries of what’s possible.
  • Growth paths: Joining ElevenLabs means joining a dynamic team with countless opportunities to drive impact - beyond your immediate role and responsibilities.
  • Learning & development: ElevenLabs proactively supports professional development through an annual discretionary stipend.
  • Social travel: We also provide an annual discretionary stipend to meet up with colleagues each year, however you choose.
  • Annual company offsite: Each year, we bring the entire team together in a new location - past offsites have included Croatia and Italy.
  • Co-working: If you’re not located near one of our main hubs, we offer a monthly co-working stipend.

Location

This role is remote and can be executed globally. If you prefer, you can work from our offices in London, New York, San Francisco, and Warsaw.

Before You Apply
️
remote Be aware of the location restriction for this remote position: Northern America, Europe
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
HPC Infrastructure Engineer @ElevenLabs
Engineering
Salary unspecified
Remote Location
Employment Type full-time
Posted 1mth ago
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
️
remote Be aware of the location restriction for this remote position: Northern America, Europe
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
Γ—
Apply to the best remote jobs
before everyone else

Access 125,000+ vetted remote jobs and get daily alerts.

4.9 β˜…β˜…β˜…β˜…β˜… from 500+ reviews

⚑ 127,064+ remote jobs, refreshed hourly

πŸ”” Real-time alerts: Apply first, direct to employer

πŸ›‘οΈ Vetted companies, no scams, true remote only

Unlock All Jobs Now

Maybe later