Senior Software Engineer, GPU Cluster Infrastructure @FAR.AI
Software Development
Salary usd 150,000 - 2..
Remote Location
Employment Type full-time
Posted 3wks ago

[Hiring] Senior Software Engineer, GPU Cluster Infrastructure @FAR.AI

3wks ago - FAR.AI is hiring a remote Senior Software Engineer, GPU Cluster Infrastructure. πŸ’Έ Salary: usd 150,000 - 275,000 per year πŸ“Location: USA, Singapore

Role Description

You'd work across the whole infrastructure stack, from scheduling to storage to monitoring to security, and bring real depth in at least one part of it. We're particularly interested in experience with:

  • Large-scale pre-training and post-training infrastructure
  • Cluster security and sandboxing
  • Distributed storage systems
  • Batch scheduling for large GPU clusters

In frontier AI research, working out the infrastructure is often part of the science. You'd work directly with researchers and other engineers to keep our large-scale experiments performant and fault-tolerant.

We're also hiring a Tech Lead Manager, GPU Cluster Infrastructure for this team. If leading a small team while staying hands-on sounds like you, take a look there instead.

What you'll do

  • Operate the Kubernetes GPU fleet day to day, handling node lifecycle, upgrades, driver and image rollouts, staged changes with safe rollback, and capacity planning.
  • Own batch scheduling and multi-tenancy, including queues, quotas, priorities, preemption, gang scheduling, and fair share across research teams.
  • Design and run the storage under the fleet, from high-performance shared filesystems for datasets and checkpoints to object storage tiers, quotas, and backups.
  • Keep multi-node training runs fault-tolerant, owning node health and automated draining, debugging NCCL and fabric problems, tracking down stragglers and flaky GPUs, and building the checkpoint and restart patterns.
  • Harden the platform, covering identity and access, network policy, secrets, workload isolation, and sandboxing for the AI agents that run on the cluster.
  • Bring new capacity online, acceptance-testing providers on fabric, NCCL, and storage throughput, holding them to their SLAs, and integrating new clusters into the platform with infrastructure as code.
  • Work directly with research teams on their infrastructure problems and turn the recurring ones into platform fixes. Share the on-call rotation, runbooks, and postmortems.

Qualifications

  • 3+ years in systems or infrastructure engineering on production Linux, running GPU, HPC, or large-scale batch platforms, with ownership of at least one system from design through operation.
  • Experience running production Kubernetes for GPU workloads with a batch layer on top (Slurm, Kueue, Volcano, or similar), including quotas, priority and preemption, and node health.
  • Ownership of infrastructure as code and observability for a production fleet, provisioning with Terraform or Ansible, deploying with Helm and ArgoCD, and monitoring with Prometheus, or their equivalents.
  • Strong programming skills in at least one language commonly used for infrastructure, such as Python, Go, Rust, or C++, with automation and services maintained as shared code.
  • Ability to write clearly for engineers, researchers, and providers, whether it's a design doc, an incident summary, or an escalation.
  • If you meet most of this and not all of it, we encourage you to apply anyway.

Requirements

  • Real depth in one or more of the following areas makes you a compelling candidate:
    • Distributed training infrastructure: multi-node PyTorch and NCCL debugging, the NVIDIA node stack (drivers, GPU Operator, DCGM), InfiniBand or RoCE fabrics, topology-aware placement.
    • Distributed storage: VAST, Weka, Lustre, Ceph, or object storage at scale; checkpoint I/O.
    • Cluster security: admission control, RBAC, node and container hardening, sandboxed runtimes (gVisor, Kata, Firecracker), and isolating autonomous agents on shared infrastructure.
    • Scheduler internals: Kubernetes scheduler plugins or custom controllers, gang scheduling, fair-share and quota, and the utilization, fairness, and latency tradeoffs between them.
    • Multi-provider platforms: scheduling and storage across clusters at different providers so users see one system, including clusters with no shared network and uneven data locality.

Benefits

  • 🩺 Health Insurance - 94% of Insurance premium paid by Organization commencing within 1 month after your start date
  • πŸ’° Retirement - 401(k) plan with up to 2% match
  • 🏝️ PTO - 25 days Paid Time Off per year, accrued weekly and up to 10 days of paid sick leave per year
  • 🚼 Paid Leave - Paid Bereavement, Family, Medical and Pregnancy Disability Leave
  • πŸ–₯️ WFH Stipend & Equipment - Work computer and stipend provided for eligible employees
  • 🍽️ Catered Meals (Berkeley Office Only) - Catered lunches and dinners on workdays at our office

Logistics

  • If based in the USA or Singapore, you will be an employee of FAR.AI (501(c)(3) research non-profit / non-profit CLG). Outside the USA or Singapore, you will be employed via an EOR organisation on behalf of FAR.AI or as a contractor.
  • Location: Both remote and in-person (Berkeley, CA or Singapore) are possible. We sponsor visas for in-person employees, and can hire remotely in most countries. For this role we prefer candidates whose working hours overlap with Berkeley.
  • Hours: Full-time (40 hours/week).
  • On-call: We don't run a formal on-call rotation yet. The team is spread across time zones and covers incidents during working hours. As the experiments we run get larger we expect to introduce one, and this role would take part in it.

If you have any questions about the role, feel free to contact us at [email protected] . Otherwise, if you don't have questions, the best way to ensure a proper review of your skills and qualifications is by applying directly via the application form. Please don't email us to share your resume (it won't have any impact on our decision). Thank you!

Before You Apply
️
remote Be aware of the location restriction for this remote position: USA, Singapore
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Senior Software Engineer, GPU Cluster Infrastructure @FAR.AI
Software Development
Salary usd 150,000 - 2..
Remote Location
Employment Type full-time
Posted 3wks ago
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
️
remote Be aware of the location restriction for this remote position: USA, Singapore
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
Γ—
Apply to the best remote jobs
before everyone else

Access 125,000+ vetted remote jobs and get daily alerts.

4.9 β˜…β˜…β˜…β˜…β˜… from 500+ reviews

⚑ 127,064+ remote jobs, refreshed hourly

πŸ”” Real-time alerts: Apply first, direct to employer

πŸ›‘οΈ Vetted companies, no scams, true remote only

Unlock All Jobs Now

Maybe later