Senior Infrastructure Engineer - GPU Compute @Boundless Networks, Inc.
Engineering
Salary unspecified
Remote Location
Employment Type full-time
Posted 2mths ago

[Hiring] Senior Infrastructure Engineer - GPU Compute @Boundless Networks, Inc.

2mths ago - Boundless Networks, Inc. is hiring a remote Senior Infrastructure Engineer - GPU Compute. πŸ’Έ Salary: unspecified πŸ“Location: Worldwide

Role Description

Boundless is coordinating GPU compute at scale as it becomes a leader in AI. As a Senior Infrastructure Engineer (GPU Compute), you'll build and operate the compute fabric that powers our AI inference workloads β€” a large, heterogeneous, globally distributed GPU fleet spanning consumer cards (including RTX 5090) and datacenter hardware. Your job is to keep that fleet full, fast, cheap, and always on: orchestrating workloads across regions and providers, squeezing every bit of performance out of the hardware, and driving down cost per GPU-hour. This role rewards engineers who want to go deep on bare-metal and GPU optimization.

You should be comfortable operating with a high degree of autonomy, navigating ambiguity, and defaulting to a strong bias for action.

What You'll Do

  • GPU Fleet Orchestration: Operate a heterogeneous, multi-region GPU fleet (consumer + datacenter, including RTX 5090) using tools like SkyPilot, Kubernetes/k3s, and cloud + on-prem providers. Build the patterns that let us schedule inference workloads across the entire fleet reliably.
  • Compute Scheduling & Utilization: Maximize GPU utilization across inference workloads. Own workload placement across spot, on-prem, and cloud capacity, keeping the "always-on inference substrate" saturated and economical.
  • Bare-Metal & GPU Optimization: Go deep on GPU performance β€” PCIe P2P, ReBAR, NUMA topology (e.g. EPYC SP5), CUDA/driver tuning, memory configuration, and network topology β€” to push throughput per node.
  • Reliability, Access & Observability: Build secure fleet access (Tailscale, Teleport), robust observability and alerting, and zero-downtime rollouts across a distributed node fleet.
  • Cost Optimization: Drive down $/GPU-hr through spot instance management, intelligent workload placement between on-prem and cloud, and resource scheduling β€” without sacrificing reliability.

Qualifications

  • 5+ years of infrastructure/DevOps experience operating large-scale production systems
  • Deep expertise in Kubernetes, Docker, and container orchestration at scale
  • Strong Linux systems administration skills
  • Proficiency in infrastructure-as-code tools (Terraform, Ansible, Pulumi)
  • Track record of managing mission-critical, high-throughput systems
  • Strong infrastructure-as-code background in heterogeneous environments
  • Proficiency in at least one common scripting or programming language (Python, Bash, TypeScript, Go, etc.)
  • Comfort navigating ambiguity with a strong bias for action

Nice to Have

  • Experience with GPU computing infrastructure (CUDA, bare-metal optimization, kernel tuning)
  • Experience operating ML training or other large-scale distributed compute infrastructure
  • Experience with GPU fleet orchestration (SkyPilot, Ray, Slurm)
  • Familiarity with fleet access and networking tooling (Tailscale, Teleport)
  • Knowledge of network optimization and topology design
  • Experience with multi-region, globally distributed systems
  • Proficiency in Rust or low-level systems programming
  • Experience with on-premises data center operations

Additional Requirements

  • Candidates must include a public GitHub profile in their application.
  • The GitHub profile should demonstrate a minimum of 1 year of activity/history.
  • Applications that do not include a GitHub profile, or show insufficient activity, will not be considered.

Benefits

  • Competitive salary + equity allocation
  • Health, dental, vision (for U.S. employees; region-adjusted globally)
  • Flexible PTO
  • Professional development and conference travel budget
  • Remote-first with regular off-sites and a high-trust, high-velocity team environment
  • We are a global team, and applicants from around the world are welcome to apply.
Before You Apply
️
worldwide Be aware of the location restriction for this remote position: Worldwide
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Senior Infrastructure Engineer - GPU Compute @Boundless Networks, Inc.
Engineering
Salary unspecified
Remote Location
Employment Type full-time
Posted 2mths ago
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
️
worldwide Be aware of the location restriction for this remote position: Worldwide
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
Γ—
Apply to the best remote jobs
before everyone else

Access 125,000+ vetted remote jobs and get daily alerts.

4.9 β˜…β˜…β˜…β˜…β˜… from 500+ reviews

⚑ 126,939+ remote jobs, refreshed hourly

πŸ”” Real-time alerts: Apply first, direct to employer

πŸ›‘οΈ Vetted companies, no scams, true remote only

Unlock All Jobs Now

Maybe later