Principal ML Performance Engineer @Proxima
Artificial Intelligence
Salary unspecified
Remote Location
Employment Type full-time
Posted 4d ago

[Hiring] Principal ML Performance Engineer @Proxima

4d ago - Proxima is hiring a remote Principal ML Performance Engineer. πŸ’Έ Salary: unspecified πŸ“Location: Northern America, Europe

Role Description

Profile and optimize training and inference for structural and generative models, including:

  • Transformers
  • Diffusion
  • Geometric deep learning

Write and tune custom kernels (CUDA, Triton) and use compilers (torch.compile, TensorRT, XLA) when beneficial.

Scale distributed training across 32-64 nodes, employing:

  • FSDP
  • DeepSpeed
  • Tensor and pipeline parallelism
  • Mixed precision

Reduce inference cost by optimizing:

  • Memory scaling for large complexes
  • Diffusion sampling efficiency
  • Batching ragged inputs
  • Maximizing throughput across up to 1000 GPUs

Manage GPU cluster efficiency on GCP, focusing on:

  • Scheduling
  • Utilization
  • Spot strategy
  • Cost reporting

Develop benchmarks and profiling tools for the research team.

Qualifications

  • Minimum of 6+ years experience in ML systems, HPC, or performance engineering
  • BS/MS/PhD in CS, EE, or related field
  • Demonstrated ability to set technical direction beyond coding
  • Deep knowledge of PyTorch internals with hands-on experience profiling and fixing real bottlenecks
  • Experience with CUDA and Triton, skilled at reading Nsight output
  • Strong understanding of memory bandwidth and occupancy
  • Experience with distributed training at multi-node scale
  • Strong proficiency in Python and C++
  • Able to name a model they made materially faster and quantify the improvement

Requirements

  • Experience in geometric deep learning, equivariant networks, or protein structure models such as AlphaFold, ESM, or RFdiffusion
  • Experience writing kernels for structure-model primitives, including triangle attention, triangle multiplicative updates, cuEquivariance, or FlashAttention for pair bias
  • Experience orchestrating large batch inference and managing Kubernetes GPU scheduling
Before You Apply
️
remote Be aware of the location restriction for this remote position: Northern America, Europe
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Principal ML Performance Engineer @Proxima
Artificial Intelligence
Salary unspecified
Remote Location
Employment Type full-time
Posted 4d ago
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 120,000+ Remote Jobs
️
remote Be aware of the location restriction for this remote position: Northern America, Europe
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 120,000+ Remote Jobs
Γ—
Apply to the best remote jobs
before everyone else

Access 120,000+ vetted remote jobs and get daily alerts.

4.9 β˜…β˜…β˜…β˜…β˜… from 500+ reviews

⚑ 124,006+ remote jobs, refreshed hourly

πŸ”” Real-time alerts: Apply first

πŸ›‘οΈ Vetted companies, no scams, true remote only

Unlock All Jobs Now

Maybe later