Forward Deployed Engineer - SRE @Andromeda Cluster
Devops
Salary unspecified
Remote Location
Employment Type full-time
Posted 2mths ago

[Hiring] Forward Deployed Engineer - SRE @Andromeda Cluster

2mths ago - Andromeda Cluster is hiring a remote Forward Deployed Engineer - SRE. πŸ’Έ Salary: unspecified πŸ“Location: USA, Canada

Role Description

This is not a generalist SRE role, and it is not a support role. You will embed directly with the teams running large-scale training and inference on our clusters. You are responsible for onboarding them, tuning their jobs, and debugging their failures alongside them, while owning the infrastructure and automation that makes those clusters reliable in the first place.

Forward deployed means you spend real time inside customer environments: reading their training code, sitting in their Slack channels, watching their runs, and shipping fixes that land in our platform. When a multi-hundred-GPU run stalls, you are the person who figures out whether it's the fabric, the driver, the scheduler, or their dataloader, then you make sure it can't happen the same way twice.

We're looking for engineers who have personally run GPU clusters in production, understand the failure modes of distributed training, and can reason about performance from network fabric β†’ kernel β†’ framework. Equally important: you can explain what you found to someone else's engineering team without condescension, and turn that conversation into a product improvement.

What You’ll Do

  • Serve as the primary technical point of contact for teams running large-scale training and inference workloads.
  • Own onboarding end to end; environment setup, orchestration choice (Slurm, Kubernetes, or direct SSH), storage layout, first successful run at scale.
  • Continue to stay engaged as their workloads grow.
  • Work inside customer environments to diagnose real failures: NCCL timeouts, stragglers, checkpoint I/O stalls, degraded links, OOM patterns, container and driver mismatches.
  • Read their code when you need to. Reproduce, isolate, fix, and write it down.
  • Profile and improve distributed training performance on live workloads.
  • Own reliability outcomes for the accounts you're deployed on.
  • Ensure the health and performance of high-speed interconnects (InfiniBand, RoCE, NVLink) that underpin distributed training.
  • Build deep visibility into GPU utilization, memory pressure, interconnect throughput, job performance, and hardware health.
  • Turn every repeated deployment problem into automation: cluster provisioning, GPU health checks and burn-in, preflight validation, self-healing, firmware/driver lifecycle management, and reusable reference configurations for common training and serving stacks.
  • Lead incident response for complex, multi-layer failures spanning hardware, networking, orchestration, and ML frameworks.
  • Own the customer-facing communication during the incident and the blameless postmortem and systemic fix after it.
  • Bring rough edges back to influence the roadmap, file hard bugs, and build missing pieces when that's the fastest path.

What We’re Looking For

  • Hands-on experience operating GPU clusters in production (NVIDIA A100/H100/H200/B200 or equivalent).
  • Production experience with InfiniBand, RoCE, or NVLink fabrics in the context of distributed training.
  • Working knowledge of how large training and inference jobs actually run.
  • Expert-level Linux experience: kernel tuning, driver management (NVIDIA drivers, CUDA toolkit), cgroup/namespace internals, container runtimes, and performance profiling at the syscall and hardware level.
  • Strong experience running Kubernetes in production with GPU workloads.
  • Strong engineering skills in Python, Go, or Bash.
  • Infrastructure-as-Code proficiency (Terraform, Helm, Ansible, or equivalent).
  • Hands-on experience building monitoring and alerting for GPU-specific telemetry (DCGM, nvidia-smi, fabric manager metrics).
  • Ability to go deep on architecture with a customer's infra team and clearly articulate tradeoffs to their leadership.
  • Proven track record leading incident response for complex distributed systems.

Strong Candidates May Have

  • Experience with high-performance parallel file systems (VAST, WEKA, Lustre, GPFS) and the checkpoint I/O and data-loading bottlenecks that come with large training runs.
  • Time spent embedded with external engineering teams, i.e., solutions architecture, professional services, deployed SRE, or technical account ownership at an infrastructure company.
  • Experience operating production inference.
  • Contributions to relevant OSS projects, or benchmarks, postmortems, and deep-dives you've published.
  • Experience working across heterogeneous providers and regions rather than a single hyperscaler.

What Success Looks Like

  • You know each of your accounts' actual technical goals including what they're training, what their scaling roadmap looks like over the next two quarters, what their real constraints are (budget, deadline, headcount, data).
  • You are the person your accounts' engineers message first, before they file a ticket.
  • Their reliability and throughput numbers are visibly better than at onboarding, and you can point to the specific changes that did it.
  • Recurring problems you found in the field exist as automation, preflight checks, or documentation.
  • You've advocated internally for at least one roadmap change on behalf of a strategic customer, and it shipped.

Why You’ll Love It Here

  • High-growth environment: Get in early at a company at the center of the AI infrastructure boom.
  • Ownership: First FDE for the solutions engineering team, you’ll get to build this function from the ground up.
  • Competitive compensation: + meaningful equity.
  • Comprehensive benefits: for you and your dependents, including healthcare, dental, and vision coverage, 401(k), and unlimited PTO.
Before You Apply
️
remote Be aware of the location restriction for this remote position: USA, Canada
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Forward Deployed Engineer - SRE @Andromeda Cluster
Devops
Salary unspecified
Remote Location
Employment Type full-time
Posted 2mths ago
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
️
remote Be aware of the location restriction for this remote position: USA, Canada
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
Γ—
Apply to the best remote jobs
before everyone else

Access 125,000+ vetted remote jobs and get daily alerts.

4.9 β˜…β˜…β˜…β˜…β˜… from 500+ reviews

⚑ 127,038+ remote jobs, refreshed hourly

πŸ”” Real-time alerts: Apply first, direct to employer

πŸ›‘οΈ Vetted companies, no scams, true remote only

Unlock All Jobs Now

Maybe later