Senior Compute Infrastructure Engineer @StackYak
Engineering
Salary competitive com..
Remote Location
πŸ‡ΊπŸ‡Έ USA Only
Employment Type full-time
Posted YDay

[Hiring] Senior Compute Infrastructure Engineer @StackYak

YDay - StackYak is hiring a remote Senior Compute Infrastructure Engineer. πŸ’Έ Salary: competitive compensation plus meaningful equity. πŸ“Location: USA

Role Description

We need someone who knows what stands between GPU capacity and production infrastructure, and can build it. That capacity does not arrive in one form:

  • Some of it comes as hardware, with everything that implies: firmware, drivers, hosts, and physical work carried out by people you will never meet in buildings you will rarely visit.
  • Some of it comes from neoclouds and other providers, where you control the software and very little else, on terms set by someone whose interests are not yours.

The interesting problem is neither of those on its own. It is the gap between them β€” how much of it can honestly be abstracted away, how much has to stay visible, and how much the rest of the company should ever have to think about. Those are open questions here, and the person in this role is the one who gets to answer them.

You should be comfortable at both ends:

  • On one side: Linux, kernel, firmware, drivers, virtualization, host hardening, and debugging a machine you cannot walk up to.
  • On the other: provider APIs, capacity that is not permanent, and automation at a scale where nothing gets configured by hand.

Workloads will run on bare metal, in VMs, and in containers. The right answer will not be the same one twice, and choosing is part of the job. This is not an architecture-only role. You will design systems and then build, debug, and operate them.

What You Will Own

  • Not tasks. Outcomes, and the authority that comes with them.
  • What the fleet is.
  • What we own versus what we rent, where we standardise and where we deliberately do not, and what it costs us either way.
  • The machines themselves.
  • Linux, kernel, firmware, drivers, CUDA and/or ROCm β€” down to the layer where the answer lives in a changelog rather than documentation.
  • Bare metal, VM, or container.
  • You own the boundary between them, and the uncomfortable fact that the right answer will not be the same one twice.
  • Isolation.
  • Provisioning and recovery.
  • Provider integration.
  • Physical reality.
  • Failures that cross layers.
  • Making it operable by other people.

What Success Looks Like

You can be handed GPU capacity β€” some of it ours, some of it rented, none of it uniform β€” along with a set of business and security requirements, and turn it into production infrastructure the rest of the company can trust and operate.

You will help us answer questions such as:

  • Should this workload run on bare metal, a VM, or a container?
  • How should we provision, rebuild, and recover machines we cannot walk up to?
  • Where does owned hardware genuinely beat rented capacity, and where are we just paying for the privilege?
  • What should customer isolation look like?
  • Where should we standardize, and where should we preserve flexibility?
  • How do we operate owned hardware and external compute as one coherent platform rather than two?
  • Which parts of the infrastructure should become product capabilities rather than internal operations?

What We Need

  • Run production Linux at a depth where the answer was in the kernel, not in the docs.
  • Operated GPU infrastructure for AI, HPC, cloud, or similarly demanding workloads.
  • Carried NVIDIA and/or AMD platforms through a driver or firmware upgrade that did not go smoothly.
  • Provisioned bare metal at a count where configuring machines by hand stopped being an option.
  • Managed machines you could not physically reach, and recovered one anyway.
  • Chosen between bare metal, KVM, and containers with a real consequence attached, and been able to say why.
  • Written Terraform and Python that other people depend on in production β€” tooling, not glue.
  • Hardened hosts and isolated workloads for customers who were not permitted to see each other.
  • Debugged a failure that could have been hardware, firmware, OS, driver, network, or workload, and established which it was.
  • Made an infrastructure decision on bad information and then lived with it long enough to learn whether you were right.

You Will Be Especially Strong If

  • You have built or operated infrastructure at a neocloud, hyperscaler, GPU cloud, HPC environment, hosting company, data-center operator, or AI infrastructure company.
  • You have run a fleet that mixed owned hardware with rented capacity, and have opinions about what that costs you.
  • You have worked with multi-GPU and multi-node systems.
  • You understand NUMA, PCIe topology, RDMA, NIC placement, and why physical topology matters for AI workloads.
  • You have specified or accepted hardware, and know what goes wrong between the purchase order and a machine that boots.
  • You have designed infrastructure meant to serve multiple customers safely.
  • You have done this outside the confines of a tightly siloed enterprise team.
  • You build things on the side because systems are interesting to you, not only because an employer assigned a ticket.

This Is Probably Not For You If

  • Your idea of infrastructure ends at Kubernetes manifests or CI/CD pipelines.
  • You mostly operate through tickets and escalations.
  • You prefer designing systems that someone else implements.
  • You require clearly bounded ownership before touching a problem.
  • You have cloud experience but little understanding of what happens underneath the VM.
  • You know GPU APIs but have never been responsible for the machines underneath them.
  • You want to work only on hardware you own and consider provider-managed capacity beneath you β€” or the reverse, and would rather never think about a physical machine again.
  • You want months to become productive in the core areas of the role.

How We Work

  • Small, senior team with direct access to the founders.
  • Strong opinions, loosely held.
  • Everyone is expected to participate in technical decisions.
  • Everyone shares responsibility for production and on-call.
  • We value people who can move between design, implementation, debugging, and operations.
  • We care much more about what you have built and operated than degrees, certifications, or academic credentials.
  • We expect people to leave ego at the door, argue the technical case, make a decision, and then execute.
  • We are remote and distributed across time zones.
  • Hiring here is a few real conversations with the people you would actually work with, not a recruiter screen followed by a panel of strangers.
  • This is an early-stage startup. The pace is high, the problems are hard, and the scope will change as we grow.

Compensation

Competitive compensation plus meaningful equity. Exact structure will depend on location, engagement model, and experience.

A Note For Agencies

We are not using external recruiters or agencies for this role, and we will not be persuaded otherwise by an email. We do not want your spam. We will not read the CVs you send, we will not reply to your follow-up, and no candidate you put in front of us creates a fee obligation of any kind. Do not contact us.

Before You Apply
️
πŸ‡ΊπŸ‡Έ Be aware of the location restriction for this remote position: USA Only
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Senior Compute Infrastructure Engineer @StackYak
Engineering
Salary competitive com..
Remote Location
πŸ‡ΊπŸ‡Έ USA Only
Employment Type full-time
Posted YDay
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
️
πŸ‡ΊπŸ‡Έ Be aware of the location restriction for this remote position: USA Only
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
Γ—
Apply to the best remote jobs
before everyone else

Access 125,000+ vetted remote jobs and get daily alerts.

4.9 β˜…β˜…β˜…β˜…β˜… from 500+ reviews

⚑ 128,512+ remote jobs, refreshed hourly

πŸ”” Real-time alerts: Apply first, direct to employer

πŸ›‘οΈ Vetted companies, no scams, true remote only

Unlock All Jobs Now

Maybe later