Support Engineer, GPU Infrastructure @Hydra Host
All Others
Salary 95,000 - 130,00..
Remote Location
πŸ‡ΊπŸ‡Έ USA Only
Employment Type full-time
Posted 3wks ago

[Hiring] Support Engineer, GPU Infrastructure @Hydra Host

3wks ago - Hydra Host is hiring a remote Support Engineer, GPU Infrastructure. πŸ’Έ Salary: 95,000 - 130,000 usd per annum πŸ“Location: USA

Role Description

You are the person a customer issue reaches when it is real. You take it from the first symptom to a resolution that holds, working from the Linux host down through the hardware and out to the facility floor.

The tier in the title is deliberate. Nearly everything that arrives at support here is already a Tier 2 or Tier 3 problem, on production hardware, with a customer's workload affected.

Most of the day is diagnosis, coordinating the people with hands on the hardware, and writing down what you found so the next person does not start over. Some of it crosses into infrastructure engineering work.

Support at Hydra is being built right now and you are one of the first hires into it. Expect less structure than you are used to, and more influence over what the structure becomes.

What you will do

  • Diagnose across the stack
  • Work Linux server issues end to end:
    • Boot and network boot failures
    • Kernel and driver problems
    • Filesystems, storage pressure, services, memory and CPU behavior, general instability
  • Diagnose hardware failures using out-of-band management, sensor data, POST and boot errors, SMART data, and vendor diagnostics:
    • IPMI, Redfish, iDRAC, iLO, or equivalent
  • Isolate server-side network problems:
    • NICs and drivers, VLANs, addressing, routing, MTU, DNS, DHCP, bonding, link state, and packet captures
  • Troubleshoot NVIDIA GPU servers:
    • GPU availability, thermal throttling, driver and VBIOS mismatch, PCIe, XID errors
    • Host-level conditions that look like GPU problems and are not
  • Separate hardware from OS from network from application from configuration before escalating
  • Own the incident, not just the ticket
  • Set severity by blast radius and communicate it
  • Notice when several tickets are one problem
  • Escalate to engineering with an evidence pack rather than a description
  • Take part in root cause analysis and post-incident review
  • Work the partner and vendor boundary
  • Drive issues with data center partners
  • Open and track hardware RMAs with OEMs
  • Validate repaired or replaced equipment before it returns to production
  • Keep the asset record accurate
  • Support server turn-ups, migrations and decommissions
  • Communicate with customers
  • Write clear updates to technically sophisticated customers
  • Ask for diagnostic information in a way that never implies the customer has misread their own situation
  • Deliver an unwelcome answer plainly when that is the honest one
  • Document what you learn
  • Write the runbook after you solve something the first time
  • Improve the runbooks and operational procedures you inherit
  • Automation and continuous improvement
  • Build scripts and small tools that take repetitive diagnostic and support work off the queue
  • Use Python, Bash, or similar to automate health checks, data collection, and routine operations
  • Improve the troubleshooting tools and workflows you inherit
  • Convert recurring manual procedures into documented ones, then into automated ones
  • Contribute to infrastructure-as-code and configuration management
  • Help improve monitoring and alerting
  • Work with engineering to find the changes that make the platform easier to operate and cheaper to support at scale

Coverage

  • Support runs across time zones and this role works a set schedule
  • A defined shift, agreed before you start
  • An escalation rotation for high severity issues outside your shift hours
  • A written handoff at the end of every shift
  • Your time zone matters to this hire

Qualifications

  • Three or more years supporting production servers, data center infrastructure, or bare metal and cloud environments
  • Strong hands-on Linux troubleshooting
  • Real experience with server hardware
  • Out-of-band management experience
  • Working TCP/IP knowledge
  • Experience working in ticketing, monitoring, incident management, or infrastructure management systems
  • Clear written English
  • Sound judgment alone in production

Helpful, not required

  • NVIDIA GPU servers at scale
  • HPC or AI training environments
  • Enterprise platforms from Dell, HPE, Supermicro, or Lenovo
  • NVMe, ZFS, Ceph, or distributed storage
  • Prometheus, Grafana, or similar observability tooling
  • NetBox or another infrastructure and asset register
  • Ansible, Terraform, or configuration management
  • Optics, transceivers, and DAC or AOC cabling
  • Working across geographically distributed third-party facilities

What success looks like

  • By 90 days:
    • Working the majority of your shift's issues without escalation
    • Every ticket you close carries an accurate category and a real closure reason
    • At least three runbooks written from issues you personally resolved
    • You know how to escalate to each facility contact on your shift without asking
  • By six months:
    • Your escalations to engineering arrive complete and are rarely handed back
    • Issues on your shift are increasingly caught before the customer reports them
    • Something that used to be a recurring ticket is gone because you removed the cause

Salary

95000 - 130000 USD Per annum

Before You Apply
️
πŸ‡ΊπŸ‡Έ Be aware of the location restriction for this remote position: USA Only
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Support Engineer, GPU Infrastructure @Hydra Host
All Others
Salary 95,000 - 130,00..
Remote Location
πŸ‡ΊπŸ‡Έ USA Only
Employment Type full-time
Posted 3wks ago
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
️
πŸ‡ΊπŸ‡Έ Be aware of the location restriction for this remote position: USA Only
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
Γ—
Apply to the best remote jobs
before everyone else

Access 125,000+ vetted remote jobs and get daily alerts.

4.9 β˜…β˜…β˜…β˜…β˜… from 500+ reviews

⚑ 126,955+ remote jobs, refreshed hourly

πŸ”” Real-time alerts: Apply first, direct to employer

πŸ›‘οΈ Vetted companies, no scams, true remote only

Unlock All Jobs Now

Maybe later