Datacenter Infrastructure Specialist @Runpod
All Others
Salary eur 105,324 - 1..
Remote Location
remote UK
Employment Type full-time
Posted 4d ago

[Hiring] Datacenter Infrastructure Specialist @Runpod

4d ago - Runpod is hiring a remote Datacenter Infrastructure Specialist. 💸 Salary: eur 105,324 - 140,432 per year 📍Location: UK

Role Description

We are looking for a Datacenter Infrastructure Specialist to be the operational linchpin of our global fleet. Reporting to the Manager of Infrastructure Capacity & Management, you will serve as the technical authority bridging our hardware partners and internal engineering teams.

As our Datacenter Infrastructure Specialist, you will own the technical lifecycle and operational health of Runpod’s rapidly expanding, high-density GPU fleet. You will be part of a team that acts as the technical anchor for our hardware partners—serving as their infrastructure advisor, technical translator, adopter, and incident commander. This role blends deep HPC systems engineering, advanced network troubleshooting, and process automation to ensure rock-solid uptime for the world's most demanding AI workloads.

This is a high-visibility, high-impact position where you will move beyond traditional ticket-closing. You will work directly with cutting-edge GPU cloud technologies, advanced RDMA fabrics, and modern observability stacks to solve complex hardware challenges at scale. If you want the autonomy to build automated infrastructure tooling and directly contribute to the resilience, scalability, and revenue velocity of Runpod’s global physical backbone, this is where you do it.

Responsibilities

  • Hardware Validation & Benchmarking:
    • Assist in validating new hardware, ensuring partner deployments meet Runpod’s specifications for distributed AI/ML workloads.
  • Uptime & SLA Enforcement:
    • Monitor fleet health to identify performance degradation.
    • Help audit downtime and provide the technical data needed to protect customer SLAs.
  • AI-Driven Operations:
    • Operate with an AI-first mindset, powering operations with the technology hosted.
    • Work with LLMs and AI agents to help automate network triage and generate dynamic runbooks for the fleet.
  • Incident Support:
    • Coordinate technical incident communications with clear updates, acting as a steady hand that translates outages into actionable resolutions.
  • Partner Technical Support:
    • Support the growth of our infrastructure partners.

Qualifications

  • 3–5 years of experience in infrastructure operations, systems reliability, or datacenter engineering.
  • Strong proficiency in standard datacenter networking and performance troubleshooting.
  • Exposure to RDMA, InfiniBand, or RoCE is highly preferred.
  • Hands-on experience with the NVIDIA Software Stack (driver installation, performance utilities) and an understanding of multi-node performance tuning.
  • Solid Linux system administration skills and experience with containerization (Docker).
  • Clear written and verbal communication skills.
  • Operational flexibility; may require participating in an on-call rotation in the future.
  • Detail-oriented and proactive when it comes to identifying potential failures before they impact customers.

Requirements

  • Experience working in a fast-paced environment where you have contributed to building operational workflows.
  • Experience managing or optimizing bare-metal High-Performance Computing environments at massive scale.
  • Experience with Grafana, Prometheus, or Datadog to monitor system health.
  • Proficiency in Python, Go (Golang), or Bash to automate repetitive infrastructure tasks and interface with internal APIs.

Benefits

  • Competitive base pay ranging from €105,324.00 to €140,432.00.
  • Meaningful equity in a fast-growing AI infra company — everyone receives stock options.
  • Generous medical, dental & vision plans — 100% coverage for all employees and partial for dependents.
  • Flexible PTO — take the time you need to recharge.
  • Most roles are remote work first with inclusive, collaborative teams utilizing Slack as the main form of internal communication.
  • $1,200 Home Office & Equipment Stipend — support to create your ideal workspace.
Before You Apply
️
remote Be aware of the location restriction for this remote position: UK
‼ Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Datacenter Infrastructure Specialist @Runpod
All Others
Salary eur 105,324 - 1..
Remote Location
remote UK
Employment Type full-time
Posted 4d ago
Apply for this position
Did not apply ✓
Applied ✓
Sent Follow-Up ✓
Interview Scheduled ✓
Interview Completed ✓
Offer Accepted ✓
Offer Declined ✓
Application Denied ✓
Unlock 125,000+ Remote Jobs
️
remote Be aware of the location restriction for this remote position: UK
‼ Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply ✓
Applied ✓
Sent Follow-Up ✓
Interview Scheduled ✓
Interview Completed ✓
Offer Accepted ✓
Offer Declined ✓
Application Denied ✓
Unlock 125,000+ Remote Jobs
×
Apply to the best remote jobs
before everyone else

Access 125,000+ vetted remote jobs and get daily alerts.

4.9 ★★★★★ from 500+ reviews

⚡ 128,483+ remote jobs, refreshed hourly

🔔 Real-time alerts: Apply first

🛡️ Vetted companies, no scams, true remote only

Unlock All Jobs Now

Maybe later