Senior Site Reliability Engineer @Parasail
All Others
Salary unspecified
Remote Location
🇺🇸 USA Only
Employment Type full-time
Posted 1wk ago

[Hiring] Senior Site Reliability Engineer @Parasail

1wk ago - Parasail is hiring a remote Senior Site Reliability Engineer. 💸 Salary: unspecified 📍Location: USA

Role Description

At Parasail, reliability is an engineering problem that spans the entire stack. A GPU fails. A provider goes down. Traffic spikes. Customers still expect their inference to work.

We’re hiring Site Reliability Engineers to build the systems that make that possible. You’ll own infrastructure across our global GPU fleet, write software that automates operations, and make the platform better at detecting, surviving, and recovering from failures.

You’ll work directly with infrastructure, platform, and inference engineers in a flat organization. We welcome SREs, software engineers, platform engineers, and systems engineers who want to build ambitious systems and take responsibility for how they perform in production.

What You’ll Do

  • Scale a global GPU fleet.
  • Build and improve the Kubernetes infrastructure behind provisioning, networking, storage, and service deployment across providers and regions.
  • Make failure survivable.
  • Design better isolation, failover, and recovery so hardware and infrastructure failures have less impact on customers.
  • Build software that runs infrastructure.
  • Automate capacity expansion, deployments, and maintenance, eliminating manual work and making changes safer.
  • Make the system understandable.
  • Develop observability and diagnostics that reveal bottlenecks, surface failures, and help engineers act quickly.
  • Own the production feedback loop.
  • Respond to incidents, get to the root cause, and turn what you learn into stronger systems.
  • Push the platform forward.
  • Work across the stack to improve performance, utilization, security, and reliability as inference demand grows.

Qualifications

  • Experience building and operating production infrastructure or distributed systems, with real ownership of reliability.
  • Strong Linux fundamentals and practical knowledge of networking, storage, and containers.
  • Hands-on experience running Kubernetes in production.
  • The ability to write maintainable software and automation to solve infrastructure problems.
  • A systematic approach to debugging problems that cross application, cluster, network, and hardware boundaries.
  • Good judgment about when to move quickly, when to simplify, and where reliability matters most.
  • The initiative to take a problem from investigation through implementation and work closely with teammates along the way.
  • Your strongest skill might be software development, distributed systems, or infrastructure operations.

Nice to Have

  • Experience with multi-region, multi-provider, or bare-metal infrastructure.
  • Familiarity with GPUs, model serving, or inference systems such as vLLM or SGLang.
  • Experience with infrastructure as code, CI/CD, observability, or automated recovery.
  • Experience building highly available services, multi-tenant platforms, or distributed data systems.

Benefits

  • The systems you build will determine how reliably and efficiently customers can run AI in production.
  • You’ll work close to the hardware, deep in distributed systems, and alongside engineers optimizing the inference stack.
  • This is a small team tackling problems at substantial scale.
  • You’ll own meaningful architecture decisions, ship improvements directly into production, and help build the foundation for the next stage of AI infrastructure.
Before You Apply
️
🇺🇸 Be aware of the location restriction for this remote position: USA Only
‼ Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Senior Site Reliability Engineer @Parasail
All Others
Salary unspecified
Remote Location
🇺🇸 USA Only
Employment Type full-time
Posted 1wk ago
Apply for this position
Did not apply ✓
Applied ✓
Sent Follow-Up ✓
Interview Scheduled ✓
Interview Completed ✓
Offer Accepted ✓
Offer Declined ✓
Application Denied ✓
Unlock 125,000+ Remote Jobs
️
🇺🇸 Be aware of the location restriction for this remote position: USA Only
‼ Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply ✓
Applied ✓
Sent Follow-Up ✓
Interview Scheduled ✓
Interview Completed ✓
Offer Accepted ✓
Offer Declined ✓
Application Denied ✓
Unlock 125,000+ Remote Jobs
×
Apply to the best remote jobs
before everyone else

Access 125,000+ vetted remote jobs and get daily alerts.

4.9 ★★★★★ from 500+ reviews

⚡ 126,895+ remote jobs, refreshed hourly

🔔 Real-time alerts: Apply first, direct to employer

🛡️ Vetted companies, no scams, true remote only

Unlock All Jobs Now

Maybe later