Senior AI Inference Engineer @StackYak
Artificial Intelligence
Salary competitive com..
Remote Location
πŸ‡ΊπŸ‡Έ USA Only
Employment Type full-time
Posted YDay

[Hiring] Senior AI Inference Engineer @StackYak

YDay - StackYak is hiring a remote Senior AI Inference Engineer. πŸ’Έ Salary: competitive compensation plus meaningful equity. πŸ“Location: USA

Role Description

We need someone who can take a model, a set of GPUs, and a production requirement and determine how that model should actually run.

  • Handle a frontier open-weight model β€” dense or mixture-of-experts, tens to hundreds of billions of parameters.
  • Make decisions on precision, quantization, tensor and pipeline and expert parallelism, memory strategy, batching, concurrency, topology, and serving runtime.
  • Benchmark, deploy, debug, tune, automate, and operate real inference systems.

This is not a research position and it is not an architecture-only role.

What You Will Own

  • How models run: placement, parallelism, precision, and memory strategy across single-GPU, multi-GPU, and multi-node.
  • The serving runtimes: vLLM, SGLang, TensorRT-LLM, and others.
  • Promises on concurrency limits, latency and TTFT targets, throughput under load.
  • Benchmarking methodology, harness, reproducibility.
  • Model lifecycle in production: loading, startup, health, upgrades, capacity, and failure handling.
  • Multi-tenant versus dedicated serving.
  • Inference incidents, including those unrelated to the serving runtime.
  • Turning knowledge into software to prevent single points of failure.

What Success Looks Like

  • Quickly produce a deployment that is sensible, measurable, reproducible, and ready to operate.
  • Help make expertise programmatic.
  • Answer questions regarding hardware compatibility, precision strategies, GPU or node splitting, serving runtimes, concurrency limits, and latency versus throughput trade-offs.

What We Need

  • Experience running large language models in production under real load.
  • Served a model across multiple GPUs and nodes.
  • Correctly sized a model against hardware.
  • Chosen quantization and precision strategies with real consequences.
  • Tuned batching and concurrency beyond easy wins.
  • Measured and defended latency, TTFT, throughput, utilization, and cost.
  • Worked in NVIDIA and/or AMD inference environments.
  • Written Python for production tooling and products.
  • Debugged Linux and systems problems below the container boundary.
  • Identified root causes of inference failures outside the serving runtime.

You Will Be Especially Strong If

  • Experience running inference infrastructure at an AI company or serious internal AI platform.
  • Worked across multiple GPU generations and vendors.
  • Understand distributed inference and networking implications.
  • Benchmark and compare serving runtimes.
  • Explain deployment configurations beyond vendor recipes.
  • Automated deployment decisions or built related infrastructure.
  • Build and test outside assigned roadmaps.

This Is Probably Not For You If

  • Your background is primarily model training or ML research.
  • You have limited experience with GPU memory, parallelism, topology, and concurrency.
  • You treat vLLM defaults as architecture.
  • You are interested in benchmarks but not in operating systems.
  • You prefer to remain the expert rather than encode knowledge into software.
  • You need narrowly defined ownership or a long runway before taking responsibility.

How We Work

  • Small, senior team with direct access to founders.
  • Strong opinions, loosely held.
  • Everyone participates in technical decisions.
  • Shared responsibility for production and on-call.
  • Value practical experience over degrees or certifications.
  • Expect technical arguments and decision-making.
  • Remote and distributed across time zones.
  • Hiring involves real conversations with potential colleagues.
  • Hiring across inference, infrastructure, and networking.
  • Fast-paced, early-stage startup environment.

Compensation

Competitive compensation plus meaningful equity. Exact structure will depend on location, engagement model, and experience.

A Note For Agencies

We are not using external recruiters or agencies for this role and will not respond to unsolicited emails.

Before You Apply
️
πŸ‡ΊπŸ‡Έ Be aware of the location restriction for this remote position: USA Only
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Senior AI Inference Engineer @StackYak
Artificial Intelligence
Salary competitive com..
Remote Location
πŸ‡ΊπŸ‡Έ USA Only
Employment Type full-time
Posted YDay
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
️
πŸ‡ΊπŸ‡Έ Be aware of the location restriction for this remote position: USA Only
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
Γ—
Apply to the best remote jobs
before everyone else

Access 125,000+ vetted remote jobs and get daily alerts.

4.9 β˜…β˜…β˜…β˜…β˜… from 500+ reviews

⚑ 128,627+ remote jobs, refreshed hourly

πŸ”” Real-time alerts: Apply first, direct to employer

πŸ›‘οΈ Vetted companies, no scams, true remote only

Unlock All Jobs Now

Maybe later