Engineering Lead, Inference Optimization @Venice.ai
Artificial Intelligence
Salary $270,000-$330,0..
Remote Location
πŸ‡ΊπŸ‡Έ USA Only
Employment Type full-time
Posted 3d ago

[Hiring] Engineering Lead, Inference Optimization @Venice.ai

3d ago - Venice.ai is hiring a remote Engineering Lead, Inference Optimization. πŸ’Έ Salary: $270,000-$330,000 πŸ“Location: USA

Role Description

Venice is the only AI platform that runs inference with zero data retention and zero training on user inputs. This is an opportunity for you to be on the bleeding edge of privacy-focused AI with a unique and dedicated team of high-agency individuals alongside you. This role requires both hands-on work as an individual contributor as well as the management of a small team. You will play a pivotal role, shaping Venice's overarching technical strategy and assembling an exceptional team to deliver peak inference performance at massive scale.

What you'll do

  • Own Venice’s technical strategy for inference performance
  • Recruit and lead the Inference Optimization Team at Venice
  • Optimize Venice's GPU infrastructure across a range of architectures (e.g. H200s, B300s)
  • Improve latency, throughput, and cost per token for LLM inference workloads
  • Build reproducible benchmarking harnesses across inference engines (e.g. vLLM, SGLang) to identify the optimal engine, quantization scheme, and parallelism strategy per workload and GPU SKU
  • Work with our inference routing system to optimize multivariate inference load-balancing algorithms
  • Evaluate emerging inference optimization techniques (custom CUDA/Triton kernels), novel attention variants, new quantization schemes, and compilation stack improvements. Hands-on kernel development experience is a strong plus.
  • Evaluate emerging inference hardware (FPGAs, ASICs, custom silicon) for viability in Venice's stack.

Qualifications

  • 8+ years in performance optimization or HPC, with deep GPU architecture and parallel programming knowledge
  • 5+ years experience leading engineering teams
  • Proficiency in Python, Rust, or Go.
  • Bonus: C++/CUDA
  • Hands-on experience with at least one production LLM inference engine (e.g. vLLM, SGLang) running at high volume in production
  • Demonstrated experience with LLM inference optimization techniques: continuous batching, PagedAttention/KV cache management, speculative decoding, quantization, CUDA graphs, and torch.compile
  • Fluency with quantization tradeoffs, both qualitative and quantitative
  • Experience with distributed inference strategies (tensor parallelism, pipeline parallelism, MoE parallelism) in multi-GPU and multi-node environments
  • Fluency with GPU profiling (Nsight Systems, Nsight Compute, PyTorch Profiler) and a bias toward measuring before optimizing
  • Bonus: diffusion/image model inference optimization, custom Triton kernels, contributions to open-source inference frameworks

Benefits

Before You Apply
️
πŸ‡ΊπŸ‡Έ Be aware of the location restriction for this remote position: USA Only
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Engineering Lead, Inference Optimization @Venice.ai
Artificial Intelligence
Salary $270,000-$330,0..
Remote Location
πŸ‡ΊπŸ‡Έ USA Only
Employment Type full-time
Posted 3d ago
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
️
πŸ‡ΊπŸ‡Έ Be aware of the location restriction for this remote position: USA Only
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
Γ—

Apply to the best remote jobs
before everyone else

Access 125,000+ vetted remote jobs and get daily alerts.

4.9 β˜…β˜…β˜…β˜…β˜… from 500+ reviews
Unlock All Jobs Now

Maybe later