AI Performance Engineer @Modular
Artificial Intelligence
Salary $180,000 - $270..
Remote Location
Employment Type full-time
Posted YDay

[Hiring] AI Performance Engineer @Modular

YDay - Modular is hiring a remote AI Performance Engineer. 💸 Salary: $180,000 - $270,000 usd 📍Location: USA, UK, Canada

Role Description

We are looking for an AI Performance Engineer to join the ASIC Performance team. You will design, implement, and tune high-performance attention kernels—and a broader set of foundational AI kernels—for new accelerator architectures.

  • Design and implement high-performance attention kernels for prefill and decode, including variants such as MHA, GQA, paged attention, sliding-window/global attention, and fused FlashAttention-style algorithms.
  • Build and optimize general AI kernels—including matrix multiplication, softmax, normalization, activation, embedding, reduction, and data-movement primitives—needed to bring up modern models on new accelerators.
  • Contribute to the design and implementation of the Attention framework, with reusable interfaces that balance portability, composability, tunability, and hardware-specific optimization.
  • Profile workloads, identify bottlenecks, establish performance models, and tune implementations against hardware limits and competitive baselines.
  • Develop benchmarks, correctness tests, regression tests, and automated performance analysis to ensure kernels remain reliable and fast across model shapes and hardware generations.
  • Partner with the rest of Modular teams to contribute to overall software stack improvements and share learnings.
  • Work directly with hardware partners when needed to understand new architectures, validate toolchains, diagnose low-level issues, and influence hardware/software interfaces.
  • Write design documents, participate in technical reviews, and share architecture and performance insights across the team.
  • Mentor teammates and raise the engineering bar for kernel quality, performance methodology, and maintainable low-level software.

Qualifications

  • 5+ years of relevant industry or research experience in high-performance computing, AI kernel development, accelerator programming, compiler engineering, or a closely related field.
  • Demonstrated experience writing and optimizing production-quality GPU or accelerator kernels, with a strong understanding of parallel algorithms, numerical behavior, memory access patterns, synchronization, and data movement.
  • Strong knowledge of modern attention algorithms and their implementation tradeoffs, including the distinct performance characteristics of prefill and decode.
  • Strong understanding of architectures beyond conventional GPUs—such as NPUs, DSPs, TPUs, or custom ASICs—and the ability to reason about unfamiliar compute units, memory hierarchies, and execution models.
  • Experience with at least one heterogeneous programming model or kernel ecosystem such as CUDA, Triton, SYCL, OpenCL, HIP, or a vendor accelerator SDK.
  • Strong performance-analysis skills and proficiency with relevant profiling, tracing, benchmarking, and debugging tools.
  • Ability to read architecture manuals, compiler output, and low-level generated code; form hypotheses from data; and systematically close performance gaps.
  • A collaborative, team-oriented approach; clear written and verbal communication; intellectual curiosity; and comfort owning ambiguous technical problems from investigation through productization.

Requirements

  • Knowledge of the Mojo programming language or experience authoring kernels and performance-sensitive libraries in Mojo.
  • Knowledge of MLIR, LLVM, and related AI compiler technologies, including dialect design, lowering pipelines, code generation, scheduling, or autotuning.
  • Experience implementing FlashAttention or other fused attention algorithms, custom attention operators, paged KV-cache kernels, or inference-specific attention optimizations.
  • Experience with modern kernel DSLs and libraries such as CuTe/CUTLASS, Triton, Pallas, or similar systems.
  • Familiarity with AI framework internals and operator integration in systems such as PyTorch, JAX, TensorFlow, vLLM, SGLang, or TensorRT-LLM.
  • Experience bringing up a model or kernel library on a new hardware platform, including working through compiler, runtime, driver, and hardware constraints.
  • Experience with performance modeling, autotuning, numerical validation, or benchmarking infrastructure.
  • Familiarity with LLM architectures such as GPT and Gemma, including grouped-query attention, mixture-of-experts, speculative decoding, and long-context inference.
  • An advanced degree in Computer Science, Electrical Engineering, Computer Engineering, or a related field.

Benefits

  • Amazing Team: We are a progressive and agile team with some of the industry’s best engineering and product leaders.
  • World-class Benefits: Your benefits package may include comprehensive healthcare coverage, retirement and savings programs, employee stock purchase opportunities, paid time off, wellbeing resources, family support programs, and learning and development opportunities.
  • Competitive Compensation: We offer very strong compensation packages, including RSU grants.
  • Team Building Events: We organize regular team onsites and local meetups in Los Altos, CA as well as different cities.
Before You Apply
️
remote Be aware of the location restriction for this remote position: USA, UK, Canada
‼ Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
AI Performance Engineer @Modular
Artificial Intelligence
Salary $180,000 - $270..
Remote Location
Employment Type full-time
Posted YDay
Apply for this position
Did not apply ✓
Applied ✓
Sent Follow-Up ✓
Interview Scheduled ✓
Interview Completed ✓
Offer Accepted ✓
Offer Declined ✓
Application Denied ✓
Unlock 130,000+ Remote Jobs
️
remote Be aware of the location restriction for this remote position: USA, UK, Canada
‼ Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply ✓
Applied ✓
Sent Follow-Up ✓
Interview Scheduled ✓
Interview Completed ✓
Offer Accepted ✓
Offer Declined ✓
Application Denied ✓
Unlock 130,000+ Remote Jobs
×
Apply to the best remote jobs
before everyone else

Access 130,000+ vetted remote jobs and get daily alerts.

4.9 ★★★★★ from 500+ reviews

⚡ 131,109+ remote jobs, refreshed hourly

🔔 Real-time alerts: Apply first, direct to employer

🛡️ Vetted companies, no scams, true remote only

Unlock All Jobs Now

Maybe later