Senior Principal AI Engineer @Cerence
Artificial Intelligence
Salary unspecified
Remote Location
πŸ‡ΊπŸ‡Έ USA Only
Employment Type full-time
Posted 2mths ago

[Hiring] Senior Principal AI Engineer @Cerence

2mths ago - Cerence is hiring a remote Senior Principal AI Engineer. πŸ’Έ Salary: unspecified πŸ“Location: USA

Role Description

A Moving Experience.

What You Will Work On:

  • Design and operate distributed training systems for large neural networks (autoregressive, diffusion, State Space Models, etc.) across GPU clusters.
  • Optimise multi-node, multi-GPU execution to maximize throughput and utilization.
  • Diagnose & resolve bottlenecks across compute, memory, and network.
  • Improve training stability and fault tolerance at scale.
  • Partner with research and applied ML teams to productionize large-model training pipelines.

Core Responsibilities

  • Distributed Training Infrastructure:
    • Build and optimize GPU cluster orchestration using: Slurm, Kubernetes, Ray, RunAI.
    • Ensure efficient scheduling, isolation, and fairness across training workloads.
  • Communication & Networking:
    • Optimize and debug distributed communication using: NCCL, RDMA, InfiniBand, NVLink.
    • Minimize networking bottlenecks that dominate end-to-end training time.
  • Training Frameworks:
    • Scale large-model training using: PyTorch, Distributed Megatron-LM, DeepSpeed.
    • Own multi-node launch configurations, failure recovery, and performance tuning.
  • Memory & Performance Optimization:
    • Apply advanced memory optimization techniques: Activation checkpointing, ZeRO (Stage 1–3) and offload strategies.
    • Balance compute, memory, and communication to push model size and batch scale.

What Success Looks Like

  • GPU utilization consistently stays high (>80–90%).
  • Training scales cleanly from single node to dozens or hundreds of GPUs.
  • Communication overhead is minimized and predictable.
  • Large training jobs run stably for days or weeks without failure.
  • New models can be trained faster, larger, and more reliably than before.

Qualifications

  • Deep hands-on experience with distributed systems or ML systems.
  • Experience running large-scale workloads on GPU clusters.
  • Production experience with PyTorch distributed training.
  • Strong understanding of parallelism strategies (data, tensor, pipeline parallelism).
  • Low-level understanding of GPU communication and networking.

Requirements

  • Critical Technical Skills:
    • GPU orchestration: Slurm, Kubernetes, Ray, RunAI.
    • Communication libraries: NCCL, RDMA, InfiniBand, NVLink.
    • Training frameworks: PyTorch Distributed, Megatron-LM, DeepSpeed.
    • Memory optimization: activation checkpointing, ZeRO offload techniques.

Common Problems You’ll Be Solving

  • Many teams fail at scale because:
    • GPU utilization is low despite large clusters.
    • Networking and communication dominate training time.
    • Training jobs crash or become unstable at large scale.
  • You will be explicitly focused on eliminating these failure modes.

Ideal Background

  • This role is a strong fit for individuals who have worked as:
    • ML Systems Engineer.
    • Distributed Systems Engineer.
    • AI Infrastructure Engineer.
    • HPC Engineer transitioning into ML.
  • Experience working with large language models or foundation models is a strong plus, but deep systems expertise is valued over pure model architecture experience.

Why This Role Matters

  • Without robust distributed training infrastructure, progress on large models stalls. This role directly enables:
    • Larger models.
    • Faster iteration cycles.
    • More reliable research-to-production pipelines.
  • You will be building the foundation that makes large-scale AI possible.
Before You Apply
️
πŸ‡ΊπŸ‡Έ Be aware of the location restriction for this remote position: USA Only
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Senior Principal AI Engineer @Cerence
Artificial Intelligence
Salary unspecified
Remote Location
πŸ‡ΊπŸ‡Έ USA Only
Employment Type full-time
Posted 2mths ago
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
️
πŸ‡ΊπŸ‡Έ Be aware of the location restriction for this remote position: USA Only
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
Γ—
Apply to the best remote jobs
before everyone else

Access 125,000+ vetted remote jobs and get daily alerts.

4.9 β˜…β˜…β˜…β˜…β˜… from 500+ reviews

⚑ 127,048+ remote jobs, refreshed hourly

πŸ”” Real-time alerts: Apply first, direct to employer

πŸ›‘οΈ Vetted companies, no scams, true remote only

Unlock All Jobs Now

Maybe later