Machine Learning Engineer, Speech - Joint Audio-Video Modeling @Cantina
Artificial Intelligence
Salary usd 200,000 - 2..
Remote Location
πŸ‡ΊπŸ‡Έ USA Only
Employment Type full-time
Posted 4d ago

[Hiring] Machine Learning Engineer, Speech - Joint Audio-Video Modeling @Cantina

4d ago - Cantina is hiring a remote Machine Learning Engineer, Speech - Joint Audio-Video Modeling. πŸ’Έ Salary: usd 200,000 - 220,000 per year πŸ“Location: USA

Role Description

We're looking for a Research / ML Engineer to join our Speech Team to build state-of-the-art speech and audio generation systems end-to-end from data specs through production inference with a focus on joint audio-video modeling.

You'll own the audio side of multimodal generation:

  • The representations (audio VAEs, neural codecs)
  • The generative backbone (diffusion / flow-matching transformers)
  • The conditioning and alignment machinery that makes characters speak, sing, and emote in sync with what's on screen

This includes:

  • Voice cloning and multi-speaker conditioning inside joint AV models
  • Cinematic dialogue with music and sound design
  • Adjacent speech tasks (controllable TTS, voice conversion) that feed the same stack

You'll drive the model ↔ data ↔ eval flywheel, partnering closely with research, video, data, and infra to ship fast, reliable, and cost-aware models. In this role you'll work at the intersection of cutting-edge research and practical engineering, contributing to the development of safe, steerable, and trustworthy AI systems.

You will thrive in this role if you:

  • See research and engineering as two sides of the same coin and enjoy owning work end-to-end
  • Are excited to work across modalities and collaborate closely with a video generation team rather than staying inside audio
  • Are results-oriented, flexible, and willing to pick up whatever moves the needle
  • Like collaborating closely with infra, data, and product to ship measurable improvements
  • Enjoy designing experiments, listening tests, and metrics that correlate with user-perceived quality
  • Are eager to learn every day, and to find and solve unique large-scale problems

Qualifications

  • Exceptional research/development experience with large-scale audio models (>8B parameters, >500k hours of data)
  • Deep hands-on experience with diffusion and/or flow-matching transformers, including practical knowledge of samplers, schedules, conditioning mechanisms, and distillation
  • Deep hands-on experience training audio VAEs, neural audio codecs, and vocoders latent/tokenizer design, reconstruction and perceptual objectives, adversarial training
  • Strong experience with multi-node, multi-GPU distributed training (FSDP/DeepSpeed or equivalent)
  • Strong software engineering skills with a proven track record of building complex systems
  • Strong with PyTorch and performance work (profiling, CUDA/Triton/C++ as needed) and writing reliable production-quality code
  • Shipped large-scale speech/audio or multimodal generative models to production
  • Background in working with large-scale ML data, and the ability to iterate on data and triangulate quality using both subjective and objective signals
  • Experience with voice cloning, speech control/steerability, or expressive speech generation
  • Notable publications and/or open-source contributions in speech/audio/ML

Requirements

  • Experience with multimodal audio-video modeling: joint AV generation of multi-shot, multi-speaker scenes with dialogue, music, and sound design generated jointly with video, and the cross-modal alignment that keeps them in sync
  • Experience with video generation: video diffusion/flow-matching transformers, video VAEs, conditioned and multi-shot generation, building data pipelines for video models
  • Streaming or real-time generation, causal distillation (e.g., Self Forcing / Self Forcing++)

Benefits

  • Competitive salary and generous company equity
  • Medical, dental, and vision insurance – 99.99% of premiums covered by Cantina
  • 42 days of paid time off, including:
    • 15 PTO days
    • 10 sick days
    • 15 company holidays
    • 2 floating holidays
  • Generous parental leave & fertility support
  • 401(k) retirement savings plan
  • Lifestyle spending account – $500/month to use however you’d like
  • Complimentary lunch and snacks for in-office employees
  • One Medical membership, and more!
Before You Apply
️
πŸ‡ΊπŸ‡Έ Be aware of the location restriction for this remote position: USA Only
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Machine Learning Engineer, Speech - Joint Audio-Video Modeling @Cantina
Artificial Intelligence
Salary usd 200,000 - 2..
Remote Location
πŸ‡ΊπŸ‡Έ USA Only
Employment Type full-time
Posted 4d ago
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
️
πŸ‡ΊπŸ‡Έ Be aware of the location restriction for this remote position: USA Only
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
Γ—

Apply to the best remote jobs
before everyone else

Access 125,000+ vetted remote jobs and get daily alerts.

4.9 β˜…β˜…β˜…β˜…β˜… from 500+ reviews
Unlock All Jobs Now

Maybe later