Machine Learning Engineer - Voice Conversion @Cantina
Artificial Intelligence
Salary usd 200,000 - 2..
Remote Location
🇺🇸 USA Only
Employment Type full-time
Posted 2mths ago

[Hiring] Machine Learning Engineer - Voice Conversion @Cantina

2mths ago - Cantina is hiring a remote Machine Learning Engineer - Voice Conversion. 💸 Salary: usd 200,000 - 220,000 per year 📍Location: USA

Role Description

We’re looking for a Research / ML Engineer to join our Speech Team to build state-of-the-art speech systems end-to-end—from data specs through production inference. You’ll drive the model ↔ data ↔ eval flywheel for VC and adjacent tasks (controllable TTS, voice design and more), partnering closely with research, data, and infra to ship fast, reliable, and cost-aware models. In this role, you will work at the intersection of cutting-edge research and practical engineering, contributing to the development of safe, steerable, and trustworthy AI systems.

You will thrive in this role if you:

  • See research and engineering as two sides of the same coin and enjoy owning work end-to-end.
  • Are results-oriented, flexible, and willing to pick up whatever moves the needle.
  • Like collaborating closely with infra, data, and product to ship measurable improvements.
  • Enjoy designing experiments, listening tests, and metrics that correlate with user-perceived quality.
  • Eager to learn every day, find and solve unique large-scale problems.

What You’ll Do

  • Model Building: Architect, implement, pre-train, fine-tune, and post-train/alignment (e.g., GRPO/DPO) for large-scale speech models.
  • Experimental Design: Design, run, and analyze scientific experiments to advance our understanding of the models.
  • Tool Development: Develop and improve dev tooling to enhance team productivity.
  • Full-Stack Contribution: Contribute to the entire stack, from low-level optimizations to high-level model design.
  • Data Ownership: Define data requirements and collaborate on acquisition, curation, augmentation, labeling quality, and synthetic data strategies.
  • Rigorous Evaluation: Design automated objective/subjective evaluations—listening tests, SV/WER/ASR-based metrics, robustness & bias checks, and red-team studies.
  • Pipeline Delivery: Harden the training → evaluation → inference pipeline; profile latency, memory, and cost; and meet production SLAs with robust monitoring and rollback.
  • Safety & Responsibility: Contribute to safety/consent guardrails and to misuse/abuse mitigation for responsible speech technology.

Qualifications

  • Exceptional research/development experience with large-scale audio models (>8B parameters, >500k hours of data).
  • Deep hands-on experience with diffusion and/or flow-matching transformers, including practical knowledge of samplers, schedules, conditioning mechanisms, and distillation.
  • Deep hands-on experience training audio VAEs, neural audio codecs, and vocoders latent/tokenizer design, reconstruction and perceptual objectives, adversarial training.
  • Strong experience with multi-node, multi-GPU distributed training (FSDP/DeepSpeed or equivalent).
  • Strong software engineering skills with a proven track record of building complex systems.
  • Strong with PyTorch and performance work (profiling, CUDA/Triton/C++ as needed) and writing reliable production-quality code.
  • Shipped large-scale speech/audio or multimodal generative models to production.
  • Background in working with large-scale ML data, and the ability to iterate on data and triangulate quality using both subjective and objective signals.
  • Experience with voice cloning, speech control/steerability, or expressive speech generation.
  • Notable publications and/or open-source contributions in speech/audio/ML.

Compensation

The anticipated annual base salary range for this role is between $200,000-$220,000 (€170,000-€190,000). When determining compensation, a number of factors will be considered, including skills, experience, job scope, location, and competitive compensation market data.

Benefits

  • Competitive salary and generous company equity
  • Medical, dental, and vision insurance – 99.99% of premiums covered by Cantina
  • 42 days of paid time off, including:
    • 15 PTO days
    • 10 sick days
    • 15 company holidays
    • 2 floating holidays
  • Generous parental leave & fertility support
  • 401(k) retirement savings plan
  • Lifestyle spending account – $500/month to use however you’d like
  • Complimentary lunch and snacks for in-office employees
  • One Medical membership, and more!
Before You Apply
️
🇺🇸 Be aware of the location restriction for this remote position: USA Only
‼ Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Machine Learning Engineer - Voice Conversion @Cantina
Artificial Intelligence
Salary usd 200,000 - 2..
Remote Location
🇺🇸 USA Only
Employment Type full-time
Posted 2mths ago
Apply for this position
Did not apply ✓
Applied ✓
Sent Follow-Up ✓
Interview Scheduled ✓
Interview Completed ✓
Offer Accepted ✓
Offer Declined ✓
Application Denied ✓
Unlock 125,000+ Remote Jobs
️
🇺🇸 Be aware of the location restriction for this remote position: USA Only
‼ Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply ✓
Applied ✓
Sent Follow-Up ✓
Interview Scheduled ✓
Interview Completed ✓
Offer Accepted ✓
Offer Declined ✓
Application Denied ✓
Unlock 125,000+ Remote Jobs
×
Apply to the best remote jobs
before everyone else

Access 125,000+ vetted remote jobs and get daily alerts.

4.9 ★★★★★ from 500+ reviews

⚡ 127,038+ remote jobs, refreshed hourly

🔔 Real-time alerts: Apply first, direct to employer

🛡️ Vetted companies, no scams, true remote only

Unlock All Jobs Now

Maybe later