Skip to content
Cantina logo

Machine Learning Engineer, Speech - Joint Audio-Video Modeling

CantinaDigital Transformation company
Remote$200K - $220KSenior
Data & AI

About the role

TL;DR

Builds state-of-the-art speech and audio generation systems with a focus on joint audio-video modeling.

  • Cantina Labs is seeking a Research / ML Engineer to join their Speech Team to build state-of-the-art speech and audio generation systems end-to-end, with a focus on joint audio-video modeling.
  • This role involves owning the audio side of multimodal generation, from representations to generative backbones and conditioning machinery.
  • Key Responsibilities Design, train, and improve audio VAEs, neural codecs, and vocoders.
  • Architect, implement, and fine-tune diffusion and flow-matching transformers for audio and video generation.
  • Design audio conditioning and cross-modal alignment within joint AV models.
  • Define data requirements and collaborate on acquisition, curation, and quality filtering.
  • Design and run experiments to advance model understanding and drive improvements.
  • Requirements Exceptional research/development experience with large-scale audio models.
  • Deep hands-on experience with diffusion and/or flow-matching transformers.
  • Strong experience with multi-node, multi-GPU distributed training.
  • Strong software engineering skills with PyTorch and production-quality code.
  • Shipped large-scale speech/audio or multimodal generative models to production.
View original posting →

Required skills

PythonPyTorchHugging FaceAWSGoogle CloudAzureDockerKubernetesGitComputer VisionNLPLLMs

Nice-to-have skills

TensorFlowscikit-learnJAX

Domain expertise

aimedia

Benefits & perks

Competitive salary, Generous company equity, Medical, dental, and vision insurance, 42 days of paid time off, Generous parental leave & fertility support, 401(k) retirement savings plan, Lifestyle spending account, Complimentary lunch and snacks, One Medical membership

Tech stack

PythonPyTorchTensorFlowscikit-learnHugging FaceAWSGoogle CloudAzureDockerKubernetesGit

Similar jobs

C

Machine Learning Engineer - Voice Conversion

F

Forward Deployed Engineer

F

Senior Forward Deployed Engineer