Machine Learning Engineer, Speech - Joint Audio-Video Modeling
CantinaDigital Transformation company
Remote$200K - $220KSenior
Data & AI
About the role
TL;DR
Builds state-of-the-art speech and audio generation systems with a focus on joint audio-video modeling.
- •Cantina Labs is seeking a Research / ML Engineer to join their Speech Team to build state-of-the-art speech and audio generation systems end-to-end, with a focus on joint audio-video modeling.
- •This role involves owning the audio side of multimodal generation, from representations to generative backbones and conditioning machinery.
- •Key Responsibilities Design, train, and improve audio VAEs, neural codecs, and vocoders.
- •Architect, implement, and fine-tune diffusion and flow-matching transformers for audio and video generation.
- •Design audio conditioning and cross-modal alignment within joint AV models.
- •Define data requirements and collaborate on acquisition, curation, and quality filtering.
- •Design and run experiments to advance model understanding and drive improvements.
- •Requirements Exceptional research/development experience with large-scale audio models.
- •Deep hands-on experience with diffusion and/or flow-matching transformers.
- •Strong experience with multi-node, multi-GPU distributed training.
- •Strong software engineering skills with PyTorch and production-quality code.
- •Shipped large-scale speech/audio or multimodal generative models to production.
Required skills
PythonPyTorchHugging FaceAWSGoogle CloudAzureDockerKubernetesGitComputer VisionNLPLLMs
Nice-to-have skills
TensorFlowscikit-learnJAX
Domain expertise
aimedia
Benefits & perks
Competitive salary, Generous company equity, Medical, dental, and vision insurance, 42 days of paid time off, Generous parental leave & fertility support, 401(k) retirement savings plan, Lifestyle spending account, Complimentary lunch and snacks, One Medical membership
Tech stack
PythonPyTorchTensorFlowscikit-learnHugging FaceAWSGoogle CloudAzureDockerKubernetesGit