Research Scientist - Model Team
Mirelo AIAI Audio company
Berlin, GermanySenior
Andreessen Horowitz
Index Ventures
TriplePoint Capital
Atlantic.vc
Data & AI
About the role
TL;DR
Develop and implement multimodal generative models for audio generation.
- •At Mirelo, you'll work at the centre of how we build the next generation of multimodal video-to-audio models.
- •This role is deeply hands-on and research-heavy: with a great H100/200-per-engineer ratio you explore and build new multimodal models and push the boundaries of what's possible in music, sound, and speech generation.
- •Key Responsibilities Design, implement and train large-scale multimodal generative models for audio generation (diffusion and/or autoregressive models).
- •Explore new modeling ideas for audio generation (music, sound, speech) while taking inspiration from the language and image domains.
- •Develop and experiment with post-training for new capabilities (fine-grained control, in/out-painting, editing,...) Conduct rigorous ablation studies, get actionable insights and communicate results to the team to discuss new research directions.
- •Contribute hands-on to all stages of model development including data curation, experimentation, evaluation, and deployment.
- •Requirements Hands-on experience in training large-scale generative models in a fast-paced research environment.
- •Deep understanding of cutting-edge methods and ML research in at least one of the domains: image, language, video or audio (specific audio experience not necessary, but nice to have).
- •Strong proficiency in PyTorch, transformer architectures, and the full ecosystem of modern deep learning.
- •Solid understanding of distributed training techniques—FSDP, low precision training, model parallelism Strong track-record in working on generative models (publications in top-tier venues, open-source contributions or applied ML projects).
Required skills
PyTorch
Domain expertise
entertainment-music
Tech stack
PyTorch