Lead research on realtime audio understanding, generation, and speech-to-speech models.
•You will lead research at the frontier of realtime conversation and human AI interaction.
•You will contribute to novel architectures, data, and evaluations for realtime audio, and translate them into state-of-the-art models used in voice agents around the world.
•Key Responsibilities Architect and develop new architectures for realtime audio understanding, generation, and speech-to-speech models.
•Contribute to frontier multimodal and multilingual datasets for pre-training and post-training.
•Set new standards for how we evaluate and benchmark our audio models.
•Requirements Strong applied mindset and ability to balance scientific novelty with product impact.
•Excited and able to work across the stack from infra, to data, to evals, to architecture.
•Deep expertise in deep generative modeling.
•Experience with large-scale training, GPU/TPU acceleration, and model optimization.