Lead research at the intersection of spatiotemporal modeling and multimodal representation learning.
•As a Senior ML Research Scientist, you will lead research at the intersection of spatiotemporal modeling and multimodal representation learning.
•Key Responsibilities Research multimodal representations that connect temporal and spatial structure with semantic understanding.
•Develop models that operate across multiple levels of granularity, from full videos and clips to regions and entities.
•Design training objectives, datasets, evaluation methods, and large-scale experiments.
•Use task-specific vision models as supervision or components, and integrate useful signals into general-purpose representations.
•Partner with research and engineering teams to scale and ship new capabilities.
•Requirements A strong research track record in computer vision, video understanding, multimodal learning, or a related field.
•Deep expertise in at least one of the following: video foundation models, self-supervised or contrastive learning, embeddings, and retrieval; or temporal modeling, object-centric learning, detection, segmentation, and tracking.
•Strong hands-on skills in Python and PyTorch or an equivalent deep learning framework.
•Proven ability to design and run rigorous experiments at scale.
•Research impact through publications, production systems, or both.