TL;DR
Define how frontier AI models are measured by designing benchmarks and running experiments.
- •Intelligence is the product lab company behind DesignArena, a platform with over 5.5M+ users.
- •We rigorously evaluate state-of-the-art multimodal AI models for frontier model providers.
- •Key Responsibilities Design new benchmarks and evaluation methodologies for frontier AI models.
- •Run experiments and analyze model behavior to understand emerging capabilities.
- •Shape public leaderboards and evaluation tools shared with frontier labs.
- •Publish research, technical reports, and analyses on AI model evaluation.
- •Requirements Strong STEM background (Computer Science, Data Science, Statistics, Math, Engineering, Physics, or related).
- •Deep curiosity about frontier AI models and their behavior.
- •Intellectual drive to be at the forefront of AI development.
- •Willingness to get hands-on and do real work.
View original posting →