TL;DR
Define how frontier AI models are measured and evaluated.
- •Define how frontier AI models are measured by designing new benchmarks, running experiments, and analyzing model behavior.
- •Your work will shape industry-trusted evaluation signals and tools shared with frontier labs.
- •Key Responsibilities Design industry-leading evaluations for frontier model performance.
- •Investigate model failures and identify emerging capabilities.
- •Publish research and technical reports shaping AI model evaluation.
- •Requirements Strong STEM background (Computer Science, Data Science, Statistics, Math, Engineering, Physics, or related).
- •Deep curiosity about frontier AI models and their behavior.
- •Genuine thirst for being on the frontier of AI development.
- •Fearlessness to get hands dirty and do real work.
View original posting →