Member of Technical Staff, Model Evaluation
MirendilAI Research company
San Francisco, United States$300K - $500KLead
Andreessen Horowitz
Kleiner Perkins
Nvidia
Data & AI
About the role
TL;DR
Build evaluation infrastructure to measure model capabilities.
- •We are looking for a research engineer to build the evaluation infrastructure that tells us whether our models are getting better in ways we care about.
- •Key Responsibilities Design and build evaluation frameworks that measure model capabilities along realistic axes, beyond standard benchmarks.
- •Build automated eval pipelines and regression-detection systems that run continuously and surface signal quickly.
- •Develop agent-assisted workflows for humans to efficiently inspect model behavior.
- •Requirements Extensive experience in AI/ML model evaluation and benchmarking.
- •Proficiency in Python and machine learning frameworks such as TensorFlow or PyTorch.
- •Strong understanding of model capabilities and evaluation metrics.
Required skills
PythonTensorFlowscikit-learnPandasNumPyAirflowAWSKubernetesCI/CDDocker
Domain expertise
ai
Benefits & perks
Equity
Tech stack
PythonTensorFlowscikit-learnPandasNumPyAirflowAWSKubernetesCI/CDDocker