Skip to content
Mirendil logo

Member of Technical Staff, Model Evaluation

MirendilAI Research company
San Francisco, United States$300K - $500KLead
Andreessen Horowitz logo
Andreessen Horowitz
Kleiner Perkins logo
Kleiner Perkins
Nvidia logo
Nvidia
Data & AI

About the role

TL;DR

Build evaluation infrastructure to measure model capabilities.

  • We are looking for a research engineer to build the evaluation infrastructure that tells us whether our models are getting better in ways we care about.
  • Key Responsibilities Design and build evaluation frameworks that measure model capabilities along realistic axes, beyond standard benchmarks.
  • Build automated eval pipelines and regression-detection systems that run continuously and surface signal quickly.
  • Develop agent-assisted workflows for humans to efficiently inspect model behavior.
  • Requirements Extensive experience in AI/ML model evaluation and benchmarking.
  • Proficiency in Python and machine learning frameworks such as TensorFlow or PyTorch.
  • Strong understanding of model capabilities and evaluation metrics.
View original posting →

Required skills

PythonTensorFlowscikit-learnPandasNumPyAirflowAWSKubernetesCI/CDDocker

Domain expertise

ai

Benefits & perks

Equity

Tech stack

PythonTensorFlowscikit-learnPandasNumPyAirflowAWSKubernetesCI/CDDocker

Similar jobs

R

Senior Machine Learning Engineer, Reliability

R

Senior Data Scientist, Creator Platform

R

Senior Software Engineer, ML Infra (Communication Safety)