Evaluations - Member of Technical Staff
SimileSimile is an
San Francisco, United States$200,000 - $400,000 USDLead
Data & AI
About the role
TL;DR
Build measurement systems to evaluate the accuracy and trustworthiness of human behavior simulations.
- •As a Member of Technical Staff, Model Evaluations at Simile, you will build the measurement systems that determine whether our simulations of human behavior are accurate, trustworthy, and useful enough to guide real-world decisions.
- •You will help shape what Simile measures, the quality bars we defend, and how evaluation evidence guides model, product, and customer decisions.
- •Key Responsibilities Build the measurement layer for behavioral simulation: Design evals, metrics, rubrics, datasets, dashboards, and workflows that measure whether Simile’s models are accurately predicting human behavior.
- •Partner with modeling to improve models: Evaluate new model versions, diagnose regressions, and identify priority areas for model-improvement cycles.
- •Contribute to product and applied evals: Build evals for qualitative responses, retrieval, survey generation, and other product surfaces.
- •Make ground truth and uncertainty legible: Develop rigorous ways to compare simulated responses against human data and behavioral datasets.
- •Automate evaluation workflows: Use modern agentic coding tools to rapidly build internal tools and inspect model outputs.
- •Requirements Strong intuition for what makes an eval meaningful, robust, and decision-relevant.
- •Understanding of modern LLM training, post-training, model evaluation, and hill-climbing.
- •Comfort reasoning about noisy data, uncertainty, sampling, distributions, and calibration.
- •Ability to build internal tools, scripts, dashboards, and automated eval pipelines quickly using Python, SQL, R, notebooks, LLM APIs, and agentic coding tools.
- •Willingness to independently drive a workstream and do the work yourself.
Required skills
PythonSQLLLMsOpenAI APIAnthropic APIGemini APILangChainHugging Face
Nice-to-have skills
MLflowSageMakerVertex AI
Domain expertise
ai
Benefits & perks
Equity, Health & Wellness, Time Off
Tech stack
PythonSQLROpenAI APIAnthropic APIGemini APILangChainHugging FaceLLMs