Ensure model quality, regression testing, and release validation for clinical AI systems.
•We are looking for an ML Evaluation Engineer to own model quality, regression testing, release validation, and production impact analysis for clinical AI systems.
•Key Responsibilities Build and maintain evaluation frameworks for clinical NLP, LLM, RAG, information extraction, and structured abstraction systems.
•Create and manage hidden test datasets that are not directly visible to model developers, reducing overfitting risk.
•Compare model versions and identify performance degradation across clinical segments, document types, clients, data sources, labels, and edge cases.
•Work with Clinical AI Data Specialists to design gold sets, hidden test sets, adjudication workflows, and label quality checks.
•Requirements 3-6+ years of experience in ML engineering, data science, model evaluation, ML QA, applied NLP evaluation, or data-heavy quality engineering.
•Strong Python and data analysis skills.
•Strong understanding of precision/recall/F1, calibration, confidence thresholds, dataset splits, leakage, overfitting, statistical testing, and error analysis.
•Experience building evaluation pipelines, benchmark suites, test harnesses, dashboards, or regression frameworks.
•Ability to work with imperfect labels, annotation disagreement, clinical ambiguity, and hidden evaluation sets.