Staff Software Development Test Engineer - AI Evaluation
TekionAutomotive Retail company
Bangalore, IndiaLead
Dragoneer Investment Group
Index Ventures
Alkeon Capital
Durable Capital
Hyundai Motor Company
Advent International
Data & AINew
About the role
TL;DR
Build and scale AI evaluation platforms and frameworks for measuring AI quality.
- •Tekion is seeking a Senior SDET
- •AI Evaluation to build and scale AI evaluation platforms and frameworks.
- •This role is crucial for measuring, trusting, and improving the quality of AI outputs across various business domains.
- •Key Responsibilities Define and build Tekion’s AI evaluation capabilities as a shared platform service.
- •Design evaluation datasets, automated scoring pipelines, and quality metrics.
- •Own systems that assess AI model quality for release and production monitoring.
- •Utilize AI and LLMs to scale evaluation processes, including automated judges and synthetic datasets.
- •Requirements 5-8 years in SDET, quality engineering, ML engineering, or data science with experience in evaluation systems.
- •Strong Python programming skills for building robust evaluation pipelines.
- •Deep understanding of ML/LLM evaluation, benchmark design, and non-deterministic systems.
- •Hands-on experience with LLM-as-judge, rubric-based scoring, or human-in-the-loop evaluation.
- •Solid grasp of LLM/agent concepts and generative failure modes.
Required skills
PythonLLMsCI/CDDockerKubernetesAWSGoogle CloudAzure
Nice-to-have skills
Helm
Domain expertise
automotiveai
Tech stack
PythonLLMsCI/CDDockerKubernetesAWSGoogle CloudAzureHelm