Build and deploy AI agent pipelines for extracting structured oncology variables from patient documents.
•Build and deploy AI agent pipelines that extract structured oncology variables from unstructured patient documents for pharmaceutical companies and cancer hospitals.
•Key Responsibilities Design and build agentic extraction pipelines that process 500+ page patient charts and output structured JSON per customer data dictionaries.
•Own accuracy end-to-end: define evaluation datasets, run precision/recall analysis per variable, identify failure modes, and improve through agent architecture changes, prompt engineering, fine-tuning, or rule-based post-processing.
•Go deep into the clinical source data
•read the actual patient charts, understand how oncologists document, learn why certain data points are ambiguous and use that understanding to improve extraction.
•Work with the clinical annotation team to build gold-standard datasets and resolve edge cases.
•Coordinate with customer data science and clinical teams to clarify dictionary definitions, review output quality, and close accuracy gaps.
•Requirements 2+ years building ML/AI systems in production Built and deployed AI agents or multi-step LLM pipelines Strong Python Practical LLM experience: prompt engineering, fine-tuning, RAG, evaluation design Built evaluation frameworks for LLM based document extraction tasks Willingness to become a domain expert in oncology data Comfortable owning customer-facing communication alongside technical delivery Can operate in high-intensity delivery sprints and manage your own time across multiple workstreams