Build, evaluate, and deploy machine learning models for voice AI.
•At Bolna, we're building tools that change the way teams leverage Voice AI.
•We're looking for a Founding Machine Learning Engineer to own the end-to-end lifecycle of building, evaluating, deploying, and improving models that power millions of production conversations.
•Key Responsibilities Build the data engine
•Design pipelines to source and clean conversational voice data across Indian languages, accents, and telephony conditions.
•Fine-tune models that ship
•Fine tune and train models to improve accuracy, speed, and reliability across different use-cases.
•Define what "good" means
•Build evaluation datasets and benchmarks for transcription accuracy, voice naturalness, interruption handling, latency, and end-to-end conversation quality.
•Set up human-in-the-loop pipelines to capture subjective quality at scale.
•Ship to production
•Work with the engineering team to deploy models into a latency-sensitive, high-volume system.
•Monitor performance in the wild, debug regressions, and iterate fast.
•Requirements 3+ years of hands-on ML experience with deep practical real-world experience in training models.
•Strong Python and PyTorch fundamentals with exposure in distributed training, and modern fine-tuning techniques (LoRA, QLoRA, DPO, RLHF, etc.).
•Training data as a first-class problem.
•Experience designing data pipelines from collection, cleaning, labeling, deduplication, augmentation and treating data quality as a core engineering discipline.
•Rigorous about evaluation.
•You know that "looks good in a demo" is not a benchmark.
•You build the evals before you trust the model.
•Speech model experience is a plus with real-time / streaming inference experience where you would have contributed to latency optimization, quantization, and distillation for production deployment.
•Bias toward shipping.
•You'd rather have a model running in production this week than a perfect one in a notebook next quarter.