Research Engineer, Large-Scale Training
Together AIAI Infrastructure company
San Francisco, United States$200,000 - $290,000Mid
General Catalyst
NVIDIA
Salesforce Ventures
Kleiner Perkins
Coatue
Prosperity7 Ventures
Data & AI
About the role
TL;DR
Research Engineer to optimize large-scale AI model training infrastructure.
- •The Model Shaping team at Together AI works on products and research for tailoring open foundation models to downstream applications.
- •We build services that allow machine learning developers to choose the best models for their tasks and further improve these models using domain-specific data.
- •In addition, we develop new methods for more efficient model training and evaluation, drawing inspiration from a broad spectrum of ideas across machine learning, natural language processing, and ML systems.
- •As a Research Engineer on the Scaling Team within Model Shaping, you will turn cutting-edge research on efficient foundation model training into robust, high-performance systems.
- •You will profile and optimize Together's training infrastructure, identify performance bottlenecks across the stack, and implement state-of-the-art techniques from both the research literature and our own scientists in production environments.
- •Your work will directly shape the fine-tuning experience of Together's customers.
- •You will rapidly bring newly released open-source models onto the Model Shaping platform, ensuring they train efficiently and reliably across diverse customer workloads.
- •Working closely with Research Scientists, you will also build the experimental infrastructure that accelerates research and enables validated ideas to be deployed reliably at scale.
- •Key Responsibilities Design, implement, and optimize core components of Together's large-scale training infrastructure.
- •Integrate new model architectures, validate training correctness and convergence, and optimize performance for production fine-tuning workloads.
- •Profile distributed training workloads to identify and eliminate bottlenecks across compute, memory, and communication.
- •Design and execute experiments to validate performance hypotheses and benchmark new approaches against state-of-the-art methods.
- •Partner closely with Research Scientists to productionize novel training methods and contribute to publications and open-source releases.
- •Requirements Demonstrated ability to independently take ambiguous performance or infrastructure problems from investigation through deployment.
- •Strong programming skills in Python and PyTorch, with an emphasis on writing efficient, maintainable code.
- •Hands-on experience training or fine-tuning large neural networks in multi-GPU or multi-node environments.
- •Solid understanding of ML systems fundamentals, including GPU architecture, mixed-precision training, and distributed training paradigms such as data, tensor, pipeline, or expert parallelism.
- •Strong communication skills and the ability to collaborate effectively with both researchers and engineers.
Required skills
PythonPyTorch
Domain expertise
ai
Benefits & perks
startup equity, health insurance
Tech stack
PythonPyTorchGit