Skip to content
Together AI logo

Research Engineer, Large-Scale Training

Together AIAI Infrastructure company
San Francisco, United States$200,000 - $290,000Mid
General Catalyst logo
General Catalyst
Prosperity7 Ventures
NVIDIA logo
NVIDIA
Salesforce Ventures logo
Salesforce Ventures
Kleiner Perkins logo
Kleiner Perkins
Coatue logo
Coatue
Data & AI

About the role

TL;DR

Research Engineer to optimize large-scale AI model training infrastructure.

  • The Model Shaping team at Together AI works on products and research for tailoring open foundation models to downstream applications.
  • We build services that allow machine learning developers to choose the best models for their tasks and further improve these models using domain-specific data.
  • In addition, we develop new methods for more efficient model training and evaluation, drawing inspiration from a broad spectrum of ideas across machine learning, natural language processing, and ML systems.
  • As a Research Engineer on the Scaling Team within Model Shaping, you will turn cutting-edge research on efficient foundation model training into robust, high-performance systems.
  • You will profile and optimize Together's training infrastructure, identify performance bottlenecks across the stack, and implement state-of-the-art techniques from both the research literature and our own scientists in production environments.
  • Your work will directly shape the fine-tuning experience of Together's customers.
  • You will rapidly bring newly released open-source models onto the Model Shaping platform, ensuring they train efficiently and reliably across diverse customer workloads.
  • Working closely with Research Scientists, you will also build the experimental infrastructure that accelerates research and enables validated ideas to be deployed reliably at scale.
  • Key Responsibilities Design, implement, and optimize core components of Together's large-scale training infrastructure.
  • Integrate new model architectures, validate training correctness and convergence, and optimize performance for production fine-tuning workloads.
  • Profile distributed training workloads to identify and eliminate bottlenecks across compute, memory, and communication.
  • Design and execute experiments to validate performance hypotheses and benchmark new approaches against state-of-the-art methods.
  • Partner closely with Research Scientists to productionize novel training methods and contribute to publications and open-source releases.
  • Requirements Demonstrated ability to independently take ambiguous performance or infrastructure problems from investigation through deployment.
  • Strong programming skills in Python and PyTorch, with an emphasis on writing efficient, maintainable code.
  • Hands-on experience training or fine-tuning large neural networks in multi-GPU or multi-node environments.
  • Solid understanding of ML systems fundamentals, including GPU architecture, mixed-precision training, and distributed training paradigms such as data, tensor, pipeline, or expert parallelism.
  • Strong communication skills and the ability to collaborate effectively with both researchers and engineers.
View original posting →

Required skills

PythonPyTorch

Domain expertise

ai

Benefits & perks

startup equity, health insurance

Tech stack

PythonPyTorchGit

Similar jobs

C

MLOps Field Engineer

I

Senior Machine Learning Engineer