Skip to content
Fireworks logo

Member of Technical Staff, AI Training Infrastructure

FireworksFireworks is the
San MateoLead
Data & AI

About the role

TL;DR

Design, build, and optimize infrastructure for large-scale AI model training.

  • As a Training Infrastructure Engineer, you'll design, build, and optimize the infrastructure that powers our large-scale model training operations.
  • Your work will be essential to developing high-performance AI training infrastructure.
  • Key Responsibilities Design and implement scalable infrastructure for large-scale model training workloads Develop and maintain distributed training pipelines for LLMs and multimodal models Optimize training performance across multiple GPUs, nodes, and data centers Implement monitoring, logging, and debugging tools for training operations Architect and maintain data storage solutions for large-scale training datasets Requirements Bachelor's degree in Computer Science, Computer Engineering, or related field, or equivalent practical experience 3+ years of experience with distributed systems and ML infrastructure Experience with PyTorch Proficiency in cloud platforms (AWS, GCP, Azure) Experience with containerization, orchestration (Kubernetes, Docker)
View original posting →

Required skills

PythonPyTorchAWSGoogle CloudAzureKubernetesDockerBash

Nice-to-have skills

LLMsComputer VisionNLP

Domain expertise

aideveloper-tools

Tech stack

PythonPyTorchAWSGoogle CloudAzureKubernetesDockerBash

Similar jobs

F

Principal Machine Learning Engineer

B

Lead Analytics Consultant (CJA)

F

Forward Deployed Engineer