Member of Technical Staff, AI Training Infrastructure
FireworksFireworks is the
San MateoLead
Data & AI
About the role
TL;DR
Design, build, and optimize infrastructure for large-scale AI model training.
- •As a Training Infrastructure Engineer, you'll design, build, and optimize the infrastructure that powers our large-scale model training operations.
- •Your work will be essential to developing high-performance AI training infrastructure.
- •Key Responsibilities Design and implement scalable infrastructure for large-scale model training workloads Develop and maintain distributed training pipelines for LLMs and multimodal models Optimize training performance across multiple GPUs, nodes, and data centers Implement monitoring, logging, and debugging tools for training operations Architect and maintain data storage solutions for large-scale training datasets Requirements Bachelor's degree in Computer Science, Computer Engineering, or related field, or equivalent practical experience 3+ years of experience with distributed systems and ML infrastructure Experience with PyTorch Proficiency in cloud platforms (AWS, GCP, Azure) Experience with containerization, orchestration (Kubernetes, Docker)
Required skills
PythonPyTorchAWSGoogle CloudAzureKubernetesDockerBash
Nice-to-have skills
LLMsComputer VisionNLP
Domain expertise
aideveloper-tools
Tech stack
PythonPyTorchAWSGoogle CloudAzureKubernetesDockerBash