About the role
TL;DR
Design and maintain large-scale backend and cloud-native infrastructure for AI model training.
- •Fireworks is seeking a Training Infrastructure Engineer to design, develop, and maintain large-scale backend and cloud-native infrastructure for their generative AI platform.
- •This role involves architecting scalable systems and collaborating with ML and product teams.
- •Key Responsibilities Architect and build scalable, resilient backend infrastructure for distributed ML training and inference.
- •Lead technical design discussions and establish best practices for large-scale ML systems.
- •Drive infrastructure optimization initiatives for compute cost, storage, and network performance.
- •Collaborate with ML, DevOps, and product teams to translate requirements into infrastructure solutions.
- •Evaluate and integrate cloud-native and open-source technologies like Kubernetes and Ray.
- •Requirements Bachelor's degree in Computer Science or related field with 4 years of experience. 4 years of experience in designing and optimizing large-scale backend infrastructure and distributed data systems in cloud environments. 4 years of experience with major server-side programming languages (Python, C++, Go, TypeScript). 3 years of experience developing and maintaining data processing and API systems. 2 years of experience with cloud-native tools like Docker and Kubernetes.
Required skills
PythonC++GoTypeScriptPostgreSQLMySQLDynamoDBSparkApache FlinkApache KafkaKubernetesDockerAWSGoogle CloudAzure
Nice-to-have skills
KubeflowMLflowgRPC
Domain expertise
aideveloper-tools
Tech stack
PythonC++GoTypeScriptPostgreSQLMySQLDynamoDBSparkApache FlinkApache KafkaKubernetesKubeflowMLflowgRPCDockerAWSGoogle CloudAzure