ML Systems Performance Engineer (MFU)
HiggsfieldGenerative AI company
Almaty, KazakhstanSenior
Accel
Menlo Ventures
GFT Ventures
AI Capital Partners
BroadLight Capital
NextEquity Partners
Data & AINew
About the role
TL;DR
Optimize ML training performance across compute, memory, and communication.
- •Higgsfield AI is seeking an ML Systems Performance Engineer to optimize the performance of large-scale AI training runs.
- •You will work at the forefront of generative AI, contributing to a company that is rapidly scaling and rewriting the future of AI-powered video creation.
- •Key Responsibilities Profile end-to-end training runs and identify bottlenecks across compute, memory, communication, storage, and orchestration.
- •Define, measure, and improve MFU, tokens/sec/GPU, scaling efficiency, training goodput, and GPU uptime.
- •Optimize distributed training and model-sharding strategies.
- •Improve collective communication and optimize data loading, preprocessing, sequence packing, and checkpointing.
- •Requirements Strong experience running and optimizing multi-GPU or multi-node training.
- •Experience with PyTorch Distributed or an equivalent training framework.
- •Understanding of GPU architecture, collective communication, and distributed-training bottlenecks.
- •Ability to debug complex performance and reliability problems.
Required skills
PythonPyTorchBashLinuxGitDockerKubernetesAWSGoogle CloudAzure
Domain expertise
ai
Benefits & perks
Competitive base salary in USD, Equity participation, Relocation support to Almaty, Company-provided equipment, meals, transportation, or other office benefits
Tech stack
PythonPyTorchAWSGoogle CloudAzureDockerKubernetesLinuxGitBash