Skip to content
Higgsfield logo

ML Systems Performance Engineer (MFU)

HiggsfieldGenerative AI company
Almaty, KazakhstanSenior
Accel logo
Accel
GFT Ventures
Menlo Ventures logo
Menlo Ventures
AI Capital Partners
BroadLight Capital
NextEquity Partners
Data & AINew

About the role

TL;DR

Optimize ML training performance across compute, memory, and communication.

  • Higgsfield AI is seeking an ML Systems Performance Engineer to optimize the performance of large-scale AI training runs.
  • You will work at the forefront of generative AI, contributing to a company that is rapidly scaling and rewriting the future of AI-powered video creation.
  • Key Responsibilities Profile end-to-end training runs and identify bottlenecks across compute, memory, communication, storage, and orchestration.
  • Define, measure, and improve MFU, tokens/sec/GPU, scaling efficiency, training goodput, and GPU uptime.
  • Optimize distributed training and model-sharding strategies.
  • Improve collective communication and optimize data loading, preprocessing, sequence packing, and checkpointing.
  • Requirements Strong experience running and optimizing multi-GPU or multi-node training.
  • Experience with PyTorch Distributed or an equivalent training framework.
  • Understanding of GPU architecture, collective communication, and distributed-training bottlenecks.
  • Ability to debug complex performance and reliability problems.
View original posting →

Required skills

PythonPyTorchBashLinuxGitDockerKubernetesAWSGoogle CloudAzure

Domain expertise

ai

Benefits & perks

Competitive base salary in USD, Equity participation, Relocation support to Almaty, Company-provided equipment, meals, transportation, or other office benefits

Tech stack

PythonPyTorchAWSGoogle CloudAzureDockerKubernetesLinuxGitBash

Similar jobs

Databricks logo

Forward Deployed Engineer

D

Senior Data Engineer

H

ML Engineer (Data Engine)