Skip to content
Mirelo AI logo

Training Infrastructure Engineer

Mirelo AIAI Audio company
Berlin, GermanySenior
Andreessen Horowitz logo
Andreessen Horowitz
Index Ventures logo
Index Ventures
Atlantic.vc
TriplePoint Capital logo
TriplePoint Capital
Data & AI

About the role

TL;DR

Focus on optimizing and managing ML training infrastructure at scale.

  • In this role, you’ll focus on the full training stack
  • profiling GPU behavior, debugging training pipelines, improving throughput, choosing the right parallelism strategies, and designing the infrastructure that lets us train models efficiently at scale.
  • Key Responsibilities Find ideal training strategies (parallelism approaches, precision trade-offs) for a variety of model sizes and compute loads Profile, debug, and optimize single and multi-GPU operations using tools like Nsight and stack trace viewers to understand what's actually happening at the hardware level Analyze and improve the whole training pipeline from start to end (efficient data storage, data loading, distributed training, checkpoint/artifact saving, logging, …) Set up scalable systems for experiment tracking, data/model versioning, experiment insights.
  • Design, deploy and maintain large-scale ML training clusters running SLURM for distributed workload orchestration Requirements Familiarity with the latest and most effective techniques in optimizing training and inference workloads—not from reading papers, but from implementing them Deep understanding of GPU memory hierarchy and computation capabilities—knowing what the hardware can do theoretically and what prevents us from achieving it Experience optimizing for both memory-bound and compute-bound operations and understanding when each constraint matters Expertise with efficient attention algorithms and their performance characteristics at different scales
View original posting →

Required skills

PyTorchDockerKubernetesCI/CDApache KafkaPostgreSQLMySQLMongoDBRedisElasticsearchAWSGoogle CloudAzureTerraformCloudFormation

Domain expertise

ai

Benefits & perks

Competitive compensation and equity

Tech stack

PyTorchDockerKubernetesCI/CDApache KafkaPostgreSQLMySQLMongoDBRedisElasticsearchAWSGoogle CloudAzureTerraformCloudFormationPulumiEC2S3RDSECSEKSGKECloud RunBigQueryRedshiftAthenaDatabricksFivetranDagsterPrefect

Similar jobs

D

Senior Data Engineer

H

ML Systems Performance Engineer (MFU)

Databricks logo

Forward Deployed Engineer