Senior ML Infrastructure Engineer
Prior LabsTabular Foundation company
Berlin, GermanySenior
Balderton Capital
Atlantic Labs
XTX Ventures
Hector Foundation
Galion.exe
Data & AI
About the role
TL;DR
Design and manage infrastructure for training tabular foundation models.
- •We spend tens of millions per year on GPU compute to train tabular foundation models.
- •The person who owns this infrastructure makes decisions worth millions of dollars: cluster architecture, scheduling efficiency, provider strategy, hardware selection.
- •A wrong call costs six figures.
- •Key Responsibilities Design and implement scalable infrastructure for training tabular foundation models.
- •Optimize cluster architecture and scheduling efficiency.
- •Develop and maintain CI/CD pipelines for model training and deployment.
- •Collaborate with data scientists and engineers to ensure smooth model training and deployment.
- •Monitor and manage cloud resources to ensure cost-effectiveness and performance.
- •Requirements Proven experience in designing and managing large-scale infrastructure.
- •Strong knowledge of cloud services (AWS, Google Cloud, Azure).
- •Experience with containerization and orchestration (Docker, Kubernetes).
- •Proficiency in CI/CD tools and practices.
- •Excellent problem-solving skills and attention to detail.
Required skills
TensorFlowPyTorchKubernetesDockerCI/CDAWSGoogle CloudAzureTerraformCloudFormationPulumiEC2S3RDSECS
Domain expertise
ai
Tech stack
TensorFlowPyTorchscikit-learnKubernetesDockerCI/CDAWSGoogle CloudAzureTerraformCloudFormationPulumiEC2S3RDSECSEKSGKECloud RunBigQueryRedshiftAthenaPostgreSQLMySQLMongoDBRedisElasticsearchDynamoDBCassandraSQLite