Software Engineer, ML Infrastructure
IdeogramText-to-image generation company
Toronto, CanadaMid
Andreessen Horowitz
Index Ventures
Redpoint Ventures
Pear VC
SV Angel
AIX Ventures
Data & AI
About the role
TL;DR
Build and optimize systems for training and serving generative AI models.
- •As a Software Engineer, ML Infrastructure at Ideogram, you'll build the systems that power training and serving for our generative AI models at scale.
- •You'll work across the stack, from designing distributed training infrastructure to optimizing inference pipelines that serve millions of users.
- •Key Responsibilities Design distributed training infrastructure Optimize inference pipelines Deploy, support, and troubleshoot in complex Linux-based computing environments Experience in worker scaling for training or inference workloads Fundamental knowledge in ML models and how they run on GPUs Requirements 1-4 years developing and shipping large-scale production infrastructure Experience designing large, highly available distributed systems with Kubernetes/GCP, and GPU/TPU workloads on those clusters Experience in deploying, supporting, and troubleshooting in complex Linux-based computing environments Experience in worker scaling for training or inference workloads Fundamental knowledge in ML models and how they run on GPUs
Required skills
KubernetesGoogle CloudLinuxCI/CDDocker
Domain expertise
ai
Benefits & perks
Competitive compensation and equity, 4 weeks of vacation, Comprehensive health, vision, and dental coverage, RRSP/401(k) with employer match, Top-of-the-line tools and tech, Autonomy to explore and experiment, A culture of learning and growth
Tech stack
KubernetesGoogle CloudLinuxCI/CDDocker