Staff/Principal DevOps Engineer, AI Inference
Lila SciencesAI & company
Cambridge, United States$192,000 - $272,000 USDLead
Nvidia
Braidwell LP
Collective Global
Flagship Pioneering
Software Engineering
About the role
TL;DR
Build and optimize infrastructure for large-scale AI model inference on GPU clusters.
- •The Staff/Principal DevOps Engineer
- •AI Inference will drive the design, implementation, and optimization of infrastructure purpose-built for serving machine learning models at scale.
- •This role bridges platform engineering, site reliability, and ML infrastructure, building the systems that power low-latency, high-throughput inference across GPU clusters and cloud accelerators.
- •Key Responsibilities Design and implement GPU/accelerator infrastructure on Kubernetes for inference workloads.
- •Build model serving platforms using frameworks like vLLM or Triton Inference Server.
- •Develop autoscaling systems for dynamic compute matching across production and research workloads.
- •Create production-grade deployment pipelines for ML models with canary rollouts and A/B testing.
- •Implement observability and performance optimization for GPU utilization and inference latency.
- •Requirements Expertise in DevOps, SRE, or Platform Engineering with significant experience operating GPU/accelerator infrastructure at scale.
- •Deep experience with Kubernetes for ML workloads, including GPU scheduling and accelerator device management.
- •Strong proficiency deploying to AWS using infrastructure-as-code (Terraform, Helm) with hands-on experience managing GPU-based compute.
- •Experience with model serving infrastructure, inference servers, and request batching.
- •Strong understanding of networking for distributed inference.
- •Strong proficiency in Python for automation and tooling.
Required skills
KubernetesAWSTerraformPythonCI/CDDocker
Nice-to-have skills
Helm
Domain expertise
deeptechai
Benefits & perks
bonus potential, early-stage equity, medical, dental, vision coverage, employer-paid life and disability insurance, flexible time off, company wide holidays, paid parental leave, educational assistance program, commuter benefits, company subsidized lunch program
Tech stack
KubernetesTerraformHelmAWSEKSEC2S3PythonDockerCI/CDGit