Customer Reliability Engineer
FluidStackGPU Cloud company
San Francisco, United StatesMid
Situational Awareness
Astro Capital (NY)
7GC & Co
Armyn Capital
Autopilot Management Company
Bare Metal Ventures
About the role
TL;DR
Ensure customer workloads reliability and debug distributed systems.
- •Own reliability for named customer workloads, debug across the full stack, and run customer-facing incident communication.
- •Key Responsibilities Own reliability for named customer workloads: their clusters, their SLAs, their escalations.
- •Debug across the full stack, hardware to fabric to scheduler, when a training run degrades.
- •Run customer-facing incident communication with technical depth and no spin.
- •Turn recurring customer pain into engineering fixes with the production teams.
- •Requirements You've supported large-scale compute customers (HPC, cloud, or AI labs) at a technical level.
- •You debug distributed systems methodically across layers you don't own.
- •You've written incident updates customers trusted more after reading.
- •You push internal teams to fix causes, not symptoms, and follow up until they do.
- •Bonus: GPU training workloads.
- •InfiniBand or RoCE.
- •Slurm or Kubernetes.
Required skills
Kubernetes
Domain expertise
ai
Tech stack
Kubernetes