Skip to content
FluidStack logo

Customer Reliability Engineer

FluidStackGPU Cloud company
San Francisco, United StatesMid
Situational Awareness
Astro Capital (NY)
7GC & Co
Armyn Capital
Autopilot Management Company
Bare Metal Ventures

About the role

TL;DR

Ensure customer workloads reliability and debug distributed systems.

  • Own reliability for named customer workloads, debug across the full stack, and run customer-facing incident communication.
  • Key Responsibilities Own reliability for named customer workloads: their clusters, their SLAs, their escalations.
  • Debug across the full stack, hardware to fabric to scheduler, when a training run degrades.
  • Run customer-facing incident communication with technical depth and no spin.
  • Turn recurring customer pain into engineering fixes with the production teams.
  • Requirements You've supported large-scale compute customers (HPC, cloud, or AI labs) at a technical level.
  • You debug distributed systems methodically across layers you don't own.
  • You've written incident updates customers trusted more after reading.
  • You push internal teams to fix causes, not symptoms, and follow up until they do.
  • Bonus: GPU training workloads.
  • InfiniBand or RoCE.
  • Slurm or Kubernetes.
View original posting →

Required skills

Kubernetes

Domain expertise

ai

Tech stack

Kubernetes

Similar jobs

C

Enterprise Sales Engineer - India

T

Senior Engineering Manager - Test Systems & Tooling

K

Senior Database Administrator - Core Infrastructure