Skip to content
FluidStack logo

Site Reliability Engineer, Compute

FluidStackGPU Cloud company
San Francisco, United StatesSenior
Situational Awareness
Astro Capital (NY)
7GC & Co
Armyn Capital
Autopilot Management Company
Bare Metal Ventures
Software Engineering

About the role

TL;DR

Own compute fleet health end to end for Fluidstack's large GPU fleet.

  • Fluidstack is building one of the largest GPU fleets in the world.
  • We need a Site Reliability Engineer to own compute fleet health end to end.
  • Key Responsibilities Build the metrics pipelines, alerting, and unified health view for the compute fleet.
  • Turn deployment/repair into a pipeline, not a procedure.
  • Design and expand the GPU qualification platform.
  • Own Redfish and BMC tooling.
  • Own end-to-end reliability, scalability, and operation of the compute fleet at-scale.
  • Requirements You treat toil as a bug.
  • You have an instinct for hardware.
  • You move toward ambiguity, not away from it.
  • You learn at a steep slope.
  • You carry a pager without flinching.
  • You're fluent with AI tooling.
  • You've shipped production automation that other teams depend on.
View original posting →

Required skills

PythonGoKubernetesCI/CDPrometheusGrafanaAnthropic APIDatabricksApache KafkaApache FlinkPostgreSQLMySQLMongoDBRedisElasticsearch

Domain expertise

ai

Benefits & perks

Competitive total compensation package (salary + equity), Retirement or pension plan, Health, dental, and vision insurance, Generous PTO policy

Tech stack

PythonGoKubernetesCI/CDPrometheusGrafanaAnthropic APIDatabricksApache KafkaApache FlinkPostgreSQLMySQLMongoDBRedisElasticsearchDynamoDBCassandraSQLiteOracleSQL ServerNeo4jCockroachDBTimescaleDBInfluxDBSupabaseFirebasePrismaAWSGoogle CloudAzure

Similar jobs

F

Senior Site Reliability Engineer

F

QA Analyst - LOIS for Word

F

Senior Database Reliability Engineer