Site Reliability Engineer, Compute
FluidStackGPU Cloud company
San Francisco, United StatesSenior
Situational Awareness
Astro Capital (NY)
7GC & Co
Armyn Capital
Autopilot Management Company
Bare Metal Ventures
Software Engineering
About the role
TL;DR
Own compute fleet health end to end for Fluidstack's large GPU fleet.
- •Fluidstack is building one of the largest GPU fleets in the world.
- •We need a Site Reliability Engineer to own compute fleet health end to end.
- •Key Responsibilities Build the metrics pipelines, alerting, and unified health view for the compute fleet.
- •Turn deployment/repair into a pipeline, not a procedure.
- •Design and expand the GPU qualification platform.
- •Own Redfish and BMC tooling.
- •Own end-to-end reliability, scalability, and operation of the compute fleet at-scale.
- •Requirements You treat toil as a bug.
- •You have an instinct for hardware.
- •You move toward ambiguity, not away from it.
- •You learn at a steep slope.
- •You carry a pager without flinching.
- •You're fluent with AI tooling.
- •You've shipped production automation that other teams depend on.
Required skills
PythonGoKubernetesCI/CDPrometheusGrafanaAnthropic APIDatabricksApache KafkaApache FlinkPostgreSQLMySQLMongoDBRedisElasticsearch
Domain expertise
ai
Benefits & perks
Competitive total compensation package (salary + equity), Retirement or pension plan, Health, dental, and vision insurance, Generous PTO policy
Tech stack
PythonGoKubernetesCI/CDPrometheusGrafanaAnthropic APIDatabricksApache KafkaApache FlinkPostgreSQLMySQLMongoDBRedisElasticsearchDynamoDBCassandraSQLiteOracleSQL ServerNeo4jCockroachDBTimescaleDBInfluxDBSupabaseFirebasePrismaAWSGoogle CloudAzure