TL;DR
Ensure the stability and resilience of Runpod's distributed platform.
Key Responsibilities Define and implement SLIs/SLOs for critical services Lead incident response and coordinate cross-team mitigation efforts Conduct blameless postmortems and ensure corrective actions are completed Requirements Strong background in reliability engineering Experience with observability systems and reliability tooling Proficiency in automation and production hardening