Lead and scale Runpod's core cloud and bare-metal environments.
•Own the critical foundational layers of our platform.
•Key Responsibilities Own Core Infrastructure & SRE: Lead multiple engineering teams responsible for Site Reliability Engineering, networking, and storage.
•Architect HPC & Global Networking: Oversee the design, scaling, and operation of Runpod's global network backbone.
•Drive Storage Engine Innovation: Direct the architecture and performance tuning of highly scalable, distributed storage systems.
•Build a High-Output Org: Hire, mentor, and grow highly technical engineering managers and senior ICs.
•Translate Scale into Strategy: Partner with Program Management and Product to forecast capacity requirements and shape technical roadmaps.
•Requirements Engineering Leadership Experience: 7+ years leading software, infrastructure, SRE, or networking teams.
•Deep Infrastructure Expertise: 8+ years building and operating large-scale distributed systems.
•HPC & Advanced Networking: Proven hands-on background or strong architectural understanding of ultra-low latency networking.
•Storage Systems Knowledge: Experience building, operating, or tuning high-performance distributed storage systems.
•SRE / DevOps Culture: Strong foundation in reliability engineering, infrastructure-as-code, container orchestration, and modern observability stacks.