Manage and optimize Runpod's global GPU datacenter infrastructure.
•Runpod is seeking a Datacenter Infrastructure Specialist to be the operational linchpin of their global GPU fleet.
•This role bridges hardware partners and internal engineering teams, focusing on the technical lifecycle and operational health of the expanding fleet.
•Key Responsibilities Validate new hardware and ensure partner deployments meet specifications for AI/ML workloads.
•Monitor fleet health, audit downtime, and provide technical data to protect customer SLAs.
•Automate network triage and generate dynamic runbooks using AI agents.
•Coordinate technical incident communications and translate outages into actionable resolutions.
•Requirements 3-5 years of experience in infrastructure operations, systems reliability, or datacenter engineering.
•Proficiency in datacenter networking and performance troubleshooting, with exposure to RDMA, InfiniBand, or RoCE.
•Hands-on experience with the NVIDIA Software Stack and multi-node performance tuning.
•Solid Linux system administration skills and experience with containerization (Docker).