Member of Technical Staff — Cluster Infrastructure & Supercomputing
RadixArkAI Infrastructure company
Palo Alto, United States$200,000 - $400,000 USDLead
Accel
Spark Capital
NVentures
Salience Capital
A&E Investments
HOF Capital
Software Engineering
About the role
TL;DR
Architect and scale AI compute clusters for frontier-level AI training and inference.
- •RadixArk is looking for a Member of Technical Staff Cluster Infrastructure to architect and scale the core compute platform that powers frontier-level AI training and inference.
- •You will design and operate highly reliable, high-performance GPU/TPU clusters, build next-generation scheduling and resource management systems, and push the limits of large-scale distributed infrastructure for AI workloads.
- •This role focuses on deep systems engineering across cluster architecture, networking, scheduling, and performance optimization.
- •Your work will directly impact how efficiently frontier AI models are trained and served.
- •Key Responsibilities Architect and scale large AI compute clusters for training and inference Design cluster management, scheduling, and resource allocation systems Optimize performance, utilization, and reliability of GPU/TPU clusters Improve fault tolerance and system resilience at scale Drive observability, monitoring, and performance profiling for cluster infrastructure Requirements 5+ years of experience in distributed systems, infrastructure, or large-scale compute platforms Strong background in distributed systems design and systems architecture Deep experience with cluster management systems (Kubernetes, Slurm, Ray, or custom schedulers) Hands-on experience with GPU/TPU infrastructure in production environments Strong Linux systems and networking fundamentals
Required skills
KubernetesGoPythonLinux
Nice-to-have skills
RustC++
Domain expertise
aideeptech
Benefits & perks
equity
Tech stack
GoRustC++PythonLinuxKubernetes