Member of Technical Staff — Inference-Multi-Hardware
RadixArkAI Infrastructure company
Palo Alto, United States$200,000 - $400,000 USDLead
Accel
Spark Capital
NVentures
Salience Capital
A&E Investments
HOF Capital
Data & AI
About the role
TL;DR
Optimize and maintain AI inference infrastructure across diverse hardware platforms.
- •RadixArk is seeking a Member of Technical Staff
- •Inference-Multi-Hardware to push the limits of performance for frontier AI systems.
- •Most performance engineering assumes a single vendor's stack.
- •This role assumes none.
- •You'll bring up, optimize, and maintain SGLang, Miles, and the RadixArk infrastructure stack across NVIDIA and AMD GPUs, Google TPUs, modern server CPUs, and a growing set of emerging AI accelerators.
- •That means porting kernels and runtimes onto unfamiliar hardware, designing the abstractions that keep one codebase fast on all of it.
- •You will be working directly with silicon and our partners, often on pre-release platforms with immature tooling.
- •Requirements 4+ years of experience in systems, performance, or ML infrastructure engineering Deep expertise in at least one accelerator programming model (CUDA, ROCm/HIP, Pallas/XLA, Triton, or a vendor SDK), with demonstrated ability to pick up new ones quickly.
- •Strong understanding of accelerator architecture: memory hierarchy, bandwidth limits, occupancy, and the tradeoffs between them Experience writing or optimizing high-performance kernels for ML workloads Experience with distributed execution and communication libraries (NCCL, RCCL, MPI, or equivalents) Proficiency in C++ and Python Strong debugging and profiling skills at the system level, including on platforms where the tooling is incomplete or unreliable Track record of performance work that shipped into production
Required skills
PythonC++LinuxGit
Nice-to-have skills
MLflowKubeflowSageMakerVertex AIONNXJAXLangChainHugging FaceTensorFlowPyTorch
Domain expertise
aideveloper-tools
Benefits & perks
equity
Tech stack
PythonC++GitLinux