Member of Technical Staff — Inference-TPU
RadixArkAI Infrastructure company
Palo Alto, United States$200,000 - $400,000Lead
Accel
Spark Capital
NVentures
Salience Capital
A&E Investments
HOF Capital
Data & AI
About the role
TL;DR
Build high-performance inference and training systems using JAX, XLA, and Pallas on TPU hardware.
- •RadixArk is looking for a Member of Technical Staff
- •TPU Systems to build high-performance inference and training systems using JAX, XLA, and Pallas.
- •You'll push model workloads to their limits on TPU hardware, working on SGLang-JAX and other critical infrastructure that enables efficient deployment of frontier models on Google's tensor processing units.
- •Key Responsibilities Build high-performance inference and training systems using JAX/XLA/Pallas, including SGLang-JAX Push large-model workloads to the limits on the newest TPU hardwares Optimize end-to-end latency and throughput for LLM serving on TPU infrastructure Design and implement SPMD strategies for efficient distributed inference and training Design and implement Pallas kernels for operations that require customized low level control for best performance Requirements 3+ years experience building production ML systems utilizing JAX/Torch, XLA, or TPU-focused frameworks.
- •Bachelor's or Master's degree in Computer Science, Electrical Engineering, or equivalent industry experience Deep understanding of XLA internals preferred: HLO, MLIR, operator fusion, SPMD partitioning, and sharding strategies.
- •Strong performance tuning instincts across compiler and runtime layers Experience with distributed inference systems (e.g.
- •SGLang, vLLM) or training frameworks (e.g.
- •Miles, Alpa, Pathways)
Required skills
JAXPythonLLMs
Domain expertise
aideveloper-tools
Benefits & perks
equity
Tech stack
PythonJAX