Member of Technical Staff — Performance
RadixArkAI Infrastructure company
Palo Alto, United States$200,000 - $400,000Lead
Accel
Spark Capital
NVentures
Salience Capital
A&E Investments
HOF Capital
Data & AI
About the role
TL;DR
Push LLM inference and training systems to the limit across real production workloads.
- •RadixArk is hiring a Member of Technical Staff
- •Performance in Palo Alto, CA
- •someone who can push LLM inference and training systems to the limit across real production workloads.
- •You’ll work on the performance-critical path of SGLang, Miles, and the RadixArk infrastructure stack: latency, throughput, GPU utilization, memory efficiency, scheduling, batching, kernel behavior, distributed execution, and cost-per-token.
- •Key Responsibilities Analyze and improve performance across SGLang, Miles, and RadixArk production deployments Benchmark LLM inference and training workloads across GPUs, TPUs, and cloud environments Optimize latency, throughput, memory usage, batching, scheduling, routing, and GPU utilization Investigate performance regressions in real customer environments Build internal tooling for profiling, tracing, benchmarking, and regression detection Requirements Strong systems engineering background, especially in performance-critical software Experience with GPU systems, distributed systems, inference serving, ML runtimes, or high-performance computing Familiarity with profiling tools, performance debugging, tracing, and benchmark methodology Comfort working with Python and C++ Understanding of LLM inference concepts such as batching, KV cache, prefill/decode, speculative decoding, MoE, long context, and P99 latency
Required skills
PythonC++LLMs
Domain expertise
aideveloper-tools
Benefits & perks
equity
Tech stack
PythonC++LLMs