Member of Technical Staff — Inference-Core Engine
RadixArkAI Infrastructure company
Palo Alto, United States$200,000 - $400,000 USDLead
Accel
Spark Capital
NVentures
Salience Capital
A&E Investments
HOF Capital
Data & AI
About the role
TL;DR
Design and optimize large-scale AI inference systems for frontier models.
- •RadixArk is seeking a Member of Technical Staff
- •Inference to push the limits of large-scale AI inference.
- •You will work on the core systems that serve frontier models at scale, optimizing performance, latency, throughput, and cost across thousands of GPUs.
- •Key Responsibilities Design and build large-scale inference systems for frontier AI models Optimize latency, throughput, and GPU utilization in production inference Develop and improve model serving architectures and runtimes Work on batching, scheduling, and memory management strategies Collaborate with kernel, compiler, and systems teams on performance optimization Requirements 5+ years of experience in systems engineering, ML infrastructure, or performance-critical backend systems Strong expertise in large-scale inference systems for LLMs or generative models Deep understanding of GPU architecture and performance characteristics Experience optimizing latency
- •and throughput-critical production systems Proficiency in Python, Rust, C++, or Go for production systems
Required skills
PythonRustC++GoLLMs
Nice-to-have skills
Hugging FaceLangChain
Domain expertise
aideeptech
Benefits & perks
equity
Tech stack
PythonRustC++GoLLMs