Skip to content
RadixArk logo

Member of Technical Staff — Inference-Core Engine

RadixArkAI Infrastructure company
Palo Alto, United States$200,000 - $400,000 USDLead
Accel logo
Accel
Spark Capital logo
Spark Capital
NVentures logo
NVentures
Salience Capital
A&E Investments
HOF Capital
Data & AI

About the role

TL;DR

Design and optimize large-scale AI inference systems for frontier models.

  • RadixArk is seeking a Member of Technical Staff
  • Inference to push the limits of large-scale AI inference.
  • You will work on the core systems that serve frontier models at scale, optimizing performance, latency, throughput, and cost across thousands of GPUs.
  • Key Responsibilities Design and build large-scale inference systems for frontier AI models Optimize latency, throughput, and GPU utilization in production inference Develop and improve model serving architectures and runtimes Work on batching, scheduling, and memory management strategies Collaborate with kernel, compiler, and systems teams on performance optimization Requirements 5+ years of experience in systems engineering, ML infrastructure, or performance-critical backend systems Strong expertise in large-scale inference systems for LLMs or generative models Deep understanding of GPU architecture and performance characteristics Experience optimizing latency
  • and throughput-critical production systems Proficiency in Python, Rust, C++, or Go for production systems
View original posting →

Required skills

PythonRustC++GoLLMs

Nice-to-have skills

Hugging FaceLangChain

Domain expertise

aideeptech

Benefits & perks

equity

Tech stack

PythonRustC++GoLLMs

Similar jobs

G

AI Interaction Evaluator (Codex / Claude Code, up to $200/hr)

S

Software Engineer II — Agentic AI Foundations

B

Senior AI Engineer, Agentic Data Enrichment