Member of Technical Staff, TPU Performance Engineering
InferactAI Inference company
SingaporeS$200,000 - S$400,000 annuallyLead
Andreessen Horowitz
Lightspeed Venture Partners
Sequoia Capital
Altimeter Capital
Redpoint Ventures
ZhenFund
Data & AI
About the role
TL;DR
Optimize vLLM for TPU inference performance.
- •We're looking for a TPU performance engineer to make vLLM a first-class inference engine on Google TPUs.
- •Key Responsibilities Build and optimize TPU backends, compiler integrations, runtime paths, and benchmarking infrastructure.
- •Work at the boundary of inference systems, kernels, compilers, and hardware architecture.
- •Improve production-relevant model serving on TPU with clear correctness, latency, and throughput benchmarks.
- •Requirements Bachelor's degree or equivalent experience in computer science, engineering, systems, machine learning, or similar.
- •Hands-on experience building or optimizing TPU workloads using JAX, XLA, Pallas, or related compiler and runtime tooling.
- •Deep understanding of TPU execution, memory behavior, compilation, and performance constraints for ML workloads.
Required skills
JAXMLflowKubeflowVertex AIONNX
Domain expertise
ai
Benefits & perks
equity, medical, dental, vision
Tech stack
JAXMLflowKubeflowVertex AIONNX