Member of Technical Staff, AMD GPU Performance Engineering
InferactAI Inference company
San Francisco, United States$200K - $400KLead
Andreessen Horowitz
Lightspeed Venture Partners
Sequoia Capital
Altimeter Capital
Redpoint Ventures
ZhenFund
Data & AI
About the role
TL;DR
Optimize vLLM for AMD GPUs to achieve high performance.
- •We're looking for an AMD GPU performance engineer to optimize vLLM for AMD GPUs.
- •Key Responsibilities Build and optimize AMD GPU backends, kernels, and runtime paths.
- •Improve performance-critical paths such as attention, GEMM, sampling, KV cache, and communication-heavy operations.
- •Develop benchmarking infrastructure using ROCm, HIP, Triton, CK, AITER, and related tools.
- •Requirements Bachelor's degree in computer science, engineering, systems, machine learning, or similar.
- •Hands-on experience optimizing AMD GPU workloads using ROCm, HIP, Triton, CK, AITER, or similar tools.
- •Deep understanding of AMD GPU execution, memory behavior, toolchains, kernel performance, and backend-specific performance constraints.
- •Experience optimizing ML kernels or inference paths such as attention, GEMM, sampling, KV cache, fused kernels, or communication-heavy runtime paths.
- •Strong performance profiling and benchmarking skills.
Required skills
MLflowKubeflowSageMakerVertex AIONNXJAX
Domain expertise
ai
Benefits & perks
equity, health insurance, 401(k) company match