Skip to content
Inferact logo

Member of Technical Staff, AMD GPU Performance Engineering

InferactAI Inference company
San Francisco, United States$200K - $400KLead
Andreessen Horowitz logo
Andreessen Horowitz
Lightspeed Venture Partners logo
Lightspeed Venture Partners
Sequoia Capital logo
Sequoia Capital
Altimeter Capital logo
Altimeter Capital
Redpoint Ventures logo
Redpoint Ventures
ZhenFund
Data & AI

About the role

TL;DR

Optimize vLLM for AMD GPUs to achieve high performance.

  • We're looking for an AMD GPU performance engineer to optimize vLLM for AMD GPUs.
  • Key Responsibilities Build and optimize AMD GPU backends, kernels, and runtime paths.
  • Improve performance-critical paths such as attention, GEMM, sampling, KV cache, and communication-heavy operations.
  • Develop benchmarking infrastructure using ROCm, HIP, Triton, CK, AITER, and related tools.
  • Requirements Bachelor's degree in computer science, engineering, systems, machine learning, or similar.
  • Hands-on experience optimizing AMD GPU workloads using ROCm, HIP, Triton, CK, AITER, or similar tools.
  • Deep understanding of AMD GPU execution, memory behavior, toolchains, kernel performance, and backend-specific performance constraints.
  • Experience optimizing ML kernels or inference paths such as attention, GEMM, sampling, KV cache, fused kernels, or communication-heavy runtime paths.
  • Strong performance profiling and benchmarking skills.
View original posting →

Required skills

MLflowKubeflowSageMakerVertex AIONNXJAX

Domain expertise

ai

Benefits & perks

equity, health insurance, 401(k) company match

Similar jobs

C

Software Engineer, Data & AI Platform

K

Sr. AI Systems & Solutions Architect, Customer Success

A

Research Scientist, All-Optical Working Memory