Member of Technical Staff, Developer Relations
InferactAI Inference company
San Francisco, United States$200K - $400KLead
Andreessen Horowitz
Lightspeed Venture Partners
Sequoia Capital
Altimeter Capital
Redpoint Ventures
ZhenFund
Marketing
About the role
TL;DR
Help developers understand and build with vLLM for AI inference.
- •We're looking for a Developer Relations Engineer to help make vLLM the default way developers understand, build, and scale AI inference.
- •Key Responsibilities Write technical deep dives, build demos, create tutorials, contribute to docs and examples, host workshops, and help developers understand topics like KV cache, continuous batching, prefix caching, prefill and decode, quantization, GPU serving, latency versus throughput, and model-server tradeoffs across vLLM and adjacent systems.
- •Shape how the broader AI infrastructure community learns, adopts, and builds with vLLM.
- •Create public artifacts that help practitioners build better systems.
- •Requirements Bachelor's degree or equivalent experience in computer science, engineering, machine learning, systems, or similar.
- •Strong technical understanding of LLM inference systems, model serving, GPU inference, distributed runtimes, scheduling, batching, quantization, or related infrastructure.
- •Ability to credibly explain systems concepts such as KV cache, PagedAttention, continuous batching, prefill / decode scheduling, prefix caching, speculative decoding, tensor parallelism, data parallelism, or latency versus throughput tradeoffs.
Required skills
LLMs
Domain expertise
ai
Benefits & perks
equity, health insurance, 401(k) company match
Tech stack
PyTorch