Member of Technical Staff, Exceptional Generalist
InferactAI Inference company
RemoteLead
Andreessen Horowitz
Lightspeed Venture Partners
Sequoia Capital
Altimeter Capital
Redpoint Ventures
ZhenFund
Data & AI
About the role
TL;DR
Work across the entire vLLM stack to optimize AI inference.
- •We're seeking exceptional generalist engineers who can work across the entire vLLM stack: from low-level GPU kernels to high-level distributed systems.
- •Key Responsibilities Work asynchronously with our San Francisco headquarters while maintaining full ownership of critical infrastructure.
- •Optimize CUDA kernels, design distributed orchestration systems, and implement new model architectures.
- •Directly impact how the world runs AI inference.
- •Requirements Bachelor's degree or equivalent experience in computer science, engineering, or similar.
- •Demonstrated ability to work autonomously and drive projects to completion without close supervision.
- •Excellent asynchronous communication skills and ability to collaborate effectively across time zones.
- •Strong track record of shipping high-impact work in complex technical environments.
- •Deep expertise in at least one of: systems programming, GPU/accelerator programming, distributed systems, or ML infrastructure.
Required skills
Python
Nice-to-have skills
RustGoC++Kubernetes
Domain expertise
ai
Benefits & perks
Competitive compensations (salary + equity), Visa sponsorship, Competitive benefits
Tech stack
RustGoC++PythonPyTorchKubernetes