Member of Technical Staff, Kernel Engineering
InferactAI Inference company
Singapore, SingaporeS$200,000 - S$400,000 annuallyLead
Andreessen Horowitz
Lightspeed Venture Partners
Sequoia Capital
Altimeter Capital
Redpoint Ventures
ZhenFund
Software Engineering
About the role
TL;DR
Optimize vLLM's inference engine for maximum performance on various accelerators.
- •We're looking for a performance engineer to optimize vLLM's inference engine.
- •You'll write kernels and low-level optimizations for various accelerators.
- •Key Responsibilities Write CUDA kernels and optimizations.
- •Collaborate with hardware vendors for new chip integration.
- •Ensure maximum performance extraction from hardware.
- •Requirements Bachelor's degree in computer science or equivalent experience.
- •Deep experience with CUDA kernels and GPU architecture.
- •Proficiency in C++ and Python.
Required skills
C++Python
Domain expertise
ai
Benefits & perks
equity, medical, dental, vision
Tech stack
C++Python