Member of Technical Staff, Inference
InferactAI Inference company
SingaporeS$200,000 - S$400,000Lead
Andreessen Horowitz
Lightspeed Venture Partners
Sequoia Capital
Altimeter Capital
Redpoint Ventures
ZhenFund
Data & AI
About the role
TL;DR
Optimize AI inference engines for large language models.
- •We're looking for an inference runtime engineer to push the boundaries of what's possible in LLM and diffusion model serving.
- •Key Responsibilities Optimize how models execute across diverse hardware and architectures.
- •Work at the core of vLLM, impacting how the world runs AI inference.
- •Implement model architectures and inference techniques from research papers.
- •Contribute performant and maintainable code and debug in complex ML codebases.
- •Requirements Bachelor's degree in computer science, engineering, or similar.
- •Deep understanding of transformer architectures and their variants.
- •Strong programming skills in Python with experience in PyTorch internals.
Required skills
PythonPyTorch
Domain expertise
ai
Benefits & perks
equity, medical, dental, vision
Tech stack
PythonPyTorch