Skip to content
Inferact logo

Member of Technical Staff, Inference

InferactAI Inference company
SingaporeS$200,000 - S$400,000Lead
Andreessen Horowitz logo
Andreessen Horowitz
Lightspeed Venture Partners logo
Lightspeed Venture Partners
Sequoia Capital logo
Sequoia Capital
Altimeter Capital logo
Altimeter Capital
Redpoint Ventures logo
Redpoint Ventures
ZhenFund
Data & AI

About the role

TL;DR

Optimize AI inference engines for large language models.

  • We're looking for an inference runtime engineer to push the boundaries of what's possible in LLM and diffusion model serving.
  • Key Responsibilities Optimize how models execute across diverse hardware and architectures.
  • Work at the core of vLLM, impacting how the world runs AI inference.
  • Implement model architectures and inference techniques from research papers.
  • Contribute performant and maintainable code and debug in complex ML codebases.
  • Requirements Bachelor's degree in computer science, engineering, or similar.
  • Deep understanding of transformer architectures and their variants.
  • Strong programming skills in Python with experience in PyTorch internals.
View original posting →

Required skills

PythonPyTorch

Domain expertise

ai

Benefits & perks

equity, medical, dental, vision

Tech stack

PythonPyTorch

Similar jobs

G

Senior Data Scientist - AI Research & Reliability

C

Software Engineer, Data & AI Platform

K

Sr. AI Systems & Solutions Architect, Customer Success