TL;DR
Integrate AI accelerators with inference frameworks and build the runtime stack.
- •Integrate Fractile's AI acceleration hardware with leading inference engines and build the runtime stack.
- •Key Responsibilities Integrate Fractile's AI acceleration hardware with leading inference engines including vLLM and SGLang Research KV cache management technologies and build proof-of-concept implementations Work closely with the runtime team to design and build a scalable, bare-bones reference inference engine Focus primarily on the transformer ML architecture Share your expertise to help shape the direction of our runtime stack Requirements Solid experience with ML inference at scale, including multi-user serving A deep understanding of paged attention and inference engines such as vLLM Familiarity with key components of the ML software ecosystem Strong software engineering skills and an instinct for clean, maintainable systems
View original posting →