Skip to content
FuriosaAI logo

Algorithm - Serving System Engineer

FuriosaAIAI Accelerators company
Seoul, South KoreaSenior
Industrial Bank of Korea
Keistone Partners
Korea Development Bank
PI Partners
Kakao Investment
DSC Investment
Data & AI

About the role

TL;DR

Implement and optimize NPU-based AI inference serving systems.

  • Our team researches methods to improve the roofline of NPU-based inference.
  • We define and explore challenging tasks that can be applied to products within 6 months to 2 years, and complete the productization phase in collaboration with the SW organization.
  • AI datacenters are currently at a major turning point.
  • New serving systems are rapidly emerging, going beyond existing ones that connect hundreds to thousands of chips, such as cluster-level serving, AFD (Attention-FFN Disaggregation), and heterogeneous computing systems.
  • Furthermore, as the era of Agentic AI begins, multi-turn sessions become fundamental, and methods to maximize KV cache reuse and compression by utilizing diverse memory hierarchies are actively proposed.
  • These ideas, while verifiable through analysis and simulation, are not confirmed until they are proven through POCs.
  • The Serving System Engineer is responsible for that proof.
  • We directly implement the concept of a state-of-the-art serving system on top of the NPU, utilize the kernel programming stack and low-level programming interfaces to maximize the hardware's performance, and verify what is and isn't possible.
  • The implementation is not the entire scope of this role.
  • The lessons learned from the final stage of proof to the actual implementation are the source of new research tasks and the best starting point, and we are responsible for proposing the next task as the culmination of that.
  • Key Responsibilities Demonstrate the core elements of the state-of-the-art serving system (AFD, KV cache reuse, etc.) that the team researches into NPU-based implementation (POC).
  • Explore and implement the latest compression techniques (quantization, KV cache compression, etc.) and verify them in the NPU environment.
  • Based on the implementation, iterate on new research tasks and propose optimization ideas.
  • Requirements 3+ years of relevant field practical experience, or equivalent experience (research/project experience).
  • Understanding of LLM inference operation principles (attention, KV cache, prefill/decode, batching, etc.).
  • Knowledge or experience with AI inference systems such as vLLM, SGLang, TensorRT-LLM.
  • Experience with CUDA, Triton, or accelerator programming.
  • Experience in solving problems without documented solutions and debugging them.
  • Ability to communicate clearly and lead collaboration with the team.
View original posting →

Required skills

PythonLLMsC++RustBash

Nice-to-have skills

Microservices

Domain expertise

ai

Tech stack

PythonC++RustBashLLMs

Similar jobs

D

Senior Technical Specialist, AI Solutions

H

ML Engineer (Data Engine)

H

ML Systems Performance Engineer (MFU)