Algorithm - Serving System Researcher
FuriosaAIAI Accelerators company
Seoul, South KoreaMid
Industrial Bank of Korea
Keistone Partners
Korea Development Bank
PI Partners
Kakao Investment
DSC Investment
Data & AI
About the role
TL;DR
Research methods to improve NPU-based inference roofline and explore challenging projects for product application.
- •Our team researches methods to improve the roofline of NPU-based inference.
- •We define and explore challenging projects that can be applied to products within 6 months to 2 years, and complete the productization stage by collaborating with SW organizations.
- •AI datacenters are at a major turning point.
- •Innovations such as cluster-level serving that connects hundreds to thousands of chips, AFD (Attention-FFN Disaggregation), and heterogeneous computing systems are rapidly emerging.
- •These systems are costly to implement in reality.
- •We believe that the ability to explore the design space with analytical modeling and simulation and verify its feasibility is the core of this field, and we are currently focusing on it.
- •The Serving System Researcher is responsible for the entire research cycle from initial setup to final verification.
- •The focus is on modeling and analysis, and the position is in charge of modeling and analysis.
- •We select the necessary tools on the fly, not tied to specific codebases or frameworks—while focusing on roofline analysis of the simulator, there is also implementation on the NPU side.
- •The task of defining this position is not a question: "On which workload does this serving system excel? What is needed to prove it most quickly and efficiently?" Responsibilities Identify and define research tasks for the full cycle of cluster-level serving systems, heterogeneous computing systems, and other advanced NPU serving technologies.
- •Explore the design space of serving architectures and algorithms through analytical modeling, roofline analysis, cost models, and simulators, quantitatively evaluate performance and cost, and provide foundations for key decision-making.
- •Design serving algorithms such as scheduling, batching, KV cache management, and disaggregation, verify them through modeling/simulation and NPU-based implementation, and develop promising ideas into products through POC/prototypes in collaboration with SW organizations.
- •Minimum Qualifications Master's degree or equivalent experience in AI/ML, Computer Architecture, or related fields.
- •Understanding of LLM inference dynamics (attention, KV cache, prefill/decode, batching, multi-chip parallelism, etc.) and performance characteristics.
- •Ability to perform performance modeling, simulation, and data-driven analysis using Python, PyTorch, etc.
- •Ability to structure ambiguous problems and lead the modeling-verification cycle proactively.
- •Ability to communicate analysis results and technical decisions clearly through documents and presentations.
- •Preferred Qualifications Experience in performance analysis/optimization of distributed systems, cluster-level software, or large-scale workloads.
- •Understanding of the internal architecture of LLM serving systems such as vLLM, SGLang, TensorRT-LLM.
- •Experience in quantitative system evaluation such as roofline analysis, analytical performance modeling, and TCO analysis.
- •Experience in optimization or development of accelerator-based systems such as GPU/NPU.
- •Familiarity with AI development tools. 이런 환경에서 일하게 됩니다 To help you determine if this position is a good fit, we'll explain our way of working.
- •Uncertainty is fundamental.
- •We embrace the unknown future technology.
- •We also consider the reason why "it doesn't work" as a valuable research result.
- •The nature of the work changes continuously.
- •Reading papers, understanding hundreds of use cases, simulator coding, actual implementation, and creating presentation materials are all tasks, and the scope of work extends from computer architecture to serving algorithms, data centers, and even NPU implementation.
- •Rather than a specific tool or framework, continuous learning about the problem itself is a driving force.
- •Research is not just about reporting.
- •Our analysis and POCs serve as the basis for product decision-making.
- •Contact [email protected]
Required skills
PythonPyTorchLLMs
Domain expertise
ai
Tech stack
PythonPyTorch