Sr. Software Engineer - Inference Engine (Platform Software)
FuriosaAIAI Accelerators company
Seoul, South KoreaSenior
Industrial Bank of Korea
Keistone Partners
Korea Development Bank
PI Partners
Kakao Investment
DSC Investment
Software Engineering
About the role
TL;DR
Develop and optimize a high-performance inference engine for LLMs on FuriosaAI NPUs.
- •Software Engineer (Inference Engine) is responsible for developing and optimizing a high-performance inference engine for Large Language Models (LLMs) and multimodal LLMs running on FuriosaAI NPUs.
- •You will proactively research and apply state-of-the-art inference optimization techniques to our inference engine, working closely with compiler and hardware teams.
- •Key Responsibilities Design and implement FuriosaAI’s next-generation inference engine for large and multimodal language models, optimized for throughput, latency, and memory efficiency.
- •Design and implement advanced inference optimizations such as speculative decoding, KV-cache management, and parallelism.
- •Design and develop capabilities for distributed and scalable inference.
- •Collaborate closely with the Compiler team to co-design and optimize execution for FuriosaAI NPUs.
- •Proactively research, evaluate, and integrate state-of-the-art inference optimization techniques.
- •Requirements BS degree in Computer Science, Engineering, or a related field, with at least 3 years of relevant industry experience.
- •Proficiency in Rust or C++ programming skill.
- •Knowledge and passion of deep learning, LLM, and/or generative AI models.
- •Excellent problem-solving and data analysis skills.
- •Strong communication and collaboration skills.
Required skills
RustC++LLMs
Nice-to-have skills
Python
Domain expertise
deeptechai
Tech stack
RustC++LLMs