Skip to content
FuriosaAI logo

Sr. Software Engineer - Inference Engine (Platform Software)

FuriosaAIAI Accelerators company
Seoul, South KoreaSenior
Industrial Bank of Korea
Keistone Partners
Korea Development Bank
PI Partners
Kakao Investment
DSC Investment
Software Engineering

About the role

TL;DR

Develop and optimize a high-performance inference engine for LLMs on FuriosaAI NPUs.

  • Software Engineer (Inference Engine) is responsible for developing and optimizing a high-performance inference engine for Large Language Models (LLMs) and multimodal LLMs running on FuriosaAI NPUs.
  • You will proactively research and apply state-of-the-art inference optimization techniques to our inference engine, working closely with compiler and hardware teams.
  • Key Responsibilities Design and implement FuriosaAI’s next-generation inference engine for large and multimodal language models, optimized for throughput, latency, and memory efficiency.
  • Design and implement advanced inference optimizations such as speculative decoding, KV-cache management, and parallelism.
  • Design and develop capabilities for distributed and scalable inference.
  • Collaborate closely with the Compiler team to co-design and optimize execution for FuriosaAI NPUs.
  • Proactively research, evaluate, and integrate state-of-the-art inference optimization techniques.
  • Requirements BS degree in Computer Science, Engineering, or a related field, with at least 3 years of relevant industry experience.
  • Proficiency in Rust or C++ programming skill.
  • Knowledge and passion of deep learning, LLM, and/or generative AI models.
  • Excellent problem-solving and data analysis skills.
  • Strong communication and collaboration skills.
View original posting →

Required skills

RustC++LLMs

Nice-to-have skills

Python

Domain expertise

deeptechai

Tech stack

RustC++LLMs

Similar jobs

F

Forward Deployed Software Engineer - Intern

D

Senior Software Engineer, Data Infrastructure

A

Software Engineer I