Software Engineer, Model Runtime
OpenAIGenerative AI company
San Francisco, United StatesSenior
Microsoft
Nvidia
Thrive Capital
Sequoia Capital
SoftBank
Amazon
Software EngineeringNew
About the role
TL;DR
Build and optimize the LLM inference runtime for custom AI silicon.
- •Build the model runtime within the inference engine that executes complex, frontier models at scale on OpenAI’s custom silicon.
- •The runtime will translate inference workloads into efficient execution, optimizing for throughput, latency, utilization, and reliability.
- •Key Responsibilities Design and implement the LLM inference runtime for frontier models running on custom silicon.
- •Build scheduling, continuous batching, memory management, KV-cache management, and execution orchestration.
- •Develop distributed execution strategies across chips, hosts, and racks.
- •Optimize end-to-end latency, throughput, memory efficiency, and hardware utilization.
- •Partner with kernel, compiler, architecture, and silicon teams to co-design interfaces.
- •Requirements Strong systems programming experience in C++, Rust, Python, or comparable performance-oriented environments.
- •Experience building or optimizing runtimes, distributed systems, compilers, kernels, or model-serving infrastructure.
- •Understanding of modern LLM inference, including prefill and decode behavior, batching, KV-cache tradeoffs, and model parallelism.
- •Ability to reason quantitatively about performance metrics.
- •Comfort profiling and debugging performance across multiple layers of a hardware-software stack.
Required skills
C++RustPythonLLMs
Nice-to-have skills
KubernetesDocker
Domain expertise
ai
Tech stack
C++RustPythonLLMsKubernetesDockerGit