Member of Technical Staff, Inference
MirendilAI Research company
San Francisco, United States$300K - $500KLead
Andreessen Horowitz
Kleiner Perkins
Nvidia
Data & AI
About the role
TL;DR
Engineer to optimize and manage inference systems for AI models.
- •We are looking for an engineer to own the inference systems that power our models in production and research.
- •Key Responsibilities Design and build high-throughput, low-latency inference serving systems for frontier models Optimize inference performance across GPU and accelerator hardware Enable and extend distributed inference frameworks to support novel architectures Implement and validate inference-time optimizations Build observability and reliability infrastructure Partner directly with teams to bring new model architectures into production Requirements Strong background in machine learning and AI Experience with inference optimization and distributed systems Proficiency in Python and deep learning frameworks Knowledge of cloud infrastructure and containerization Excellent problem-solving and communication skills
Required skills
PythonTensorFlowPyTorchKubernetesAWSDocker
Domain expertise
ai
Benefits & perks
Equity
Tech stack
PythonTensorFlowPyTorchKubernetesAWSDocker