DevOps / Platform Engineer
Ineffable IntelligenceReinforcement Learning company
London, United KingdomMid
Sequoia Capital
Lightspeed Venture Partners
NVIDIA
DST Global
Google
Index
Software Engineering
About the role
TL;DR
Define and maintain the infrastructure for a groundbreaking AI research platform.
- •Join us in our London office to shape the platform architecture and infrastructure for our superlearner mission.
- •Key Responsibilities Manage Kubernetes clusters and deploy internal tools Handle GPU scheduling at scale with KAI scheduler and Kueue Ensure large-scale infrastructure resilience and observability Requirements Experience with Kubernetes and containers Hands-on experience with GPU scheduling tools Strong background in cloud infrastructure, particularly Google Cloud
Required skills
KubernetesGoogle CloudPythonRust
Nice-to-have skills
GrafanaDatadog
Domain expertise
ai
Tech stack
KubernetesGoogle CloudGrafanaDatadogPythonRust