Staff Site Reliability Engineer (AI Platform)
ManychatChat Marketing company
Amsterdam, NetherlandsLead
Software Engineering
About the role
TL;DR
Own AI infrastructure reliability, performance, and cost, shaping its architecture from the ground up.
- •Manychat is seeking a Staff Site Reliability Engineer to own the reliability, performance, and cost of its AI infrastructure.
- •This role involves shaping the architecture of a young AI platform from the ground up, focusing on AI-native production concerns.
- •Key Responsibilities Own reliability and performance of AI infrastructure, including AI Gateway and inference services.
- •Design and evolve the AI Gateway for routing, failover, rate limiting, and caching.
- •Build observability for AI systems with SLOs per model and provider.
- •Drive cost optimization and FinOps for AI workloads.
- •Scale AI expertise across the organization by setting standards and coaching teams.
- •Requirements 5+ years in SRE/platform/infrastructure engineering with production ownership at scale.
- •Hands-on experience operating LLM-backed systems in production.
- •Deep cloud-native background (AWS, Kubernetes, Terraform/IaC, CI/CD).
- •Strong observability practice and experience defining SLOs for non-deterministic systems.
- •Proven cost-optimization work.
Required skills
KubernetesAWSTerraformCI/CDDockerPrometheusGrafanaOpenAI APIAnthropic APILLMsPythonGit
Domain expertise
aideveloper-tools
Benefits & perks
Hybrid onboarding, Relocation support, Comprehensive health insurance, Professional development budget, Flexible benefits package, Hybrid work, Generous, flexible time off, In-office perks (free meals and snacks), Company-funded sport activities, Annual offsites and team-building events
Tech stack
PythonAWSKubernetesTerraformCI/CDPrometheusGrafanaDockerOpenAI APIAnthropic APILLMsGit