Lead reliability strategy and architecture for AI agent infrastructure on AWS.
•Yuno is seeking a Staff Site Reliability Engineer to define the technical direction for reliability across its AI-native operating system for global commerce.
•This role involves evolving the architecture of a production platform that provisions, deploys, and manages AI agents at scale on AWS.
•Key Responsibilities Define reliability strategy, including SLO culture, error-budget policy, and incident practices.
•Drive architectural decisions for the AI agent infrastructure and messaging layer.
•Own cloud infrastructure, automate provisioning with IaC, and ensure platform scalability.
•Build comprehensive observability, monitoring, tracing, and alerting systems.
•Lead incident response, conduct blameless postmortems, and mentor other engineers.
•Requirements Experience with event-driven architecture and messaging systems (Kafka, NATS, RabbitMQ).