Sets reliability architecture for business and manufacturing systems.
•Anduril Industries is a defense technology company transforming military capabilities with advanced technology.
•This Staff SRE role sets the reliability architecture for the systems that run Anduril’s business and manufacturing operations, focusing on proactive reliability engineering rather than ticket response.
•Key Responsibilities Set the reliability architecture for production environments, including observability, deployment systems, incident management, and capacity planning.
•Design and operate the observability platform for production visibility at scale.
•Own deployment infrastructure and release-safety mechanisms for fast, stable shipping.
•Define and govern SLO frameworks to make reliability measurable and actionable.
•Identify systemic reliability risks and drive infrastructure investments to eliminate failure classes.
•Requirements 10+ years of experience in SRE, production engineering, or infrastructure engineering, with architecture/platform-wide scope.
•Experience designing and owning reliability infrastructure used by multiple engineering teams.
•Deep technical fluency in distributed systems, Kubernetes, cloud platforms (AWS, GCP, Azure), networking, and storage.
•Proficiency in systems programming languages like Go, Python, or Rust.
•Experience defining SRE standards and influencing adoption across engineering teams.