Set reliability architecture for business and manufacturing operations systems.
•Anduril Industries is a defense technology company focused on transforming military capabilities with advanced technology.
•The CorpTech Platform team builds the internal engineering foundations that power Anduril's corporate and hardware systems, driving towards an autonomous enterprise through AI integration in the engineering lifecycle.
•Key Responsibilities Set the reliability architecture for production environments, including observability, deployment systems, incident management, and capacity planning.
•Design and operate the observability platform for real production visibility at scale.
•Own deployment infrastructure and release-safety mechanisms to enable fast, stable shipping.
•Define and govern SLO frameworks to make reliability measurable and actionable.
•Identify systemic reliability risks and drive infrastructure investments to eliminate failure classes.
•Requirements 10+ years of experience in site reliability engineering, production engineering, or infrastructure engineering at an architecture or platform-wide scope.
•Demonstrated experience designing and owning reliability infrastructure used by multiple engineering teams.
•Deep technical fluency across distributed systems, container orchestration (Kubernetes), cloud platforms (AWS, GCP, or Azure), networking, and storage.
•Proficiency in systems programming languages like Go, Python, or Rust.
•Experience defining SRE standards and influencing adoption across engineering teams.