Lead SRE to ensure production reliability and scale healthcare AI platform.
•Heidi is seeking a Lead Site Reliability Engineer to own production systems, lead a small SRE team, and ensure the reliability and scalability of their healthcare AI platform.
•This role requires hands-on involvement in incident response, on-call duties, and system reliability.
•Key Responsibilities Lead on-call and incident response, ensuring clear communication and rapid service restoration.
•Improve operational reliability by identifying recurring issues and driving fixes through automation and process improvements.
•Own and operate Kubernetes clusters and cloud infrastructure, strengthening observability with dashboards and alerts.
•Reduce operational toil through automation and simplify runbooks, while supporting safe change management.
•Lead and grow the SRE team, including hiring, onboarding, and career development.
•Requirements 7+ years in SRE, DevOps, or platform engineering roles, with formal team leadership experience.
•Proven track record of hiring, coaching, and growing engineers.
•Deep experience supporting production systems, including on-call rotations and debugging live systems under pressure.
•Strong experience operating cloud infrastructure at scale (AWS preferred) and hands-on Kubernetes experience.
•Infrastructure as code (Terraform) and monitoring tools (Datadog, Prometheus) experience.