Define and implement SRE practices to enhance reliability and reduce outages.
•We're hiring our first Site Reliability Engineer to join the central DevOps team and bring SRE principles to Talkiatry's engineering organization.
•Key Responsibilities Define and roll out an SRE practice for a six-team organization: SLOs/SLIs, error budgets, and reliability standards that teams genuinely adopt.
•Build and improve observability—metrics, logging, distributed tracing, dashboards, and alerting—so that more incidents are detected by monitoring before anyone outside engineering notices.
•Drive down outage frequency by surfacing systemic reliability risks and partnering with teams to remediate them at the root.
•Reduce toil through automation, infrastructure-as-code, and self-service tooling that teams can own and extend themselves.
•Own the health and usability of our observability tooling, providing documentation and training where necessary.
•Requirements 7+ years in software or infrastructure engineering, with substantial hands-on SRE or production reliability experience.
•A track record of reducing incidents and improving detection—the outcomes this role is judged on.
•Hands-on experience defining SLOs/SLIs and using error budgets to guide engineering decisions.
•Deep observability expertise across metrics, logging, tracing, and alerting (e.g., Datadog, Prometheus, Grafana, or similar).
•Strong experience operating production systems on AWS.