Owns the major incident lifecycle end to end, driving response and process improvement.
•We are hiring a Staff Incident Manager to own the major incident lifecycle end to end.
•You are the person who takes control when something is broken in production, brings the right people together, drives the response to resolution, and makes sure we are measurably better after every incident than we were before it.
•Key Responsibilities Act as incident commander on major and critical incidents, owning coordination, decision-making cadence, and escalation from detection through resolution.
•Drive down time to detect, time to engage, and time to recover; own MTTR as a headline metric.
•Run rotations, escalation policies, paging hygiene, and alert quality for the on-call program.
•Own internal stakeholder alignment during an event and drive clear, accurate merchant-facing updates.
•Run blameless postmortems and ensure action items are concrete, owned, and tracked to closure.
•Requirements Proven experience running major incident response in a production environment, ideally as an incident commander or in a dedicated incident management function.
•Strong working knowledge of modern distributed systems and cloud-native operations.
•Experience defining or maturing an incident management practice from the ground up.
•Fluency with observability and incident tooling (metrics, logging, tracing, alerting, paging).
•Hands-on experience with Datadog is strongly preferred; familiarity with a paging platform and a status page tool is expected.
•Excellent written and verbal communication skills.
•Calm, decisive judgment during high-pressure events.