XOC & Incident Management
FluidStackGPU Cloud company
Austin, United StatesSenior
Situational Awareness
Astro Capital (NY)
7GC & Co
Armyn Capital
Autopilot Management Company
Bare Metal Ventures
Operations & Strategy
About the role
TL;DR
Manage and optimize the incident management process for a large-scale data center operation.
- •Stand up and run the fleet operations center that watches every site 24/7: alarms, tickets, escalations, and communications.
- •Key Responsibilities Own the incident management process end to end, from first alert to postmortem, across facility and compute events.
- •Write the runbooks, escalation trees, and severity definitions the whole fleet operates on.
- •Drive incident metrics (time to acknowledge, time to resolve, repeat rate) down with process and tooling, not headcount.
- •Requirements You've run a NOC, GOC, or mission-control function and owned its performance numbers.
- •You've written incident processes that other people still use after you left.
- •You stay structured when several things break at once.
- •You write postmortems that change how the organization operates, not just what it apologizes for.
- •Bonus: Data center or utility operations center experience.
- •PagerDuty or ServiceNow-class tooling.
- •SRE-style incident frameworks.
Nice-to-have skills
PagerDuty
Tech stack
PagerDuty