Staff Site Reliability Engineer
FilevineLegal Case company
RemoteLead
Insight Partners
Accel
Halo Fund
Meritech
Stepstone
Album Ventures
Software EngineeringNew
About the role
TL;DR
Senior technical authority shaping production systems and driving reliability strategy.
- •As a Staff Site Reliability Engineer, you will be the senior technical authority on the SRE team, shaping engineering culture and defining the technical standard for production systems.
- •You will bridge the gap between business goals and technical execution, with a forward-looking perspective on how AI and machine learning drive reliability.
- •Key Responsibilities Define and execute the technical strategy for Observability & Alerting and Platform Infrastructure.
- •Lead the evolution of reliable, scalable, secure, and efficient cloud platforms and distributed systems.
- •Champion SLIs, SLOs, error budgets, capacity planning, and automation.
- •Lead the organization through complex production incidents and drive permanent engineering improvements.
- •Build self-service platform capabilities to reduce toil and improve engineering velocity.
- •Requirements 12+ years of experience in software engineering, infrastructure, platform engineering, or SRE, with 6+ years in SRE and 3+ years leading complex technical initiatives.
- •Expert-level depth in observability and platform infrastructure, with broad expertise in incident response, capacity planning, automation, and reliability engineering.
- •Advanced experience with Kubernetes and an observability platform (e.g., New Relic, Datadog).
- •Strong software-engineering ability in Python, Go, Bash, or similar, with experience building production tooling and automation.
- •Proven ability to mentor engineers and communicate technical risk clearly.
Required skills
PythonGoBashKubernetesNew RelicDatadogTerraform
Nice-to-have skills
CloudFormationPulumiAWSGoogle CloudAzureCI/CDDockerLinux
Domain expertise
legaltechaiproductivity
Tech stack
PythonGoBashKubernetesNew RelicDatadogTerraform