Senior Site Reliability Engineer
SanityContent Management company
RemoteSenior
ICONIQ Growth
Bullhound Capital
Heavybit
Threshold
Shopify
Blue Cloud Ventures
Software Engineering
About the role
TL;DR
Build and operate scalable, reliable infrastructure for a leading AI Content Operating System.
- •Sanity is seeking a Senior Site Reliability Engineer to ensure their AI Content Operating System platform is scalable, fast, safe, and inspiring to use.
- •This role involves partnering with development teams to build and operate robust infrastructure.
- •Key Responsibilities Design, build, and operate shared platform foundations including GCP infrastructure, Kubernetes, networking, and CI/CD.
- •Diagnose and troubleshoot complex distributed systems at high request volume.
- •Ensure observability and analyze the behavior of the platform stack.
- •Improve on-call rotations and incident response processes.
- •Mentor engineers and raise the technical bar through code and design reviews.
- •Requirements 5+ years of experience as part of an SRE on-call rotation.
- •Experience with Kubernetes for orchestrating containerized applications.
- •Experience building CI/CD pipelines and using observability stacks (e.g., Prometheus).
- •Analytical approach to designing, diagnosing, and optimizing infrastructure.
- •Experience managing scalable, highly available cloud-based applications.
Required skills
KubernetesGoogle CloudCI/CDPrometheusElasticsearchPostgreSQLBashLinuxDocker
Domain expertise
developer-toolsai
Benefits & perks
Comprehensive health plans, Stock options
Tech stack
KubernetesPrometheusElasticsearchPostgreSQLGoogle CloudCI/CDDockerBashLinux