Senior CloudOps Engineer
CloudZeroCloud Cost company
Boston, United StatesSenior
Matrix Partners
BlueCrest Capital Management
Innovius Capital
Threshold Ventures
Underscore VC
G20 Ventures
Software Engineering
About the role
TL;DR
Own reliability, performance, and observability of critical systems at scale.
- •CloudZero is seeking a Senior Site Reliability Engineer to own the reliability, performance, and observability of critical systems, empowering engineering teams to ship features that help customers understand and optimize their cloud spend.
- •This role involves real infrastructure work at scale, focusing on engineering solutions rather than firefighting.
- •Key Responsibilities Own the reliability practice for CloudZero's real-time ingestion path, including SLOs and architectural improvements.
- •Build observability into systems to proactively identify and debug issues.
- •Develop reliability tooling, production Python code, and maintain CloudFormation/SAM modules.
- •Automate deployments, scaling, backups, and limit changes.
- •Partner with Product Engineering to design resilient services and optimize for cost and performance.
- •Requirements Strong production Python experience.
- •Experience operating asynchronous, event-driven systems (Kafka, Kinesis, SQS, etc.). 5+ years of experience building and operating distributed systems in AWS.
- •Infrastructure as Code experience (CloudFormation, SAM, Terraform, or Pulumi).
- •Hands-on experience instrumenting systems in monitoring tools (Sumo Logic, Datadog, Prometheus, Splunk).
Required skills
PythonAWSCloudFormationApache KafkaDatadogPrometheusGemini API
Nice-to-have skills
AzureGoogle CloudTerraformPulumiAnthropic APIGitHub Actions
Domain expertise
developer-toolsai
Tech stack
PythonAWSAzureGoogle CloudCloudFormationDatadogPrometheusAnthropic APIGemini APIGitHub Actions