Senior Software Engineer, SRE / Observability Tooling
GliaDigital Customer company
RemoteSenior
Insight Partners
Entrepreneurs Roundtable Accelerator
Tola Capital
Wildcat Capital Management
RingCentral Ventures
Grassy Creek
Software Engineering
About the role
TL;DR
Develop and maintain observability and SRE tooling for Glia's cloud-native infrastructure.
- •As a Site Reliability Engineer on the Observability Team, you will focus on building SRE and observability tooling for Glia's cloud-native infrastructure.
- •Key Responsibilities Develop standards, infrastructure, and automation for dashboards, alerts, and monitors as code.
- •Partner with development teams to establish production readiness and operational readiness.
- •Build tooling and templates for defining, measuring, and reporting on Service Level Objectives (SLOs) and Service Level Indicators (SLIs).
- •Develop tooling to automate observability and operational workflows, eliminating manual toil for engineering teams.
- •Build and improve incident response tooling and workflows to help teams resolve outages faster and learn from them.
- •Requirements Expert-level proficiency with AWS and Kubernetes (EKS), particularly in areas of observability, networking, and auto-scaling.
- •Experience with modern observability platforms (e.g., DataDog, Prometheus) and a deep understanding of metrics, logging, and tracing.
- •Deep, practical understanding of Site Reliability Engineering (SRE) principles (SLOs, error budgets, toil reduction).
- •Demonstrable experience analyzing and troubleshooting large-scale distributed systems.
- •Strong software development skills in a language like Python or Go, used to build operational tools, services, or automation.
- •Expertise in designing and operating robust CI/CD pipelines for a microservices architecture (e.g., using ArgoCD, Github Actions, Helm).
- •A systematic, data-driven approach to problem-solving and root cause analysis.
- •Proficiency in using AI tools thoughtfully, maintaining ownership of the final output while recognizing the tools' limitations.
Required skills
AWSKubernetesDatadogCI/CDPythonGoArgoCDGitHub ActionsHelm
Nice-to-have skills
PrometheusNode.jsJavaScriptReactJavaA/B TestingProduct Analytics
Domain expertise
fintech
Tech stack
AWSKubernetesRabbitMQDatadogGitHub ActionsArgoCDJenkinsHelmTerraformPythonNode.jsJavaScriptReactJavaGo