Skip to content
M

Sr. Site Reliability Engineer

MeridianlinkDigital Lending company
RemoteSenior
Centerbridge Partners
Silversmith Capital Partners
Avery Dennison
Serent Capital
Software Engineering

About the role

TL;DR

Owns reliability, scalability, and observability of financial SaaS applications and infrastructure.

  • We are seeking a Senior Site Reliability Engineer to join our cloud engineering team.
  • You will own the reliability, scalability, and observability of our critical financial SaaS applications and infrastructure, working across cloud platforms to ensure our customers experience is seamless, secure, and performant services.
  • This is a high-impact role for someone who is passionate about building resilient systems and preventing outages before they happen.
  • Key Responsibilities Design, implement, and maintain Service Level Objectives (SLOs) and Service Level Indicators (SLIs) across all critical systems; ensure we meet or exceed targets consistently Lead observability strategy by designing comprehensive monitoring, logging, and tracing architectures; select and deploy observability tools that provide deep visibility into system behavior Build and own runbooks, incident response procedures, and post-incident review processes; mentor the team on incident management and blameless postmortems Architect and deploy cloud infrastructure on AWS or Azure; implement infrastructure-as-code practices and ensure high availability, disaster recovery, and business continuity Develop automation and AIOps capabilities to reduce toil, accelerate incident detection, and enable self-healing systems; implement intelligent alerting to minimize false positives Requirements 7+ years in Site Reliability Engineering, DevOps, platform engineering, or closely related roles with significant responsibility for production systems Expert-level experience with Azure or AWS (or both); deep knowledge of compute, networking, storage, and managed services; experience managing infrastructure at scale Demonstrated expertise in observability: designing and implementing monitoring, alerting, logging, and distributed tracing solutions; hands-on with observability platforms (e.g., Prometheus, Grafana, ELK, Datadog, New Relic, or similar) Strong background in SLOs, SLIs, and SLAs; experience defining meaningful objectives and building systems to meet them; understanding of error budgets and their role in prioritization Proficiency in Python, PowerShell, bash, etc. scripting languages for production automation, tooling, and systems programming; ability to write clean, maintainable code for operational workflows
View original posting →

Required skills

PythonAWSAzurePrometheusGrafanaDatadogNew RelicTerraformDockerKubernetesCI/CDLinuxBash

Nice-to-have skills

CloudFormationAnsible

Domain expertise

fintech

Tech stack

PythonAWSAzurePrometheusGrafanaDatadogNew RelicTerraformCloudFormationAnsibleDockerKubernetesCI/CDLinux

Similar jobs

M

Senior Site Reliability Engineer

I

Site Reliability Engineer II

R

Avionics Automation Test Engineer II