Skip to content
AlayaCare logo

Senior Site Reliability Specialist (SRE)

AlayaCareHome Health company
Montreal, CanadaSenior
Inovia Capital
Generation Investment Management
Klass Capital
CDPQ
Investissement Québec
BDC Capital
Software Engineering

About the role

TL;DR

Senior SRE responsible for scaling cloud infrastructure, improving reliability, and developing tooling.

  • We are seeking a Senior Site Reliability Specialist to join our SRE team.
  • Reporting to the Engineering Manager, you is responsible for scaling AWS cloud infrastructure, evolving Kubernetes deployment pipelines, improving monitoring, alerting, and resiliency, and developing tooling that enables product teams to deliver safely and efficiently.
  • For acquired Azure-based products the focus is on monitoring, alert triage, and runbook-driven incident response rather than greenfield platform design.
  • This role owns shared platform services across cloud regions, including databases, messaging, logging, search, and tenant provisioning.
  • The Senior SRE is expected to lead major infrastructure initiatives and proof-of-concept efforts, contribute to technical planning and prioritization, as well as partners with the Product teams to reduce operational incidents.
  • The SRE team also develops and operates AI-driven tools to streamline runbooks, accelerate incident response, and generate operational insights from platform telemetry to improve reliability and reduce manual effort.
  • Key Responsibilities Design, build, and maintain infrastructure and platform services, including Kubernetes and observability tooling.
  • Monitor production systems, troubleshoot issues, and improve logging, monitoring, alerting, and runbooks.
  • Partner with Product, Engineering, and development teams to translate requirements into reliable and operable infrastructure solutions.
  • Requirements Bachelor’s or advanced degree in computer science, computer engineering, or related practical fields with demonstrated experience. 5+ years of hands-on experience with AWS in a multi-account, multi-region environment: EKS, AWS Organizations, IAM, and KMS.
  • Strong proficiency with Terraform and Infrastructure as Code workflows.
  • Practical experience running workloads on Docker and Kubernetes in production.
View original posting →

Required skills

AWSKubernetesTerraformDockerLinuxPythonGoBashIAMPostgreSQLMySQLCI/CDGit

Nice-to-have skills

New RelicPagerDutyArgoCDAzure

Domain expertise

healthcare

Benefits & perks

Equity, Health benefits, Telemedicine, Lifestyle spending accounts, Parental leave top-up, Family support programs

Tech stack

AWSKubernetesTerraformDockerLinuxPythonGoBashNew RelicPagerDutyPostgreSQLIAMArgoCDAzure

Similar jobs

T

Senior Dev Ops Engineer

T

Senior Site Reliability and Infrastructure Engineer

S

Senior Full-Stack Software Engineer