Site Reliability Engineer
Epidemic SoundMusic Licensing company
Stockholm, SwedenSenior
Creandum
EQT
Blackstone
Alecta
AMF
TIN Fonder
Software Engineering
About the role
TL;DR
Build and operate the platform for engineering teams, ensuring reliability, scalability, and security.
- •As a Site Reliability Engineer at Epidemic Sound, you will be a core member of the central platform team that builds and operates the platform the rest of Engineering ships on.
- •Your goal is to make the reliable way the easy way, enabling product teams to build and ship safely.
- •Key Responsibilities Build and operate the platform our services run on, including GKE clusters and Terraform for cloud definition.
- •Own the path from commit to production, managing CI/CD, GitOps, and progressive-delivery patterns.
- •Strengthen the networking and routing layer, managing traffic, firewalls, and network policies.
- •Govern access and guardrails through IAM, policy-as-code, and break-glass paths.
- •Grow reliability and observability through alert hygiene, runbooks, SLOs, metrics, and tracing.
- •Requirements Solid grasp of Kubernetes fundamentals, including controllers, core components, and CNI/networking.
- •Experience with infrastructure as code (Terraform), delivery (CI/CD, GitOps), and traffic management.
- •Understanding of networking, routing, VPC, firewalls, network policies, and IAM.
- •Operational depth in monitoring, troubleshooting distributed systems, and Unix/Linux.
- •Agentic development mindset, actively using AI agents in work.
Required skills
KubernetesTerraformCI/CDIAMPrometheusGrafanaLinux
Nice-to-have skills
GKEArgoCDGoogle Cloud
Domain expertise
developer-toolsmediaentertainment-music
Tech stack
KubernetesGKETerraformCI/CDArgoCDIAMPrometheusGrafanaGoogle Cloud