Site Reliability Engineer
Kong Inc.API Management company
Milan, ItalyMid
Tiger Global Management
Balderton Capital
Andreessen Horowitz
Index Ventures
Teachers' Venture Growth
137 Ventures
Software Engineering
About the role
TL;DR
Site Reliability Engineer responsible for cloud infrastructure, reliability, and performance.
- •The Site Reliability Engineering team is responsible for architecting and operating Kong's large-scale cloud infrastructure, ensuring world-class reliability and performance for customer applications.
- •This role focuses on maintaining uptime and enabling product engineering teams to ship features with confidence.
- •Key Responsibilities Build and maintain infrastructure as code using tools like Terraform and Ansible.
- •Implement monitoring, logging, and alerting systems for high uptime.
- •Resolve production incidents and drive blameless post-mortems.
- •Write automation to reduce operational toil and improve system efficiency.
- •Collaborate with developers on reliability and scalability best practices.
- •Requirements Experience operating production workloads on a major cloud provider (AWS, GCP, Azure).
- •Proficiency in Golang, Python, or Bash.
- •Hands-on experience with Docker and Kubernetes.
- •Knowledge of Infrastructure as Code principles.
- •Familiarity with CI/CD concepts and tools.
Required skills
TerraformAnsibleGoPythonBashDockerKubernetesCI/CDAWSGoogle CloudAzure
Nice-to-have skills
GitLab CIJenkinsPrometheusGrafana
Domain expertise
developer-toolsai
Tech stack
TerraformAnsibleGoPythonBashDockerKubernetesCI/CDGitLab CIJenkinsPrometheusGrafanaAWSGoogle CloudAzure