Skip to content
Lambda logo

Senior Site Reliability Engineer - Core Cloud Platform

LambdaAI Infrastructure company
On-siteSenior
TWG Global
US Innovative Technology Fund
Andra Capital
SGW
Nvidia logo
Nvidia
B Capital
Software Engineering

About the role

TL;DR

Improve reliability and scalability of core cloud platform infrastructure.

  • Lambda's Core Cloud Platform is responsible for compute provisioning and infrastructure orchestration across our data centers.
  • This role focuses on enhancing the reliability, scalability, and operational maturity of these systems as Lambda expands its fleet and customer base.
  • Key Responsibilities Operate and scale critical platform services across Lambda's data centers.
  • Improve the reliability of compute provisioning, Instance lifecycle, and regional orchestration systems.
  • Build monitoring, alerting, and tracing for service health and latency.
  • Define SLIs, SLOs, and operational readiness standards.
  • Lead production incident response and postmortems.
  • Requirements 7+ years of experience in site reliability, infrastructure, or production software engineering.
  • Deep experience operating Kubernetes in production.
  • Proficient with Terraform or similar infrastructure-as-code tools.
  • Experience with observability platforms like Prometheus, Grafana, or Datadog.
  • Ability to build production-quality tooling in Go or Python.
View original posting →

Required skills

KubernetesTerraformGoPythonPrometheusGrafanaDatadogCI/CD

Nice-to-have skills

Linux

Domain expertise

aideveloper-tools

Benefits & perks

equity, health insurance, dental coverage, vision coverage, 401k, paid time off

Tech stack

KubernetesTerraformHelmPrometheusGrafanaDatadogGoPython

Similar jobs

N

Consulting Engineer

Z

Software Engineer Intern (Summer 2027)

F

Infrastructure Engineer (Linux)