Senior Site Reliability Engineer - Core Cloud Platform
LambdaAI Infrastructure company
On-siteSenior
Nvidia
TWG Global
US Innovative Technology Fund
Andra Capital
SGW
B Capital
Software Engineering
About the role
TL;DR
Improve reliability and scalability of core cloud platform infrastructure.
- •Lambda's Core Cloud Platform is responsible for compute provisioning and infrastructure orchestration across our data centers.
- •This role focuses on enhancing the reliability, scalability, and operational maturity of these systems as Lambda expands its fleet and customer base.
- •Key Responsibilities Operate and scale critical platform services across Lambda's data centers.
- •Improve the reliability of compute provisioning, Instance lifecycle, and regional orchestration systems.
- •Build monitoring, alerting, and tracing for service health and latency.
- •Define SLIs, SLOs, and operational readiness standards.
- •Lead production incident response and postmortems.
- •Requirements 7+ years of experience in site reliability, infrastructure, or production software engineering.
- •Deep experience operating Kubernetes in production.
- •Proficient with Terraform or similar infrastructure-as-code tools.
- •Experience with observability platforms like Prometheus, Grafana, or Datadog.
- •Ability to build production-quality tooling in Go or Python.
Required skills
KubernetesTerraformGoPythonPrometheusGrafanaDatadogCI/CD
Nice-to-have skills
Linux
Domain expertise
aideveloper-tools
Benefits & perks
equity, health insurance, dental coverage, vision coverage, 401k, paid time off
Tech stack
KubernetesTerraformHelmPrometheusGrafanaDatadogGoPython