Skip to content
Lambda logo

HPC Support Engineer

LambdaAI Infrastructure company
United StatesSenior
TWG Global
US Innovative Technology Fund
Andra Capital
SGW
Nvidia logo
Nvidia
B Capital
Software Engineering

About the role

TL;DR

Troubleshoot complex HPC infrastructure and platform issues for AI cloud.

  • Lambda is seeking an experienced HPC Support Engineer to join their AI cloud infrastructure team.
  • This role involves troubleshooting complex infrastructure and platform issues, identifying and fixing process gaps, and collaborating with engineering teams.
  • Key Responsibilities Serve as a senior technical escalation point for infrastructure and platform issues.
  • Perform root-cause analysis across distributed systems, clusters, and GPU infrastructure.
  • Proactively identify and fix process, tooling, and documentation gaps.
  • Collaborate with engineering teams to implement permanent fixes for customer pain points.
  • Participate in an on-call rotation and own major incidents.
  • Requirements 3+ years of hands-on HPC experience in an administration, support, or engineering role.
  • Strong understanding and experience supporting Linux in a system administration role.
  • Proven experience in HPC environments, with preference for Kubernetes and/or Slurm.
  • Strong coding ability and CI/CD experience.
  • Proficiency with monitoring/logging tools (Prometheus, Grafana, Datadog).
View original posting →

Required skills

LinuxKubernetesPythonBashPrometheusGrafanaDatadogCI/CDGit

Nice-to-have skills

Docker

Domain expertise

aideveloper-tools

Benefits & perks

Health insurance, Dental insurance, Vision insurance, 401k Plan with company match, Flexible paid time off

Tech stack

LinuxKubernetesPythonBashPrometheusGrafanaDatadogDockerGit

Similar jobs

O

Senior Dev Ops Engineer

O

Senior Cloud Ops Engineer

O

Software Engineer