HPC Support Engineer
LambdaAI Infrastructure company
United StatesSenior
Nvidia
TWG Global
US Innovative Technology Fund
Andra Capital
SGW
B Capital
Software Engineering
About the role
TL;DR
Troubleshoot complex HPC infrastructure and platform issues for AI cloud.
- •Lambda is seeking an experienced HPC Support Engineer to join their AI cloud infrastructure team.
- •This role involves troubleshooting complex infrastructure and platform issues, identifying and fixing process gaps, and collaborating with engineering teams.
- •Key Responsibilities Serve as a senior technical escalation point for infrastructure and platform issues.
- •Perform root-cause analysis across distributed systems, clusters, and GPU infrastructure.
- •Proactively identify and fix process, tooling, and documentation gaps.
- •Collaborate with engineering teams to implement permanent fixes for customer pain points.
- •Participate in an on-call rotation and own major incidents.
- •Requirements 3+ years of hands-on HPC experience in an administration, support, or engineering role.
- •Strong understanding and experience supporting Linux in a system administration role.
- •Proven experience in HPC environments, with preference for Kubernetes and/or Slurm.
- •Strong coding ability and CI/CD experience.
- •Proficiency with monitoring/logging tools (Prometheus, Grafana, Datadog).
Required skills
LinuxKubernetesPythonBashPrometheusGrafanaDatadogCI/CDGit
Nice-to-have skills
Docker
Domain expertise
aideveloper-tools
Benefits & perks
Health insurance, Dental insurance, Vision insurance, 401k Plan with company match, Flexible paid time off
Tech stack
LinuxKubernetesPythonBashPrometheusGrafanaDatadogDockerGit