Skip to content
FluidStack logo

Software Engineer, Networking

FluidStackGPU Cloud company
San Francisco, United StatesSenior
Situational Awareness
Astro Capital (NY)
7GC & Co
Armyn Capital
Autopilot Management Company
Bare Metal Ventures
Software Engineering

About the role

TL;DR

Own network fleet health end to end, build active debugging tooling, and turn repair into a pipeline, not a procedure.

  • The Software Engineer, Networking will own network fleet health end to end, build active debugging tooling, and turn repair into a pipeline, not a procedure.
  • Key Responsibilities Define the realtime monitoring requirements, build the alerting lifecycle, and ship the dashboards that give every on-call engineer a true picture of network state across all sites.
  • Link diagnostics, remote command execution across the fleet, and repair visualization
  • the tools that turn a network fault from a mystery into a solvable problem, fast.
  • Build the automation that takes a network failure from detection through parts management and return to service.
  • Ticket integration, repair lifecycle pipelines, transceiver and optics tracking
  • owned, not improvised.
  • Build the frameworks that gate new sites and hardware into production.
  • You define what a healthy network looks like before it carries traffic.
  • Own end-to-end reliability, scalability, and operation of the network at-scale.
  • Requirements You treat toil as a bug.
  • If diagnosing a link failure requires SSHing into boxes and running commands by hand, you build the tool that does it for you.
  • You think in systems.
  • You understand how a transceiver fault, a misconfigured route, and a power event each propagate differently
  • and you build tooling that can tell them apart.
  • You move toward ambiguity, not away from it.
  • You walk into the fog, build the map, and explain it to everyone else.
  • You learn at a steep slope.
  • You reach real competence in an unfamiliar domain fast.
  • We value this over existing expertise.
  • You carry a pager without flinching.
  • You run the incident, write the postmortem, fix the systemic cause, and move on.
  • You're fluent with AI tooling.
  • LLM APIs, MCP servers, and agentic frameworks, and you drive Claude Code, Cursor, or similar every day.
  • You've shipped production network tooling or automation that other teams depend on, and you're comfortable in any language using AI coding tools.
  • You’ve developed automation tools in Go and Python, are experienced in link diagnostics, optical networks, and understanding network monitoring (gNMI, gRPC, NETCONF, SONiC)
View original posting →

Required skills

PythonGogRPCLLMsKubernetesDockerCI/CDAWSGoogle CloudAzureTerraformJenkinsPrometheusGrafanaGit

Domain expertise

telecom

Benefits & perks

Competitive total compensation package (salary + equity), Retirement or pension plan, Health, dental, and vision insurance, Generous PTO policy

Tech stack

PythonGogRPCLLMsKubernetesDockerCI/CDAWSGoogle CloudAzureTerraformJenkinsPrometheusGrafanaGitLinux

Similar jobs

O

Dev Ops Engineer

A

Security Architect

A

Software Engineer II