The below is a starting point. We always make space for exceptional people, so if you don't fit this role exactly, tell us where you would .
- You've carried a pager for a production network and can walk into an outage, find the fault, fix it or drive the RMA, and write the postmortem without someone walking you through it.
- You've written Python or Go scripts that replaced a manual, repetitive network task, from link diagnostics to config pushes to fleet-wide command execution.
- You understand how a transceiver fault, a misconfigured route, and a power event each show up differently in the data, and you can tell them apart from the signals alone.
- You've worked hands-on with link diagnostics, optics, and network monitoring protocols such as gNMI, gRPC, NETCONF, and SONiC, and you're comfortable at the CLI on switches and routers across a fleet.
- You treat toil as a bug. If a repair step means SSHing into ten boxes by hand, you script it once and never do it by hand again.
- You reach real competence in an unfamiliar part of the stack fast, and you document what you learn so the next on-call engineer doesn't start from zero.
- Bonus: RMA and repair lifecycle automation. Large-scale datacenter fabric (BGP, ECMP, spine-leaf). Out-of-band network management.
Fluency with AI coding tools such as Claude Code or Cursor to move faster on scripts and tooling.
Search Production Engineer, Network jobs near San Francisco, CA → Browse all live jobs
This posting was published by Fluidstack on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.