The below is a starting point. We always make space for exceptional people, so if you don't fit this role exactly, tell us where you would.
- You treat toil as a bug. Manual steps in a repair workflow are a backlog item, not a job description.
- You have an instinct for hardware. You're comfortable reasoning about failure modes at the firmware and silicon level, not just the software stack above it.
- You move toward ambiguity, not away from it. You walk into the fog, build the map, and explain it to everyone else.
- You learn at a steep slope. You reach real competence in an unfamiliar domain fast. We value this over existing expertise.
- You carry a pager without flinching. You run the incident, write the postmortem, fix the systemic cause, and move on.
- You're fluent with AI tooling. LLM APIs, MCP servers, and agentic frameworks, and you drive Claude Code, Cursor, or similar every day.
- You've shipped production automation that other teams depend on, and you're comfortable in any language using AI coding tools.
- Bonus: Hardware lifecycle management and RMA automation. BMC/Redfish or IPMI tooling. GPU qualification or burn-in frameworks.
Workflow and orchestration engines (Temporal, Cadence). Metrics and alerting pipelines (Prometheus, Grafana). Go or Python...
Total compensation may also include equity in the form of stock options.
Search Software Engineer, Compute (GPU) jobs near San Francisco, CA → Browse all live jobs
This posting was published by Fluidstack on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.