Experience: 3+ years
What You'll Do We are seeking a Software Engineer with good experience in backbone infrastructure that build, scale and optimize GPU-first cloud infrastructure that enables High Performance Computing and demanding AI workloads.
In this role you will be responsible for: - Design, develop and maintain software for GPU/CPU compute infrastructure with focus on performance, scalability, and reliability. - Implement and develop services for baremetal and VM instancing. - Develop distributed systems for managing and orchestrating compute resources across various SKU's. - Troubleshoot and debug complex issues in a production and development environment. - On-call and incident ownership - Collaboration across multiple teams and drive ambiguity in requirements or solutions on RFC's.
You - 3+ years of experience working with Go (Golang) or Python in production environments. - 3+ years of experience with bare metal & virtualization hardware management and configuration. - Are comfortable working in Linux environments and debugging issues at the OS, hardware, and networking layers. - Can independently troubleshoot complex systems and communicate effectively across software, infrastructure, and vendor teams.
Nice to Have - Familiarity with GPU Infrastructure or high-performance computing environments. - Experience with Slurm or Kubernetes-based cluster management. - Experience with core public cloud internals (Virtualization, KVM, QEMU, Security and Fleet health) - Experience with durable execution platforms like Temporal.
Search Software Engineer - Compute jobs near San Francisco Office (Fremont St) (Remote) → Browse all live jobs
This posting was published by Lambda on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.