Crusoe

Principal Systems Software Engineer

$260K–$340KFull-time · San Francisco, CA - US
✓ Verified live on the employer's own system · added 5 days ago
Save search
Senior · 12+ yrs exp

Requirements

Education: Bachelor's degree or related field

Experience: 12+ years

Skills & tools

Machine LearningEmbeddedTeam LeadershipDevopsRecordkeepingTroubleshootingCloud PlatformsLinux

Benefits — mentioned in this posting

401(k) / retirement
Apply on company site ↗ See your fit → free

Full job description

As the Principal Systems Software Engineer , you will serve as the visionary lead for Crusoe's next-generation AI infrastructure. This is a role for an industry-recognized expert who has already "seen the movie" at hyperscale and is ready to redefine the I/O path for the age of generative AI. You aren't just building a cloud; you are designing the fluid fabric that unifies Bare-Metal-as-a-Service (BMaaS), Intelligent IaaS, and Elastic CaaS into a single, high-performance pool of intelligence.

In this position, you will bridge the gap between silicon and software, advising executive leadership on critical hardware/software co-design pivots while remaining hands-on enough to lead elite R&D teams in shipping production-grade kernel and orchestration code. We are looking for a master of the I/O path who can push massive-scale training workloads to the theoretical limits of hardware.

This is a full-time position.

- Bare-Metal-as-a-Service (BMaaS): Architect systems that deliver raw GPU throughput via zero-latency InfiniBand/RDMA fabrics for massive-scale training.

- Intelligent IaaS: Design highly optimized, thin virtualization layers using KVM or custom micro-VMs to provide enterprise-grade isolation without the "virtualization tax."

- Elastic CaaS: Build a high-performance container substrate (utilizing Kubernetes or Slurm) that allows AI workloads to burst and scale across heterogeneous GPU nodes.

- Mastering the I/O Path: Lead the architectural design of our internal cloud fabric, drawing on experience from top-tier hyperscalers to drive the technical roadmap for SR-IOV, RDMA, and virtualized GPU scheduling.

- Advanced R&D Leadership: Lead elite workstreams to prototype and productionize novel methods for managing memory, networking, and compute that don't yet exist in standard cloud distributions.

- Technical Strategy & Documentation: Draft white papers and RFCs that define the next two years of Crusoe's compute and networking stack.

- High-Level Debugging: Work alongside Staff and Senior engineers to resolve complex race conditions in the I/O path and optimize kernel-level memory pinning for GPU clusters.

- Industry Influence: Represent Crusoe in open-source communities and industry forums to influence the global direction of cloud-native AI infrastructure.

- Hyperscale Provenance: 12+ years of experience designing and shipping core infrastructure at a major hyperscaler (e.g., OCI, AWS, Azure, GCP) or a specialized HPC cloud.

- Deep Systems Authority: Authoritative knowledge of the Linux kernel, virtualization internals (KVM, QEMU, Firecracker), and high-performance networking (RoCE v2, InfiniBand).

- Hardware-Software Co-Design: Proven ability to design software that maximizes the performance of NVIDIA/AMD GPUs and high-speed NICs.

- R&D Leadership: Experience leading cross-functional teams through high-ambiguity projects and delivering production-ready, mission-critical systems.

- Industry Contributions: A portfolio of significant contributions to the field, which may include patents, major open-source contributions, or published research in distributed systems.

- Communication Mastery: The rare ability to explain the nuances of memory-mapped I/O to an engineer and the business value of a new fabric architecture to the Board.

- Mandatory Education: A Bachelor's or Master's degree in Computer Science, Computer Engineering, or a related analytical field (or equivalent professional experience).

- Patent Holder: Possession of patents related to network virtualization, GPU scheduling, or distributed file systems.

- Open Source Leadership: Maintainer status or significant contributions to the Linux Kernel, Kubernetes, or specialized HPC projects.

- AI/ML Workload Expertise: Direct experience optimizing infrastructure for Large Language Model (LLM) training and inference at scale.

- Peer-reviewed publications at top systems venues (OSDI, SOSP, NSDI, SC), patents, IETF RFCs

- Significant open-source contributions, or industry white papers on distributed systems, high-performance networking, or AI infrastructure

- 401(k) Retirement plan with company match up to 4% of salary

$260,000 - $340,000 + Significant Equity & Bonus.

Compensation is determined by the applicant's depth of expertise, previous impact at scale, and alignment with our architectural goals.

More jobs at Crusoe

Similar jobs near San Francisco, CA - US

Tell me when more Principal Software Engineer jobs post near San Francisco, CA We re-check every listing against the employer’s own board — no résumé needed.

Search Principal Systems Software Engineer jobs near San Francisco, CA - US → Browse all live jobs

This posting was published by Crusoe on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.