Experience: 5+ years
Reflection's Compute Platform team keeps our compute layer healthy and highly available. We run a Kubernetes-based platform distributed across multiple neo-clouds, where multi-cloud scheduling, node health, and performance debugging at scale present genuinely hard systems problems.
As Compute Platform Lead, you'll provide front-line leadership of the team that builds and operates this layer. You'll build, mentor, and grow a team of strong systems engineers, guide the technical and architectural decisions across multi-cloud scheduling, cluster management, and next-generation GPU deployments, and work closely with our training teams to co-design fault tolerance, node health checks, and remediation.
You'll stay close enough to the systems to make targeted contributions as an individual contributor and to maintain a deep understanding of the compute fleet our largest training runs depend on. Managing vendors - and the important deals that come with them - is a core part of the job.
- Build, mentor, and grow a high-performing team of systems engineers. Coach and support your reports in understanding, and pursuing, their professional growth.
- Provide front-line leadership of engineering efforts to keep the compute fleet reliable and highly available - multi-cloud scheduling, cluster management, and the path to next-generation GPUs and increasingly larger cluster sizes.
- Stay hands-on: become familiar with the team's technical stack enough to make targeted contributions as an individual contributor.
- Manage day-to-day execution: prioritize the team's work and manage projects in a highly dynamic, fast-paced environment.
- Guide technical and architectural decisions, emphasizing scalability, robustness, and reliability - automatic remediation, topology-aware scheduling, capacity planning, rapid hardware debugging, and cluster-wide monitoring and performance benchmarking.
- Work closely with our training teams to co-design fault tolerance, node health checks, and remediation, and manage the vendor relationships and important deals the compute fleet depends on.
- Prepare the fleet for what's next: next-generation GPUs and larger clusters, and - longer term - multi-cloud storage, petabyte-scale data replication, and GPU-to-GPU network performance.
- Raise the bar for technical judgment, prioritization, communication, and execution in a fast-moving environment.
- Experience building, mentoring, and growing systems or infrastructure teams while staying technically hands-on. (Comfortable growing into leading a team of ~10 quickly if you haven't managed at that scale before.)
- Deep systems-level engineering experience with a focus on cluster-wide behavior and maintenance.
- Strong coding ability and the credibility to earn the technical trust of a strong team.
- Depth in at least one of orchestration, storage, or GPU hardware - with the ability to learn the rest. Deep GPU knowledge beyond standard Kubernetes (e.g., NCCL) is a plus, not a prerequisite.
- Cloud storage expertise - managing high-performance data products (like VAST) across multiple data centers and handling datasets and checkpointing at scale - is a plus.
- Experience managing vendors, including negotiating and operating important deals.
- Ability to guide strategy and drive execution across a multi-cloud, large-fleet environment, and to partner effectively with research and training teams.
Search Member of Technical Staff - Engineering Lead, Compute Platform jobs near San Francisco, CA → Browse all live jobs
This posting was published by Reflectionai on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.