Reflectionai

Member of Technical Staff - Mid-Training Infra

Full-time · San Francisco, CA
✓ Verified live on the employer's own system · added 139 days ago
Save search

Skills & tools

EmbeddedDistributed SystemsMachine LearningTroubleshooting
Apply on company site ↗ See your fit → free

Full job description

- Design, build, and operate large-scale GPU infrastructure for high-throughput model inference and mid-training workloads.

- Develop systems that power synthetic data generation and reinforcement learning pipelines at scale.

- Build high-performance inference platforms capable of serving and evaluating models across thousands of GPUs.

- Optimize throughput, latency, and GPU utilization for large language model inference and rollout workloads.

- Build infrastructure that supports reinforcement learning pipelines, including large-scale rollout generation, evaluation, and policy improvement loops.

- Work closely with research teams to support distributed RL workloads and large-scale model evaluation infrastructure.

- Improve performance of model execution through kernel-level optimization, model parallelism strategies, and GPU runtime improvements.

- Develop distributed systems that enable large-scale synthetic data generation and RL-driven training workflows.

- Diagnose and resolve performance bottlenecks across inference runtimes, GPU kernels, networking, and distributed compute systems.

- Experience deploying and operating large-scale GPU systems for inference or model serving.

- Several years of hands-on experience building and running production infrastructure.

- Strong understanding of GPU performance characteristics and optimization techniques.

- Experience working with modern inference frameworks such as SGLang, Megatron, or similar high-performance LLM runtimes.

- Familiarity with distributed reinforcement learning infrastructure or rollout generation systems.

- Experience optimizing throughput for large-scale model execution workloads.

- Experience working with GPU kernels or low-level performance optimization.

- Familiarity with infrastructure used for synthetic data pipelines or RL training workflows.

- Experience debugging performance issues across GPU, networking, and distributed execution layers.

More jobs at Reflectionai

Similar jobs near San Francisco, CA

Tell me when more Research Engineer, Mid-Training jobs post near San Francisco We re-check every listing against the employer’s own board — no résumé needed.

Search Member of Technical Staff - Mid-Training Infra jobs near San Francisco, CA → Browse all live jobs

This posting was published by Reflectionai on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.