Education: Doctorate
Experience: 5+ years
You will architect and optimize the core training infrastructure that powers our models. This includes RL training loops, distributed GPU systems, and large-scale data pipelines.
You will work closely with researchers to transform new ideas into reliable, scalable training systems.
- Designing and optimizing large-scale training loops and data pipelines.
- Implementing state-of-the-art techniques and ensuring they are numerically stable and computationally efficient.
- Building internal tooling for launching, monitoring, and reproducing complex experiments.
- Diagnosing deep bottlenecks across the training stack (GPU memory issues, communication overhead, dataloader stalls).
- Translating research prototypes into reusable, production-grade infrastructure.
- Communication optimization (NCCL, RDMA, GPU interconnects)
- You are a strong software engineer who speaks the language of machine learning.
- You may not have a PhD, but you know how to implement a research paper.
- You have deep experience in at least one of the following: Distributed Training & Inference or Data Infrastructure
- You care deeply about performance, numerical stability, and reproducibility.
- You thrive in high-agency environments and enjoy solving hard technical problems.
Search Member of Technical Staff - Research Software Engineer jobs near New York, NY → Browse all live jobs
This posting was published by Reflectionai on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.