Education: Doctorate or related field
- Data Curation & Readiness: Design scalable systems for preparing high-quality multimodal datasets for frontier foundation model training.
- Training Efficiency: Develop algorithms and systems that improve the scalability, efficiency, and cost of large-scale pre-training and post-training.
- Inference Efficiency: Advance techniques that improve inference performance, reduce deployment cost, and enable efficient serving across cloud and edge platforms.
- Open-Source AI Infrastructure: Develop reusable infrastructure and contribute brand new model support to NVIDIA's open-source GenAI training platform.
- MS or Ph.D in Computer Science, AI, Applied Mathematics, or a related field (or equivalent experience).
- Strong foundation in machine learning, deep learning, and optimization.
- Excellent software engineering skills, including Python and PyTorch.
- Experience building high-performance software for large-scale AI systems.
- Strong analytical, debugging, and performance optimization skills.
Experience in some of the following areas is highly desirable:
- Large-Scale Training: Distributed training at scale, including Megatron-LM, Megatron Bridge, FSDP, TP/PP/CP/DP, heterogeneous or per-module parallelism, optimizer research, and efficient sparse or long-context attention.
- LLM/VLM Post-Training: Supervised fine-tuning (SFT), reinforcement learning for LLMs (e.g., PPO, GRPO, asynchronous RL), and large-scale RL frameworks such as NeMo-RL.
- Inference Efficiency: Model compression techniques including quantization (FP8, NVFP4, INT4), pruning, knowledge distillation, neural architecture search, and diffusion or non-autoregressive language models.
- Open-Source AI Infrastructure: Contributing to open-source AI frameworks such as Megatron-LM, Megatron Bridge, NeMo-RL, or Hugging Face Transformers along with experience in GPU performance optimization, distributed systems, latency/throughput analysis, and profiling of large-scale AI workloads.
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4.
Applications for this job will be accepted at least until August 9, 2026.
Search Senior AI Systems and Algorithms Engineer jobs near US, CA, Santa Clara → Browse all live jobs
This posting was published by NVIDIA | NVIDIA on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.