Education: Doctorate or related field
Experience: 5+ years
- Implement quantized and sparse recipes in inference engines (vLLM, TRT-LLM, SGLang)
- Own model export pipelines (ModelOpt, Megatron-LM HuggingFace), ensuring quantized checkpoints serialize correctly for downstream serving
- Build prototypes and benchmarking harnesses to evaluate recipe throughput/interactivity before full optimization
- Develop data analysis tooling and visualizations for numerics debugging
- Improve developer productivity across the team: CI, build systems, training infrastructure, pipeline friction
- Strong software engineering fundamentals: concise, well-tested code; fluent with AI-assisted tooling
- Experience with ML accelerators with a basic understanding of how certain ML layers affect execution time
- Familiarity with PyTorch internals (custom ops, autograd, export) or equivalent framework
- Experience reading, modifying, or contributing to a large open-source codebase
- MS/PhD in Computer Science or related field, or equivalent experience.
- Demonstrated ability to move fast with ambiguous requirements, with strong written and verbal communication
- Experience contributing to inference serving frameworks (vLLM, TRT-LLM, SGLang) or Triton kernel development
- Track record of debugging numerical issues across mixed-precision boundaries
- Deep experience with model compression techniques: PTQ, QAT, structured/unstructured sparsity
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4.
Applications for this job will be accepted at least until July 26, 2026.
Search Senior Software Engineer, Quantized Inference jobs near US, CA, Santa Clara → Browse all live jobs
This posting was published by NVIDIA | NVIDIA on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.