Education: Bachelor's degree
Experience: 12+ years
- Track and evaluate innovations in leading open-source LLM inference frameworks - identify performance-critical features and algorithmic improvements relevant to NVIDIA edge AI hardware
- Analyze how new model architectures and inference algorithms (attention variants, MoE routing, speculative decoding, multi-token prediction, quantized inference) map onto NVIDIA GPU architecture - identify mismatch, fallback paths, and optimization opportunities
- Characterize multi-node inference behavior: collective communication primitives (NCCL/RCCL), topology-aware all-reduce strategies, and parallelism efficiency on edge cluster configurations
- Produce performance analysis reports mapping theoretical hardware limits (memory bandwidth, FLOP/s, interconnect throughput) to observed inference throughput, latency, and utilization
- Own the model validation workflow for new model releases: architecture compatibility assessment, inference recipe development, performance characterization, and publication to developer recipe sites
- Develop and maintain developer-facing inference recipes: keep them accurate as frameworks evolve, automate staleness detection, and build feedback loops from CI results to recipe updates
- Engage with community and partners on model bring-up questions; serve as the technical point of contact for hardware-specific inference issues related to partner concerns
- BS, MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or equivalent experience.
- 12+ years of software engineering with depth in GPU computing, ML systems, or high-performance inference
- Strong Python or C++ programming, software design, and software engineering skills.
- Hands-on experience with GPU kernel development or optimization (CUDA/C++, Triton, or equivalent) - you understand how thread blocks, memory hierarchy, and warp execution affect real-world performance
- Working knowledge of LLM inference internals: attention mechanisms, KV-cache management, continuous batching, quantization formats, and tensor parallelism
- Container engineering expertise: multi-architecture Docker or OCI builds, layer optimization, runtime configuration, NVIDIA Container Toolkit
- Strong analytical skills: ability to form a performance hypothesis, design an experiment, interpret results, and communicate findings clearly
Widely considered to be one of the technology world's most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. As you plan your future, see what we can offer to you and your family www.nvidiabenefits.com/
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 224,000 USD - 356,500 USD for Level 5, and 272,000 USD - 431,250 USD for Level 6.
Applications for this job will be accepted at least until July 28, 2026.
Search Senior Software Engineer - GPU Local AI Platforms jobs near US, TX, Austin → Browse all live jobs
This posting was published by NVIDIA | NVIDIA on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.