NVIDIA | NVIDIA

Inference Performance Engineer, Agent Driven Inference Optimization

Full-time · US, CA, Remote
✓ Verified live on the employer's own system · added 3 days ago
Save search

Requirements

Education: Bachelor's degree or related field

Skills & tools

RecruitingMachine LearningRecipe DevelopmentElectricalEmbeddedPythonC Plus PlusCommunications
Apply on company site ↗ See your fit → free

Full job description

- Distill your performance instincts into reusable skills, workflows, and evidence-backed methodologies that AI agents can complete autonomously. Review agent-generated experiments, validate findings, and curate best-known configurations.

- Performance improvement of AI inference workloads that methodically increase throughput-per-GPU and user interactivity by exploring configuration options, parallelism techniques, batching, KV cache handling, quantization, and speculative decoding settings.

- Measure and optimize both aggregated and disaggregated serving architectures across TensorRT-LLM, SGLang, vLLM, and Dynamo on NVIDIA's latest GPU platforms.

- Profile workloads using Nsight Systems, kernel traces, and internal analysis tools. Use roofline and speed-of-light analysis to find credible headroom and drive fixes from hypothesis to measured wins.

- Land improvements upstream: serving framework patches, optimized kernels, and deployment recipes that advance the public Pareto frontier while maintaining strict model correctness.

- Collaborate with TensorRT-LLM, SGLang, vLLM, kernel, benchmarking, and GPU architecture teams to convert profiling insights into delivered performance improvements.

- BS, MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, Applied Math, or a related field, or equivalent experience.

- Must have: Extensive knowledge of the efficiency and optimization involved in AI model execution, covering continuous batching, throughput-latency tradeoffs, KV cache and memory limitations, parallel processing techniques, MoE serving, quantization, and meeting serving SLAs.

- Must have: Hands-on experience benchmarking and profiling GPU workloads using tools such as Nsight Systems, Nsight Compute, CUPTI, or PyTorch profiler, and interpreting kernel-level performance data.

- Strong Python engineering skills and the ability to navigate and modify large C++/CUDA serving codebases.

- Rigorous experimental methodology with controlled single-variable comparisons, reproducible benchmarks, and evidence-backed optimization decisions.

- Strong written and verbal communication skills to explain performance tradeoffs clearly to both humans and documentation for autonomous systems.

- Direct contributions to TensorRT-LLM, vLLM, SGLang, FlashInfer, Dynamo, or comparable inference frameworks.

- Experience with disaggregated serving, wide expert-parallel MoE inference, KV cache transfer, or NCCL/NIXL/NVSHMEM communication at multi-node scale.

- CUDA kernel authorship or optimization experience on Hopper/Blackwell architectures, focusing on Tensor Cores, TMA, and warp specialization.

- Proven results on public inference benchmarks such as MLPerf Inference or SemiAnalysis InferenceX.

- Experience building or operating agentic AI workflows to automate engineering tasks.

Widely considered to be one of the technology world's most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. As you plan your future, see what we can offer to you and your family www.nvidiabenefits.com/

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 124,000 USD - 195,500 USD for Level 2, and 152,000 USD - 241,500 USD for Level 3.

Applications for this job will be accepted at least until August 11, 2026.

More jobs at NVIDIA | NVIDIA

Similar jobs near US, CA, Remote

Tell me when more Sr. Inference Optimization Engineer (local / edge runtime) jobs post near US, CA, Remote We re-check every listing against the employer’s own board — no résumé needed.

Search Inference Performance Engineer, Agent Driven Inference Optimization jobs near US, CA, Remote → Browse all live jobs

This posting was published by NVIDIA | NVIDIA on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.