Education: Bachelor's degree or related field
Experience: 4+ years
We are looking for a performance-obsessed AI Infrastructure Engineer to push LLM inference to its absolute limits on Intel's next-generation GPU architectures. In this role, you will dive deep into the inference stack and redefine peak performance. You will work end-to-end across the stack: profiling bottlenecks, writing custom GPU kernels, and upstreaming your optimizations directly into industry-standard serving frameworks like vLLM and SGLang.
Your optimizations will be instrumental in unlocking the full potential of Intel hardware for state-of-the-art generative AI workloads.
What You Will Do - Drive Inference Performance: Own the end-to-end optimization pipeline for running state-of-the-art LLMs on Intel GPUs. - Deep Stack Optimization: Profile, diagnose, and resolve cross-stack performance bottlenecks. - Kernel Development and Integration: Design, write, and optimize custom high-performance kernels for critical attention mechanisms, MoE, quantization, and operator fusions. - Open Source Leadership: Upstream your architectural improvements and hardware backends directly into open-source repositories like vLLM, SGLang, and PyTorch, acting as a bridge between the hardware teams and the open-source community. - Shape the Hardware Roadmap: Apply roofline analysis and systematic profiling to decompose bottlenecks.
You will partner with our architecture and compiler teams to shape future GPU roadmaps based on real-world GenAI workload data.
- Show passion about AI infrastructure and performance optimization.
- Bachelors Degree in Computer Science, Software Engineering, Artificial Intelligence/Machine Learning, or related field and 4+ years experience, Masters Degree and 3+ years, OR PhD. - 3+ years of relevant software engineering experience in GPU computing, AI systems, or high-performance computing (HPC). - Proficiency in modern C++ and Python.
You are comfortable reading and modifying complex systems-level code.
Preferred
Qualifications - Understanding of CPU/GPU architecture. - Understanding of modern LLM architectures and inference paradigms: attention mechanisms, KV caching, continuous batching, speculative decoding, and prefill-decode disaggregation. - Prior open-source contributions to inference engines (vLLM, SGLang, PyTorch, llama.cpp). - Hands-on experience writing and optimizing custom GPU kernels using Triton, SYCL, CUDA/CUTLASS, or other DSLs. - Experience with scale-out inference orchestration across multi-node topologies. - You leverage AI coding agents daily to accelerate your own workflow and benchmark generation.
Your expertise will play a vital role in advancing Intel's AI technology. We invite you to bring your skills, experience, and passion for AI to make an impact-apply today.
Additional Locations: US, California, Folsom, US, Oregon, Hillsboro, US, Texas, Austin
Annual Salary Range for jobs which could be performed in the US: $170,500.00-315,490.00 USD
Work Model for this Role This role will be eligible for our hybrid work model which allows employees to split their time between working on-site at their assigned Intel site and off-site. * Job posting details (such as work model, location or time type) are subject to change.
Search AI Infrastructure Engineer jobs near US, Texas, Austin → Browse all live jobs
This posting was published by Intel on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.