We’re looking for someone who specializes in taking trained models and squeezing every last millisecond out of them. You understand the full stack from model weights to serving infrastructure — quantization, KV cache optimization, kernel-level acceleration, batching strategies — and you know which lever to pull for which problem.
You’ve worked with vLLM, SGLang, or similar frameworks at scale and have strong opinions about where they fall short.
This posting is aimed at experienced engineers and researchers who’ve operated at a senior to senior-staff level at big tech, a leading AI lab, or a high-traffic inference team. Everyone at Nuance is MTS — we don’t run title ladders — but we’re hiring people who have already done this work at scale.
Our stack is more complex than a standard LLM deployment: we’re serving a full-duplex multimodal system that must satisfy strict real-time latency constraints. There’s a lot of unsolved optimization work here, and we need someone who finds that genuinely exciting.
Search Member of Technical Staff - Model Optimization and Inference (Experienced) jobs near Seattle, WA → Browse all live jobs
This posting was published by Nuance Labs on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.