Nuance Labs

Member of Technical Staff - Model Optimization and Inference (Experienced)

$250K–$350KFull-time · Seattle, WA
✓ Verified live on the employer's own system · added 66 days ago
Save search

Skills & tools

EmbeddedHiringMachine LearningStudent AssessmentPython
Apply on company site ↗ See your fit → free

Full job description

We’re looking for someone who specializes in taking trained models and squeezing every last millisecond out of them. You understand the full stack from model weights to serving infrastructure — quantization, KV cache optimization, kernel-level acceleration, batching strategies — and you know which lever to pull for which problem.

You’ve worked with vLLM, SGLang, or similar frameworks at scale and have strong opinions about where they fall short.

This posting is aimed at experienced engineers and researchers who’ve operated at a senior to senior-staff level at big tech, a leading AI lab, or a high-traffic inference team. Everyone at Nuance is MTS — we don’t run title ladders — but we’re hiring people who have already done this work at scale.

Our stack is more complex than a standard LLM deployment: we’re serving a full-duplex multimodal system that must satisfy strict real-time latency constraints. There’s a lot of unsolved optimization work here, and we need someone who finds that genuinely exciting.

  • Own end-to-end inference optimization across our model stack — LLMs, audio models, and diffusion-based components
  • Evaluate, deploy, and extend inference serving frameworks (vLLM, SGLang, TensorRT-LLM, etc.) for our specific workloads
  • Apply and develop quantization techniques (INT8, INT4, GPTQ, AWQ, and beyond) to reduce memory footprint and increase throughput without meaningfully degrading quality
  • Significant hands-on experience with LLM inference optimization — you’ve shipped work on KV caching, memory layout, attention kernels, or batching strategies in a production or high-traffic research context
  • Proven proficiency with inference serving frameworks — vLLM, SGLang, TensorRT-LLM, or similar — including going well beyond default configurations and adapting them to non-standard workloads
  • Experience optimizing diffusion model inference (latency reduction, step distillation, caching, or kernel-level work)
  • Strong Python and PyTorch skills; comfort reading and writing CUDA or Triton kernels is a significant plus
  • Familiarity with speculative decoding or other inference-time acceleration techniques
  • Hands-on experience with post-training quantization (GPTQ, AWQ, or similar) and a clear sense of quality/performance tradeoffs
  • Experience deploying real-time AI systems with hard latency SLAs
  • Prior work at an AI lab, inference startup, or on a high-traffic model serving platform

More jobs at Nuance Labs

Similar jobs near Seattle, WA

Tell me when more Sr. Inference Optimization Engineer (local / edge runtime) jobs post near Seattle, WA We re-check every listing against the employer’s own board — no résumé needed.

Search Member of Technical Staff - Model Optimization and Inference (Experienced) jobs near Seattle, WA → Browse all live jobs

This posting was published by Nuance Labs on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.