Nuance Labs

Member of Technical Staff - Model Optimization and Inference (New Grad)

$200K–$300KFull-time · Seattle, WA
✓ Verified live on the employer's own system · added 60 days ago
Save search
Entry level

Requirements

Education: Bachelor's degree or related field

Skills & tools

EmbeddedMachine LearningStudent AssessmentPythonResearch
Apply on company site ↗ See your fit → free

Full job description

We’re looking for someone who’s excited about taking trained models and squeezing every last millisecond out of them. You understand — or want to deeply understand — the full stack from model weights to serving infrastructure: quantization, KV cache optimization, kernel-level acceleration, batching strategies. You’ve worked with vLLM, SGLang, or similar frameworks (through coursework, research, internships, or open-source) and have opinions about where they fall short.

This posting is aimed at early-career engineers finishing or recently finished with a BS, MS, or PhD. We don’t require a PhD — we care about systems intuition, engineering chops, and the appetite to go deep.

Our stack is more complex than a standard LLM deployment: we’re serving a full-duplex multimodal system that must satisfy strict real-time latency constraints. There’s a lot of unsolved optimization work here, and we want someone who finds that genuinely exciting and is ready to grow fast alongside people who’ve built these systems before.

  • Contribute to end-to-end inference optimization across our model stack — LLMs, audio models, and diffusion-based components
  • Work with inference serving frameworks (vLLM, SGLang, TensorRT-LLM, etc.) and extend them for our specific workloads
  • Apply quantization techniques (INT8, INT4, GPTQ, AWQ, and beyond) to reduce memory footprint and increase throughput without meaningfully degrading quality
  • BS, MS, or PhD in CS, ML, or a related field — completed or in the final stretch
  • Strong fundamentals in LLM inference or ML systems — KV caching, memory layout, attention kernels, batching, or serving — picked up through coursework, research, internships, or open-source. You don’t need to have shipped at production scale yet; you do need to learn fast and go deep.
  • Exposure to inference serving frameworks (vLLM, SGLang, TensorRT-LLM, or similar) — even at a research or hobby level
  • Strong Python and PyTorch skills; familiarity with CUDA or Triton is a significant plus
  • Curiosity about diffusion inference, speculative decoding, quantization, or other inference-time acceleration techniques
  • Internship or research experience with LLM inference, ML systems, or model serving
  • Contributions to open-source inference frameworks (vLLM, SGLang, TensorRT-LLM, etc.)
  • CUDA / Triton kernel work, even at a research or hobby scale
  • Publications or research projects in MLSys, model compression, or inference optimization

More jobs at Nuance Labs

Similar jobs near Seattle, WA

Search Member of Technical Staff - Model Optimization and Inference (New Grad) jobs near Seattle, WA → Browse all live jobs

This posting was published by Nuance Labs on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.