Education: Bachelor's degree or related field
We’re looking for someone who’s excited about taking trained models and squeezing every last millisecond out of them. You understand — or want to deeply understand — the full stack from model weights to serving infrastructure: quantization, KV cache optimization, kernel-level acceleration, batching strategies. You’ve worked with vLLM, SGLang, or similar frameworks (through coursework, research, internships, or open-source) and have opinions about where they fall short.
This posting is aimed at early-career engineers finishing or recently finished with a BS, MS, or PhD. We don’t require a PhD — we care about systems intuition, engineering chops, and the appetite to go deep.
Our stack is more complex than a standard LLM deployment: we’re serving a full-duplex multimodal system that must satisfy strict real-time latency constraints. There’s a lot of unsolved optimization work here, and we want someone who finds that genuinely exciting and is ready to grow fast alongside people who’ve built these systems before.
Search Member of Technical Staff - Model Optimization and Inference (New Grad) jobs near Seattle, WA → Browse all live jobs
This posting was published by Nuance Labs on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.