- Own evaluation pipelines — design, build, and automate offline and live evals that keep our speech and multimodal models honest in production.
- Harness the data — create tooling for safe, versioned, privacy-aware dataset curation and discovery.
- Ship models, not slide decks — partner with research and infra to prototype, train, and deploy state-of-the-art voice models that power Sesame’s real-time companion https://www.sesame.com/research/crossing_the_uncanny_valley_of_voice?utm_source=chatgpt.com experience.
- Squeeze silicon — scale training and inference for LLM-class workloads; chase latency, throughput, and cost until the graphs flatten.
- Wire up monitoring and live evals — surface quality regressions before users or PMs notice.
- Move at startup speed — take ideas from whiteboard to production in days, not quarters; leave a clean trail of tests and dashboards behind.
- Proven software engineer who loves ML; comfortable writing production code across the stack.
- Hands-on experience training or fine-tuning large language or other large-scale models with a variety of techniques.
- Evaluation expert — you’ve designed metrics and harnesses that actually predict user happiness.
- Deep knowledge of the ML lifecycle: dataset ops, training pipelines, eval frameworks, deployment, and monitoring.
- History of shipping complex projects to production—especially user-facing, online ML systems—despite shifting requirements and surprise roadblocks.
- High agency and the judgment to know when to sprint solo vs. pull in the squad.
- Track record of setting technical direction, driving consensus, and partnering smoothly with product, infra, and research.
Search Research Engineer jobs near San Francisco → Browse all live jobs
This posting was published by Sesame on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.