Data is one of the most critical inputs to frontier AI. In this role, you'll work across the full data stack - from infrastructure and web-scale data pipelines to modeling and evaluation. You'll develop a deep understanding of how data influences model performance and leverage that knowledge to advance the frontier of multimodal datasets to create new capabilities for AI.
- Design and build high-quality datasets for model training and run controlled modeling experiments to measure their impact on model performance and behavior.
- Engineer web-scale data pipelines and build systems to annotate and ensure data quality at scale.
- Develop techniques for post-training and synthetic data generation to improve model quality and intelligence.
- Experience building or working with large multilingual datasets
- Experience with generative models (speech, text, or multimodal).
- Ability to help guide human annotation and evaluation across multiple languages.
- Strong applied ML background with a focus on data-centric approaches.
- Excitement for building scalable systems that bridge research and production.
Search Research Engineer, Audio Understanding jobs near *HQ - San Francisco, CA → Browse all live jobs
This posting was published by Cartesia on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.