Cartesia

Applied Researcher, Audio

Full-time · *HQ - San Francisco, CA
✓ Verified live on the employer's own system · added 327 days ago
Save search Apply on company site ↗ See your fit → free

Full job description

You will be responsible for leading research at the frontier of realtime conversation and human AI interaction. You will contribute across the stack to novel architectures, data, and evals for realtime audio, and translate them into state-of-the-art models used in voice agents around the world.

- Architect and develop new architectures for realtime audio understanding, generation, and speech-to-speech models that reason jointly over multiple modalities in realtime

- Contribute to frontier multimodal and multilingual datasets for pre-training and post-training, including curating data mixes and developing new methods for synthetic data generation and annotation

- Set new standards for how we evaluate and benchmark our audio models

- Strong applied mindset and ability to balance scientific novelty with product impact.

- Excited and able to work across the stack from infra, to data, to evals, to architecture to solve customer problems and build state-of-the-art models.

- Deep expertise in deep generative modeling. Previous experience in audio understanding, audio generation, speech-to-speech, or language modeling preferred but not required.

- Experience with large-scale training, GPU/TPU acceleration, and model optimization.

More jobs at Cartesia

Similar jobs near *HQ - San Francisco, CA

Tell me when more Machine Learning Researcher, Audio jobs post near San Francisco We re-check every listing against the employer’s own board — no résumé needed.

Search Applied Researcher, Audio jobs near *HQ - San Francisco, CA → Browse all live jobs

This posting was published by Cartesia on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.