Cantina

Machine Learning Engineer, Speech - Joint Audio-Video Modeling

$200K–$220KFull-time · Remote (U.S. or Europe)
✓ Verified live on the employer's own system · added 10 days ago
Save search
Mid-level · 5+ yrs exp

Requirements

Experience: 5+ years

Skills & tools

Speech TherapyMultimodalTransformer ModelsMachine LearningIntegration TestingTeam LeadershipResearchProgramming

Benefits — mentioned in this posting

Health, dental & visionPaid time offFamily / parental leave401(k) / retirement
Apply on company site ↗ See your fit → free

Full job description

About the Role:

We're looking for a Research / ML Engineer to join our Speech Team to build state-of-the-art speech and audio generation systems end-to-end from data specs through production inference with a focus on joint audio-video modeling.

You'll own the audio side of multimodal generation: the representations (audio VAEs, neural codecs), the generative backbone (diffusion / flow-matching transformers), and the conditioning and alignment machinery that makes characters speak, sing, and emote in sync with what's on screen. That includes voice cloning and multi-speaker conditioning inside joint AV models, cinematic dialogue with music and sound design, and adjacent speech tasks (controllable TTS, voice conversion) that feed the same stack.

You'll drive the model ↔ data ↔ eval flywheel, partnering closely with research, video, data, and infra to ship fast, reliable, and cost-aware models. In this role you'll work at the intersection of cutting-edge research and practical engineering, contributing to the development of safe, steerable, and trustworthy AI systems.

You will thrive in this role if you:

- See research and engineering as two sides of the same coin and enjoy owning work end-to-end.

- Are excited to work across modalities and collaborate closely with a video generation team rather than staying inside audio.

- Are results-oriented, flexible, and willing to pick up whatever moves the needle.

- Like collaborating closely with infra, data, and product to ship measurable improvements.

- Enjoy designing experiments, listening tests, and metrics that correlate with user-perceived quality.

- Are eager to learn every day, and to find and solve unique large-scale problems.

What You'll Do:

- Audio Representations: Design, train, and improve the audio VAEs, neural codecs, and vocoders our generative models sit on top of latent design, reconstruction and perceptual objectives, compression-vs-fidelity tradeoffs.

- Model Building: Architect, implement, pre-train, fine-tune, and post-train/alignment (e.g., GRPO/DPO) diffusion and flow-matching transformers for large-scale audio and video generation.

- Joint Audio-Video Modeling: Design the audio conditioning and cross-modal alignment inside joint AV models, audio latents alongside video latents, reference-audio and multi-speaker conditioning, multi shot generation audio/video modeling.

- Experimental Design: Design, run, and analyze scientific experiments to advance our understanding of the models.

- Data Ownership: Define data requirements and collaborate on acquisition, curation, AV-sync and quality filtering, annotation quality, and synthetic data strategies for paired audio-video and speech corpora.

- Rigorous Evaluation: Design automated objective/subjective evaluations audio fidelity and intelligibility metrics, AV-sync, listening and viewing tests, robustness & bias checks, and red-team studies.

- Inference Efficiency: Drive distillation, step-count reduction, quantization, and kernel/memory optimization to meet interactive latency and cost targets.

- Pipeline Delivery: Harden the training → evaluation → inference pipeline; profile latency, memory, and cost; and meet production SLAs with robust monitoring and rollback.

- GPU Scaling: Partner with infrastructure to run distributed training/inference on cloud fleets and productionize models with reliability and observability.

- Project Leadership: Independently lead small research projects while collaborating on larger team initiatives, including cross-team work with video generation.

- Tool Development: Develop and improve dev tooling to enhance team productivity.

- Safety &

Responsibility: Contribute to safety/consent guardrails, watermarking, and misuse/abuse mitigation for responsible voice and likeness technology.

What You'll Bring:

- Exceptional research/development experience with large-scale audio models (>8B parameters, >500k hours of data).

- Deep hands-on experience with diffusion and/or flow-matching transformers, including practical knowledge of samplers, schedules, conditioning mechanisms, and distillation.

- Deep hands-on experience training audio VAEs, neural audio codecs, and vocoders latent/tokenizer design, reconstruction and perceptual objectives, adversarial training.

- Strong experience with multi-node, multi-GPU distributed training (FSDP/DeepSpeed or equivalent).

- Strong software engineering skills with a proven track record of building complex systems.

- Strong with PyTorch and performance work (profiling, CUDA/Triton/C++ as needed) and writing reliable production-quality code.

- Shipped large-scale speech/audio or multimodal generative models to production.

- Background in working with large-scale ML data, and the ability to iterate on data and triangulate quality using both subjective and objective signals.

- Experience with voice cloning, speech control/steerability, or expressive speech generation.

- Notable publications and/or open-source contributions in speech/audio/ML.

- Strongly preferred:

- Experience with multimodal audio-video modeling : joint AV generation of multi-shot, multi-speaker scenes with dialogue, music, and sound design generated jointly with video, and the cross-modal alignment that keeps them in sync.

- Experience with video generation : video diffusion/flow-matching transformers, video VAEs, conditioned and multi-shot generation, building data pipelines for video models.

- Streaming or real-time generation, causal distillation (e.g., Self Forcing / Self Forcing++).

Compensation:

The anticipated annual base salary range for this role is between $200,000-$220,000 (€170,000-€190,000). When determining compensation, a number of factors will be considered, including skills, experience, job scope, location, and competitive compensation market data.

Benefits for U.S.-based roles:

- Competitive salary and generous company equity

- Medical, dental, and vision insurance - 99.99% of premiums covered by Cantina

- 42 days of paid time off, including:

- 15 PTO days

- 10 sick days

- 15 company holidays

- 2 floating holidays

- Generous parental leave & fertility support

- 401(k) retirement savings plan

- Lifestyle spending account - $500/month to use however you'd like

- Complimentary lunch and snacks for in-office employees

- One Medical membership, and more!

More jobs at Cantina

Similar jobs near Remote (U.S. or Europe)

Search Machine Learning Engineer, Speech - Joint Audio-Video Modeling jobs near Remote (U.S. or Europe) → Browse all live jobs

This posting was published by Cantina on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.