Unlocking Lossless Speedups in LLMs via Discrete Diffusion
Autoregressive models deliver high-quality generation but remain slow due to sequential decoding. Diffusion language models promise much faster inference through parallel generation, yet current systems such as Mercury2, Diffusion Gemma, and Nemotron Diffusion still lag behind frontier autoregressive models in quality.
This talk presents a discrete diffusion approach that unlocks provably lossless speedups in LLM inference. The method is a drop-in replacement for existing training pipelines, preserves model quality, and significantly improves throughput across levels of parallelism during inference. We show that it outperforms existing diffusion baselines qualitatively while running substantially faster than autoregressive decoding.
This talk is part of Cohere Labs in Conversation, a limited series of talks, in which Cohere Labs scientists and engineers host external researchers for technical talks and Q&A discussions on subjects related to our current explorations at Cohere Labs. We look forward to sharing these talks with you, giving you a glimpse into the problems we're exploring, and learning together from some of the greatest minds in the field.
Speaker Details
Subham Sahoo
Saurabh Dash
Event Topic
TechnologyRelevant Audiences
All State and Local Government, All Federal Government