Diffusion LLMs get up to 40x faster inference
A caching technique cuts redundant computation in diffusion-based language models, enabling up to 40x faster inference with no meaningful accuracy loss.
A caching technique cuts redundant computation in diffusion-based language models, enabling up to 40x faster inference with no meaningful accuracy loss.