Today's digest
How it works
Many cheap reasoning drafts are sampled via diffusion. A reward model scores each step, and the best steps across all drafts are stitched into one composite answer.