Transformers that skip easy tokens save compute
A new transformer design lets layers skip unimportant tokens, cutting compute by half with no quality loss.
A new transformer design lets layers skip unimportant tokens, cutting compute by half with no quality loss.