Speculative Decoding: Getting 2-3x More Out of Every GPU Pass
Watch a GPU while a large language model generates text and you’ll see something strange: for most of every forward
Continue readingSpeculative Decoding: Getting 2-3x More Out of Every GPU Pass