Speculative Decoding in vLLM: Draft, Verify, and Cut Latency Without Losing the Distribution
Large language model inference has an awkward performance profile: the GPU does enormous math, then waits. Every token requires a
Large language model inference has an awkward performance profile: the GPU does enormous math, then waits. Every token requires a