Decoding
The token-by-token generation phase of inference, after the prompt is processed. Sequential and memory-bound, it's where most serving latency lives.
The token-by-token generation phase of inference, after the prompt is processed. Sequential and memory-bound, it's where most serving latency lives.