Signal & Noise
Menu

Paged attention

vLLM's virtual-memory-style KV-cache management: non-contiguous blocks eliminate fragmentation, fitting far more concurrent requests per GPU.

Related terms

← Back to the full glossary