Signal & Noise
Menu

Speculative decoding

A small draft model proposes several tokens; the large model verifies them in one parallel pass. Identical output, 2–3x faster — a pure serving win.

Related terms

← Back to the full glossary