Autoregressive
Generating output one token at a time, each conditioned on everything before it. Why LLM latency scales with answer length, and why the first token takes longest to appear.
Generating output one token at a time, each conditioned on everything before it. Why LLM latency scales with answer length, and why the first token takes longest to appear.