Signal & Noise
Menu

Mixture of Experts (MoE)

An architecture that activates only a few expert subnetworks per token, giving big-model quality at a fraction of the compute — but all experts must still fit in memory.

Related terms

← Back to the full glossary