Mixture of Experts (MoE)
An architecture that activates only a few expert subnetworks per token, giving big-model quality at a fraction of the compute — but all experts must still fit in memory.
An architecture that activates only a few expert subnetworks per token, giving big-model quality at a fraction of the compute — but all experts must still fit in memory.