Distillation
Training a small model on a large model's outputs so the small one approximates the big one's behavior on a narrower task at a fraction of the cost.
Training a small model on a large model's outputs so the small one approximates the big one's behavior on a narrower task at a fraction of the cost.