Distillation
Training a smaller model to imitate a larger one, capturing much of its ability in a cheaper, faster package.
It is how providers ship small models that punch above their size for everyday tasks.
Frequently asked questions
How can a small AI model be almost as good as a big one?
Through distillation, where a smaller model is trained to imitate a larger one and captures much of its ability in a cheaper, faster package. It is how providers ship compact models that punch above their size for everyday tasks.
What is the difference between distillation and quantization?
Distillation trains a brand-new, smaller model to copy a bigger one's behavior, while quantization just compresses an existing model's numbers to save memory. One builds a smaller student, the other slims down the same model.
Why would a company distill a model instead of just using the big one?
Smaller distilled models are faster and cheaper to run, which matters a lot when serving many users. For common tasks they deliver most of the quality at a fraction of the cost and speed of the large original.
New to all this? Start with what an AI agent really is, browse the full glossary, or explore the learning hub.
← Back to the glossary