MLOps
Knowledge Distillation
Training a small student to match a large teacher's soft output distribution, usually with a temperature-scaled KL term.
Why interviewers ask about it
Distil on *unlabelled production data* so the student trains on the true serving distribution. Often 90–95% of teacher quality at 1–5% of the cost.
Related terms
This term is part of the free AI/ML Engineer interview preparation module - browse the full glossary for every definition.