Subject
1 entry
Knowledge Distillation
Bookmarks
Task-Specific Knowledge Distillation for BERT
A tutorial on task-specific knowledge distillation for BERT using Hugging Face Transformers and Amazon SageMaker — compressing a 109M-parameter BERT-base into a 4M-parameter student with 90%+ performance retention. Demonstrates that you don't need a massive model in production if you can distill task knowledge from one.
