Subject
6 entries
Bert
Bookmarks
BERTopic: BERT-Based Topic Modeling
BERTopic is the leading open-source topic modeling library using sentence embeddings and clustering rather than word co-occurrence statistics — produces coherent, human-readable topics that LDA-style models often can't. The changelog tracks its evolution as the library added new backends and features.
A Simple but Effective BERT Model for Dialog State Tracking on Resource-Limited Systems
A BERT-based dialog state tracking model optimized for resource-limited systems — achieving competitive performance on MultiWOZ while reducing model size and inference cost. Demonstrates that careful architecture choices can close the gap between full-scale models and edge-deployable alternatives.
Pre-Trained Models: Past, Present and Future
Comprehensive survey of large-scale pre-trained models (PTMs) tracing the evolution from BERT and GPT through four research frontiers: architecture, contextual use, efficiency, and interpretability. Required reading for understanding how self-supervised pre-training became the unified backbone of modern AI.
Gradient Explanations for HuggingFace BERT Classification
A tutorial by Victor Dibia on generating gradient-based explanations for HuggingFace BERT text classification models in TensorFlow 2.0 — visualizing which tokens most influenced the model's prediction. Explainability for transformer classifiers was a practical gap in 2022 since attention maps alone are insufficient.
BERTopic: The Future of Topic Modeling
A Pinecone explainer on BERTopic — a topic modeling library that uses transformer embeddings instead of bag-of-words statistics, producing semantically coherent topics. BERTopic largely made LDA obsolete for practitioners who have access to modern embeddings.
Task-Specific Knowledge Distillation for BERT
A tutorial on task-specific knowledge distillation for BERT using Hugging Face Transformers and Amazon SageMaker — compressing a 109M-parameter BERT-base into a 4M-parameter student with 90%+ performance retention. Demonstrates that you don't need a massive model in production if you can distill task knowledge from one.
