Subject
1 entry
Vision Transformers
Bookmarks
Probing Vision Transformers
A research repository probing the internal representations and attention mechanisms of Vision Transformers (ViT, DeiT, DINO). Notable finding: self-supervised DINO produces more salient attention maps than supervised models, suggesting better spatial semantics from self-supervision.
