Subject
1 entry
Masked Image Modeling
Bookmarks
EVA: Exploring the Limits of Masked Visual Representation Learning at Scale
EVA is a 1-billion-parameter vision foundation model from BAAI that achieves state-of-the-art on image classification, detection, and segmentation by pretraining a ViT to reconstruct masked CLIP features. Initializing CLIP's vision tower from EVA dramatically stabilizes training — an important practical finding for building large multimodal systems.
