Skip to main content
Ryan Orban

Ryan Orban

Subject
1 entry

Masked Image Modeling

Bookmarks

  1. EVA: Exploring the Limits of Masked Visual Representation Learning at Scale

    EVA is a 1-billion-parameter vision foundation model from BAAI that achieves state-of-the-art on image classification, detection, and segmentation by pretraining a ViT to reconstruct masked CLIP features. Initializing CLIP's vision tower from EVA dramatically stabilizes training — an important practical finding for building large multimodal systems.

All bookmarks