Skip to main content
Ryan Orban

Ryan Orban

Subject
3 entries

Vision Language

Bookmarks

  1. Language Models as Models of the Visual World

    Research showing language models can use linear projections of image representations as soft prompts for vision-language tasks — without tuning the LM or image encoder. An early signal for the efficiency of frozen model feature reuse in multimodal architectures.

  2. Image as a Foreign Language: BEIT-3 Pretraining for All Vision and Vision-Language Tasks

    Wang, Bao, Dong et al. at Microsoft introduce BEIT-3, a general-purpose multimodal foundation model that treats images as a 'foreign language' and applies masked language modeling uniformly across images, text, and image-text pairs. BEIT-3 achieves state-of-the-art across seven vision and vision-language benchmarks including COCO, ImageNet, VQA, and NLVR2.

  3. Flamingo: A Visual Language Model for Few-Shot Learning

    DeepMind's Flamingo (2022) bridges a frozen vision encoder and a frozen large language model with cross-attention layers, enabling powerful few-shot vision-language capabilities without retraining either component. It set new few-shot records on image captioning and VQA benchmarks by treating visual inputs as just another type of context for an LM.

All bookmarks