Subject
3 entries
Vision Language
Bookmarks
Language Models as Models of the Visual World
Research showing language models can use linear projections of image representations as soft prompts for vision-language tasks — without tuning the LM or image encoder. An early signal for the efficiency of frozen model feature reuse in multimodal architectures.
Image as a Foreign Language: BEIT-3 Pretraining for All Vision and Vision-Language Tasks
Wang, Bao, Dong et al. at Microsoft introduce BEIT-3, a general-purpose multimodal foundation model that treats images as a 'foreign language' and applies masked language modeling uniformly across images, text, and image-text pairs. BEIT-3 achieves state-of-the-art across seven vision and vision-language benchmarks including COCO, ImageNet, VQA, and NLVR2.
Flamingo: A Visual Language Model for Few-Shot Learning
DeepMind's Flamingo (2022) bridges a frozen vision encoder and a frozen large language model with cross-attention layers, enabling powerful few-shot vision-language capabilities without retraining either component. It set new few-shot records on image captioning and VQA benchmarks by treating visual inputs as just another type of context for an LM.
