Subject
2 entries
Pretraining
Bookmarks
Image as a Foreign Language: BEIT-3 Pretraining for All Vision and Vision-Language Tasks
Wang, Bao, Dong et al. at Microsoft introduce BEIT-3, a general-purpose multimodal foundation model that treats images as a 'foreign language' and applies masked language modeling uniformly across images, text, and image-text pairs. BEIT-3 achieves state-of-the-art across seven vision and vision-language benchmarks including COCO, ImageNet, VQA, and NLVR2.
What Language Model Architecture and Pretraining Objective Work Best for Zero-Shot Generalization?
Wang, Roberts, Scao et al. (BigScience Architecture Group) conduct a large-scale comparison of model architectures (causal decoder, non-causal decoder, encoder-decoder) and pretraining objectives (autoregressive, masked LM) for zero-shot generalization. The key finding: causal decoders + autoregressive LM win at zero-shot; non-causal decoders + masked LM + multitask fine-tuning win overall.
