Subject
7 entries
Multimodal
Bookmarks
Twelve Labs — Video Understanding API
Twelve Labs provides a video understanding API that lets developers search, retrieve, and understand video content semantically — as if the model could actually watch and comprehend it. Fills the gap between text search and the dense information in video.
Improving Multimodal Interactive Agents with RLHF
The Interactive Agents Team at DeepMind (arXiv:2211.11602, 2022) applies RLHF to agents that must understand language instructions and act in visual environments, using human preference feedback to train a reward model for PPO-based RL. The result demonstrates that RLHF substantially improves instruction-following in embodied multimodal settings beyond what supervised learning alone achieves.
ASIF: Coupled Data Turns Unimodal Models to Multimodal Without Training
Norelli, Fumero, Maiorca, Rodolà et al. (Sapienza University, arXiv:2210.01738, 2022) show that any two unimodal models can be composed into a zero-shot multimodal system by finding approximate shared nearest neighbors across their embedding spaces, given only a small set of coupled pairs. The result challenges the assumption that multimodal capability requires joint training.
Language Models as Models of the Visual World
Research showing language models can use linear projections of image representations as soft prompts for vision-language tasks — without tuning the LM or image encoder. An early signal for the efficiency of frozen model feature reuse in multimodal architectures.
Image as a Foreign Language: BEIT-3 Pretraining for All Vision and Vision-Language Tasks
Wang, Bao, Dong et al. at Microsoft introduce BEIT-3, a general-purpose multimodal foundation model that treats images as a 'foreign language' and applies masked language modeling uniformly across images, text, and image-text pairs. BEIT-3 achieves state-of-the-art across seven vision and vision-language benchmarks including COCO, ImageNet, VQA, and NLVR2.
Flamingo: A Visual Language Model for Few-Shot Learning
DeepMind's Flamingo (2022) bridges a frozen vision encoder and a frozen large language model with cross-attention layers, enabling powerful few-shot vision-language capabilities without retraining either component. It set new few-shot records on image captioning and VQA benchmarks by treating visual inputs as just another type of context for an LM.
multimodal.art
multimodal.art is a gallery and community for AI-generated art from multimodal models — DALL-E, CLIP-guided diffusion, and related text-to-image systems. Captured the early 2022 moment before Stable Diffusion democratized AI image generation.
