Subject
1 entry
Soft Prompts
Bookmarks
Language Models as Models of the Visual World
Research showing language models can use linear projections of image representations as soft prompts for vision-language tasks — without tuning the LM or image encoder. An early signal for the efficiency of frozen model feature reuse in multimodal architectures.
