Subject
1 entry
Text to Video
Bookmarks
Make-A-Video: Text-to-Video Generation without Text-Video Data
Singer et al. (Meta AI, 2022) introduce Make-A-Video, a text-to-video generation system that learns spatiotemporal motion from unlabeled video while keeping semantic knowledge from paired image-text data. It sidesteps the absence of large-scale video-caption datasets by decoupling what to generate from how things move.
