Subject
2 entries
Video Generation
Bookmarks
Make-A-Video: Text-to-Video Generation without Text-Video Data
Singer et al. (Meta AI, 2022) introduce Make-A-Video, a text-to-video generation system that learns spatiotemporal motion from unlabeled video while keeping semantic knowledge from paired image-text data. It sidesteps the absence of large-scale video-caption datasets by decoupling what to generate from how things move.
Speech Driven Talking Head Generation via Attentional Landmarks Based Representation
This paper introduces an attentional landmark-based representation for generating realistic talking head video from speech audio, using facial landmarks as a compact intermediate representation that bridges audio and visual domains. The approach decouples appearance generation from motion modeling, improving generalization across identities.
