Skip to main content
Ryan Orban

Ryan Orban

Subject
2 entries

Video Generation

Bookmarks

  1. Make-A-Video: Text-to-Video Generation without Text-Video Data

    Singer et al. (Meta AI, 2022) introduce Make-A-Video, a text-to-video generation system that learns spatiotemporal motion from unlabeled video while keeping semantic knowledge from paired image-text data. It sidesteps the absence of large-scale video-caption datasets by decoupling what to generate from how things move.

  2. Speech Driven Talking Head Generation via Attentional Landmarks Based Representation

    This paper introduces an attentional landmark-based representation for generating realistic talking head video from speech audio, using facial landmarks as a compact intermediate representation that bridges audio and visual domains. The approach decouples appearance generation from motion modeling, improving generalization across identities.

All bookmarks