<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Diffusion on Ryan Orban</title><link>https://ryanorban.com/categories/diffusion/</link><description>Recent content in Diffusion on Ryan Orban</description><generator>Hugo</generator><language>en-us</language><managingEditor>me@ryanorban.com (Ryan Orban)</managingEditor><webMaster>me@ryanorban.com (Ryan Orban)</webMaster><copyright>Ryan Orban</copyright><lastBuildDate>Thu, 29 Sep 2022 00:00:00 +0000</lastBuildDate><atom:link href="https://ryanorban.com/categories/diffusion/index.xml" rel="self" type="application/rss+xml"/><item><title>Make-A-Video: Text-to-Video Generation without Text-Video Data</title><link>https://ryanorban.com/notes/make-a-video-text-to-video-generation/</link><pubDate>Thu, 29 Sep 2022 00:00:00 +0000</pubDate><author>me@ryanorban.com (Ryan Orban)</author><guid>https://ryanorban.com/notes/make-a-video-text-to-video-generation/</guid><description>&lt;h3 id="summary" class="scroll-mt-8 group"&gt;
 Summary
 
 &lt;a href="#summary"
 class="no-underline hidden opacity-50 hover:opacity-100 !text-inherit group-hover:inline-block"
 aria-hidden="true" title="Link to this heading" tabindex="-1"&gt;
 &lt;svg
 xmlns="http://www.w3.org/2000/svg"
 width="16"
 height="16"
 fill="none"
 stroke="currentColor"
 stroke-linecap="round"
 stroke-linejoin="round"
 stroke-width="2"
 class="lucide lucide-link w-4 h-4 block"
 viewBox="0 0 24 24"
&gt;
 &lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71" /&gt;
 &lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71" /&gt;
&lt;/svg&gt;

 &lt;/a&gt;
 
&lt;/h3&gt;
&lt;p&gt;Uriel Singer, Adam Polyak, and the Meta AI team present Make-A-Video — a system for generating short video clips from text prompts, trained without paired text-video data. The key insight is a clever decomposition of the problem: the model learns &lt;em&gt;what the world looks like&lt;/em&gt; from large-scale text-image pairs (where paired data is abundant), then separately learns &lt;em&gt;how things move&lt;/em&gt; from unlabeled video footage. This avoids the fundamental data bottleneck that had stalled text-to-video generation research: video paired with captions is expensive to collect at scale, but raw video and image-caption datasets exist in enormous quantities.&lt;/p&gt;</description></item></channel></rss>