Subject
36 entries
Generative Ai
Bookmarks
Websim: AI-Generated Interactive Web Experiences
Websim is a platform for AI-generating interactive web experiences — describe a simulation, game, or tool and it renders a working browser app. Community-driven with sharing, trending projects, and social features. Now at websim.com.
Genie: Generative Interactive Environments
Genie is a Google DeepMind foundation world model that generates playable, action-controllable interactive environments from a single image prompt — photo, sketch, or AI-generated. Trained on unlabeled internet videos without action annotations, it discovers a transferable latent action vocabulary.
Fine-tuning language models: a practical overview
A clear introductory overview of LLM fine-tuning from the GenAI Guidebook — covers why and when to fine-tune, the mechanics of weight updates, and the major techniques including LoRA and RLHF.
Generative AI's Act Two — Sequoia Capital
Sequoia's 'Generative AI Act Two' argues that one year after the ChatGPT moment, the real work begins: moving from demos and experiments to applications that deliver genuine business value. A VC thesis on where the durable opportunities in the generative AI wave actually lie.
Cradle: AI Platform for Protein Engineering
Cradle is an AI platform for protein engineering and design — uses generative models to suggest protein sequence modifications that improve properties like stability, expression, and activity. Targets the biotech workflow where engineers iteratively improve proteins through wet lab cycles.
The Data Moat Myth for Generative AI
A 2022 tweet arguing that the 'data moat' premise for GenAI startups is a false promise — foundational models commoditize the data advantage that once differentiated ML companies. An early articulation of why the 'ChatGPT for X with proprietary data' pitch was flawed.
Prompt Parrot — Replicate
Prompt Parrot is a Replicate-hosted model that generates creative variations of text-to-image prompts. A simple tool for exploring the prompt space around a seed idea when using Stable Diffusion or similar generators.
Magic3D — High-Resolution Text-to-3D Content Creation
Magic3D is NVIDIA Research's text-to-3D content creation method that generates high-resolution 3D meshes from text prompts using a two-stage coarse-to-fine optimization. An important step beyond NeRF-based generation toward production-usable 3D assets.
AI and I: The Age of Artificial Creativity
A NessLabs essay on what AI creative tools mean for human creativity and knowledge work — arguing that AI doesn't replace creativity but reshapes what creative work consists of. Published at the start of the generative AI wave in late 2022.
An Image is Worth One Word: Personalizing Text-to-Image Generation Using Textual Inversion
Tel Aviv University and NVIDIA paper introducing Textual Inversion — learning a single new text embedding token that represents a user-provided concept, enabling that concept to be composed into any text prompt. Showed that the embedding space of text-to-image models is richly structured and can be expanded with just 3-5 example images.
Stable Diffusion Parameters Guide
A practical overview of the key parameters for controlling AI image generation in Stable Diffusion — steps, CFG scale, sampler, seed, and dimensions. A useful reference from the early days when these knobs were being collectively figured out.
Generative AI: A Creative New World (Sequoia Capital)
Sequoia Capital's October 2022 essay framing generative AI as a new creative platform wave — arguably the most-cited VC framing of the early generative AI moment. Published right as Stable Diffusion proliferated and just before ChatGPT, it set the terms for how VCs discussed the space for the next two years.
NEW MODELS
New Models is a media project by Laboria Cuboniks collective members exploring the cultural and aesthetic dimensions of AI — intersecting tech criticism, art, and speculative theory. Sits at the AI-culture overlap rather than pure technical coverage.
Generrated: DALL-E 2 Image Gallery
Generrated is a gallery of 7,040 images generated with DALL-E 2 prompts — a browsable collection for studying what prompt structures produce what visual outputs. A practical prompt engineering reference disguised as an art gallery.
AI Generative Art Tools (Pharmapsychotic)
Pharmapsychotic's comprehensive catalog of AI generative art tools — covering text-to-image, image editing, upscaling, and animation tools available as of late 2022. One of the most widely-shared resource lists during the early Stable Diffusion era.
The AI Unbundling (Stratechery)
Ben Thompson's September 2022 analysis arguing that AI would unbundle integrated products by making the creation layer cheap — anyone could build a specialized alternative to an incumbent's bundled offering. One of the clearest early strategic framings of generative AI's market impact.
Online Art Communities Begin Banning AI-Generated Images
Andy Baio's September 2022 report documenting the wave of art platform bans on AI-generated images — ArtStation, DeviantArt, Newgrounds, and others responding to artist backlash over Stable Diffusion. The first major cultural reckoning over what generative image models meant for creative communities.
AI Content Generation, Part 1: Machine Learning Basics (Jon Stokes)
Jon Stokes' accessible introduction to machine learning as the foundation of AI content generation — Part 1 of a series aimed at readers with no ML background. Stokes covers the statistical learning framing without requiring math, making it one of the better on-ramps for non-technical audiences.
OpenArt: AI Image Discovery and Prompt Library
OpenArt is a community gallery and prompt discovery platform for AI-generated images, showing each image alongside the prompt that created it. Launched during DALL-E 2's invite-only period, it became an early resource for learning what prompts produce what visual results.
Lexica: Stable Diffusion Search Engine
Lexica is a search engine for Stable Diffusion images and the prompts that generated them. It became the go-to reference for empirical prompt knowledge — what prompts produce what visual aesthetics — and later added its own image generation feature.
DALL-E 2 vs. Midjourney vs. Stable Diffusion Comparison
A viral August 2022 Twitter mega-thread comparing DALL-E 2, Midjourney, and Stable Diffusion across photography, illustration, and abstract styles. Captures the precise moment all three image synthesis tools were accessible simultaneously — each with a distinctive aesthetic 'sound.'
Gene Kogan's Stable Diffusion Collage Tool (WIP)
Gene Kogan's WIP collage tool built on Stable Diffusion — a hybrid workflow where human collage composition meets AI generation. Kogan is a prominent artist-researcher at the machine learning × creative practice intersection, and this was shared days after SD's open release.
Lucid Sonic Dreams: Audio-Reactive GAN Visualization
Lucid Sonic Dreams syncs StyleGAN2-generated visuals to music, producing audio-reactive video from a song and style choice. Made GAN art accessible to musicians and VJs as a simple pip-installable Python package.
Stable Diffusion: CompVis Open-Source Release
The CompVis open-source release of Stable Diffusion — the text-to-image model that democratized AI image generation. This marked the moment state-of-the-art image synthesis left the walled gardens of DALL-E 2 and Midjourney and became freely runnable on consumer hardware.
Stable Diffusion Initial Release Announcement
The announcement tweet for Stable Diffusion's initial model checkpoint release in August 2022 — weights available for research upon request, with a more permissive release and inpainting coming. Captures the moment the open-source AI image generation era began.
Fontjoy: Deep Learning Font Pairing
Fontjoy uses deep learning to generate font pairings — one click finds combinations of Google Fonts that work together aesthetically. A practical tool for non-designers who need readable typography without expert knowledge.
Craiyon (formerly DALL-E mini)
Craiyon, formerly DALL-E mini, is a free web-based text-to-image generator that went viral before Stable Diffusion's release. It gave millions their first hands-on experience with AI image generation, despite producing lower-quality images than commercial alternatives.
Brandmark: AI Logo Generation
Brandmark is an AI-powered logo and brand identity generator — type a name and industry, get a logo, business cards, and social graphics in under a minute. An early example of generative AI applied to commercial design that found product-market fit before the diffusion model wave.
The DALL-E 2 Prompt Book
Unofficial visual reference guide by Guy Parsons (dallery.gallery) covering how to prompt DALL-E 2 across photography, illustration styles, art history movements, 3D artwork, and editing techniques. A practical taxonomy of the prompt space for the first widely-accessible diffusion image model.
The DALL-E 2 Prompt Book
The DALL-E 2 Prompt Book was an early community-produced guide to prompt engineering for image generation — cataloging styles, artists, modifiers, and composition techniques that reliably produce specific visual outputs. A snapshot of the craft before the field became saturated with guides.
How Sber Built ruDALL-E — Interview with Sergei Markov
Serokell's interview with Sergei Markov of SberDevices about building ruDALL-E — a 12B parameter Russian-language text-to-image model. Covers the engineering and research challenges of training massive multimodal models, plus the open-source culture argument in ML.
AD-NeRF: Audio Driven Neural Radiance Fields for Talking Head Synthesis
AD-NeRF generates photorealistic talking-head video directly from audio using neural radiance fields, bypassing the 2D landmarks or 3D face model intermediaries used by prior methods. By conditioning an implicit neural function on audio features and rendering via volume rendering, it achieves both head and upper body generation with free-viewpoint control.
Alien Dreams: CLIP-Guided Image Generation
UC Berkeley ML blog's 2021 post on using CLIP for guided image generation — an early exploration of the CLIP+VQGAN/diffusion pipeline that preceded Stable Diffusion. Historically significant as a snapshot of generative AI before it became mainstream.
Deep Daze: Text to Image with CLIP and Siren
Deep Daze is Phil Wang's early text-to-image tool combining OpenAI's CLIP with Siren (implicit neural representations) — one of the first accessible open-source implementations of text-guided image generation, predating DALL-E and Stable Diffusion by over a year.
State-of-the-Art Image Generative Models (2021)
Aran Komatsuzaki's March 2021 survey of state-of-the-art image generative models — covering BigGAN, VQVAE-2, DALL-E, CLIP, and diffusion models just as they were emerging. A historical snapshot of the field one year before diffusion models took over.
Automatic Sample Layout (VAE)
Kyle McDonald's EYEO 2016 demonstration of automatic audio sample organization using a Variational Autoencoder — the VAE learns a latent representation of sounds and arranges them in 2D space so similar samples cluster together. An early application of generative models to creative tools.
