Skip to main content
Ryan Orban

Ryan Orban

Subject
18 entries

Image Generation

Bookmarks

  1. ControlNet: Precise Spatial Control for Diffusion Models

    ControlNet adds fine-grained spatial control to Stable Diffusion — use edge maps, depth maps, pose skeletons, or sketches to precisely direct where objects and structures appear in generated images. A major step beyond text-only prompting for image generation.

  2. Gauss: Native macOS Stable Diffusion App

    Gauss is a native macOS app for running Stable Diffusion locally — built by Jake Teton-Landis using Swift/SwiftUI and Apple's Core ML stack to run image generation on Apple Silicon without Python. One of the early native Mac SD apps before AUTOMATIC1111 and ComfyUI dominated.

  3. Stable Diffusion Parameters Guide

    A practical overview of the key parameters for controlling AI image generation in Stable Diffusion — steps, CFG scale, sampler, seed, and dimensions. A useful reference from the early days when these knobs were being collectively figured out.

  4. The Illustrated Stable Diffusion

    Jay Alammar's visual explainer of how Stable Diffusion works under the hood — covering latent diffusion, the CLIP text encoder, and the U-Net denoiser. Alammar's illustrated series is one of the best entry points for building intuition about complex ML architectures.

  5. MagicPrompt-Stable-Diffusion

    MagicPrompt-Stable-Diffusion is a GPT-2-based model fine-tuned to generate effective prompts for Stable Diffusion. It solves the prompt engineering problem for image generation: given a simple idea, it produces elaborate prompt text that reliably produces better images.

  6. Generrated: DALL-E 2 Image Gallery

    Generrated is a gallery of 7,040 images generated with DALL-E 2 prompts — a browsable collection for studying what prompt structures produce what visual outputs. A practical prompt engineering reference disguised as an art gallery.

  7. AI Generative Art Tools (Pharmapsychotic)

    Pharmapsychotic's comprehensive catalog of AI generative art tools — covering text-to-image, image editing, upscaling, and animation tools available as of late 2022. One of the most widely-shared resource lists during the early Stable Diffusion era.

  8. Online Art Communities Begin Banning AI-Generated Images

    Andy Baio's September 2022 report documenting the wave of art platform bans on AI-generated images — ArtStation, DeviantArt, Newgrounds, and others responding to artist backlash over Stable Diffusion. The first major cultural reckoning over what generative image models meant for creative communities.

  9. CLIP Interrogator

    CLIP Interrogator by pharmapsychotic is a Google Colab tool that reverse-engineers what prompt would produce a given image — using CLIP to describe an image in terms that Stable Diffusion understands. The go-to tool in 2022 for figuring out how to replicate an image style.

  10. Krea.ai: AI Prompt Explorer and Gallery

    Krea.ai is a prompt exploration and gallery tool for Stable Diffusion — browse millions of AI-generated images with their prompts, and use those prompts as starting points for your own generations. The community discovery layer that was missing from Stable Diffusion's launch.

  11. Textual Inversion for Stable Diffusion

    A patch enabling textual inversion in Stable Diffusion — the technique of learning a new text token that represents a custom concept (a person, style, or object) from just 3-5 example images. Textual inversion was the first practical method for personalizing Stable Diffusion without full fine-tuning.

  12. Upscayl: Free and Open-Source AI Image Upscaler

    Upscayl is a free, open-source AI image upscaler for Linux, macOS, and Windows built with a Linux-first philosophy. It wraps Real-ESRGAN and similar super-resolution models in a polished desktop UI, making AI upscaling accessible without command-line knowledge.

  13. OpenArt: AI Image Discovery and Prompt Library

    OpenArt is a community gallery and prompt discovery platform for AI-generated images, showing each image alongside the prompt that created it. Launched during DALL-E 2's invite-only period, it became an early resource for learning what prompts produce what visual results.

  14. Stable Diffusion: CompVis Open-Source Release

    The CompVis open-source release of Stable Diffusion — the text-to-image model that democratized AI image generation. This marked the moment state-of-the-art image synthesis left the walled gardens of DALL-E 2 and Midjourney and became freely runnable on consumer hardware.

  15. Craiyon (formerly DALL-E mini)

    Craiyon, formerly DALL-E mini, is a free web-based text-to-image generator that went viral before Stable Diffusion's release. It gave millions their first hands-on experience with AI image generation, despite producing lower-quality images than commercial alternatives.

  16. The DALL-E 2 Prompt Book

    Unofficial visual reference guide by Guy Parsons (dallery.gallery) covering how to prompt DALL-E 2 across photography, illustration styles, art history movements, 3D artwork, and editing techniques. A practical taxonomy of the prompt space for the first widely-accessible diffusion image model.

  17. The DALL-E 2 Prompt Book

    The DALL-E 2 Prompt Book was an early community-produced guide to prompt engineering for image generation — cataloging styles, artists, modifiers, and composition techniques that reliably produce specific visual outputs. A snapshot of the craft before the field became saturated with guides.

  18. State-of-the-Art Image Generative Models (2021)

    Aran Komatsuzaki's March 2021 survey of state-of-the-art image generative models — covering BigGAN, VQVAE-2, DALL-E, CLIP, and diffusion models just as they were emerging. A historical snapshot of the field one year before diffusion models took over.

All bookmarks