Skip to main content
Ryan Orban

Ryan Orban

Subject
21 entries

Computer Vision

Bookmarks

  1. SlouchSniper: AI Posture App

    SlouchSniper uses on-device computer vision to monitor posture via webcam and dims your screen when you slouch, restoring it when you sit up. A behavioral nudge approach to building postural habits — no cloud, one-time purchase.

  2. A Visual Guide to Vision Transformers

    A visual guide to Vision Transformers (ViT) — explains how the transformer architecture is adapted for images, covering patch embeddings, position encodings, and attention in visual domains with diagrams. Good complement to the original ViT paper for building intuition.

  3. EditAnything: Segment Anything + Stable Diffusion for Image Editing

    EditAnything combines Meta's Segment Anything Model with Stable Diffusion to enable precise region-based image editing — click to select any object, then replace or transform it with a text prompt. One of the first practical applications of SAM.

  4. Human Motion Diffusion Model (MDM)

    Tevet et al. (Tel Aviv University, arXiv:2209.14916, 2022) apply diffusion models to human motion generation, producing MDM — a transformer-based denoiser that generates realistic motion sequences from text descriptions or action labels. It matters because it extends the generative power of diffusion to a structured temporal domain, enabling controllable motion editing that prior methods couldn't match.

  5. Twelve Labs — Video Understanding API

    Twelve Labs provides a video understanding API that lets developers search, retrieve, and understand video content semantically — as if the model could actually watch and comprehend it. Fills the gap between text search and the dense information in video.

  6. NeRF: Neural Radiance Fields

    The original NeRF (Neural Radiance Fields) project page from Matthew Tancik's site — the foundational 2020 paper that represents 3D scenes as neural functions and synthesizes novel views via volumetric rendering. A landmark that spawned a field.

  7. Make-A-Video: Text-to-Video Generation without Text-Video Data

    Singer et al. (Meta AI, 2022) introduce Make-A-Video, a text-to-video generation system that learns spatiotemporal motion from unlabeled video while keeping semantic knowledge from paired image-text data. It sidesteps the absence of large-scale video-caption datasets by decoupling what to generate from how things move.

  8. Will Transformers Take Over Artificial Intelligence?

    Quanta Magazine's 2022 look at whether transformer architectures will dominate all of AI — following their success in NLP and early incursion into image classification. A useful time-capsule of the moment when the transformer paradigm started feeling inevitable.

  9. Block-NeRF: City-Scale Neural Neighborhoods

    Block-NeRF from Waymo and UC Berkeley extends Neural Radiance Fields to city-scale scenes by dividing them into individually trained blocks that stitch together. A step toward photorealistic neural reconstruction of entire neighborhoods from street-level imagery.

  10. Speech Driven Talking Head Generation via Attentional Landmarks Based Representation

    This paper introduces an attentional landmark-based representation for generating realistic talking head video from speech audio, using facial landmarks as a compact intermediate representation that bridges audio and visual domains. The approach decouples appearance generation from motion modeling, improving generalization across identities.

  11. AD-NeRF: Audio Driven Neural Radiance Fields for Talking Head Synthesis

    AD-NeRF generates photorealistic talking-head video directly from audio using neural radiance fields, bypassing the 2D landmarks or 3D face model intermediaries used by prior methods. By conditioning an implicit neural function on audio features and rendering via volume rendering, it achieves both head and upper body generation with free-viewpoint control.

  12. Alien Dreams: CLIP-Guided Image Generation

    UC Berkeley ML blog's 2021 post on using CLIP for guided image generation — an early exploration of the CLIP+VQGAN/diffusion pipeline that preceded Stable Diffusion. Historically significant as a snapshot of generative AI before it became mainstream.

  13. CS231n: Convolutional Neural Networks for Visual Recognition

    Stanford CS231n: Convolutional Neural Networks for Visual Recognition — Andrej Karpathy's course that became the de facto entry point into deep learning for computer vision. The lecture notes remain among the best written explanations of CNNs, backprop, and training practice.

  14. DeepDream: How Alexander Mordvintsev Excavated the Computer's Hidden Layers

    The story behind Google's DeepDream — how researcher Alexander Mordvintsev discovered that running gradient ascent on a convolutional network's hidden layers produces psychedelic imagery that reveals what features the network learned. A landmark moment in neural network interpretability.

  15. Superpixel Segmentation — IVRL

    The IVRL lab at EPFL's research page on superpixel segmentation — home of the SLIC algorithm, which became the dominant superpixel method due to its speed and perceptual uniformity. Superpixels are a fundamental preprocessing step in classical computer vision pipelines.

  16. Yann LeCun: Making Facebook's AI Predict What Happens in Videos

    New Scientist interview with Yann LeCun on Facebook AI Research's goal to build models that predict what will happen in videos — covering what AI can and can't do in 2015, and LeCun's view on unsupervised learning as the key unsolved problem.

  17. What a Deep Neural Network Thinks About Your Selfie

    Andrej Karpathy trained a VGGNet on 2 million Instagram selfies to learn what makes a selfie 'good' — using likes-per-follower as the quality signal. Beyond the entertainment value, it's a sharp demonstration of how supervised learning can proxy for human aesthetic judgment at scale.

  18. OverFeat: Integrated Recognition, Localization, and Detection

    OverFeat is NYU CILVR lab's deep learning framework for object recognition, localization, and detection using convolutional networks — won the ImageNet 2013 localization task. An important 2014 artifact of Yann LeCun's group.

  19. WebCamMesh — 3D Webcam Visualization

    WebCamMesh uses WebGL and the browser's webcam API to create a real-time 3D mesh from the live video feed — mapping webcam pixel brightness to vertex displacement. An early example of combining browser camera access with WebGL for real-time visual effects.

  20. Gesticulate — Gesture Interface

    Gesticulate was a 2012 gesture-recognition interface project — likely an early prototype using webcam or depth sensor (Kinect era) input for gestural control. Part of the wave of natural user interface experimentation following Microsoft Kinect's 2010 success.

  21. Kittydar: JavaScript Cat Face Detection

    Kittydar is a JavaScript library for cat face detection in images — a real implementation of neural-network-based object detection in the browser, released in 2012 when running ML in JavaScript was novel and this kind of project showed what was becoming possible.

All bookmarks