Skip to main content
Ryan Orban

Ryan Orban

Subject
5 entries

Deepmind

Bookmarks

  1. In-Context Reinforcement Learning with Algorithm Distillation

    Algorithm Distillation trains a causal transformer on sequences of RL learning histories so the model can improve its policy entirely in-context without gradient updates. A key step toward meta-learning agents that get better at RL through experience rather than parameter updates.

  2. Building Safer Dialogue Agents (DeepMind / Sparrow)

    DeepMind's blog post on Sparrow — a dialogue agent trained with reinforcement learning from human feedback and rules to be helpful, harmless, and honest. An early published account of RLHF-based safety fine-tuning for conversational AI.

  3. Introduction to Graph Neural Networks with JAX/jraph

    DeepMind's interactive introduction to Graph Neural Networks using JAX and jraph — covers message passing, graph classification, and node prediction with runnable code. One of the clearest practical GNN tutorials available as a Google Colab notebook.

  4. A Generalist Agent

    Presents Gato, a single transformer that acts as a multi-modal, multi-task, multi-embodiment generalist policy across 600+ tasks including Atari, robotic manipulation, image captioning, and dialogue using identical weights. Demonstrates that scaling language model principles to non-text domains produces surprisingly capable generalist behavior.

  5. Flamingo: A Visual Language Model for Few-Shot Learning

    DeepMind's Flamingo (2022) bridges a frozen vision encoder and a frozen large language model with cross-attention layers, enabling powerful few-shot vision-language capabilities without retraining either component. It set new few-shot records on image captioning and VQA benchmarks by treating visual inputs as just another type of context for an LM.

All bookmarks