Subject
5 entries
Deepmind
Bookmarks
In-Context Reinforcement Learning with Algorithm Distillation
Algorithm Distillation trains a causal transformer on sequences of RL learning histories so the model can improve its policy entirely in-context without gradient updates. A key step toward meta-learning agents that get better at RL through experience rather than parameter updates.
Building Safer Dialogue Agents (DeepMind / Sparrow)
DeepMind's blog post on Sparrow — a dialogue agent trained with reinforcement learning from human feedback and rules to be helpful, harmless, and honest. An early published account of RLHF-based safety fine-tuning for conversational AI.
Introduction to Graph Neural Networks with JAX/jraph
DeepMind's interactive introduction to Graph Neural Networks using JAX and jraph — covers message passing, graph classification, and node prediction with runnable code. One of the clearest practical GNN tutorials available as a Google Colab notebook.
A Generalist Agent
Presents Gato, a single transformer that acts as a multi-modal, multi-task, multi-embodiment generalist policy across 600+ tasks including Atari, robotic manipulation, image captioning, and dialogue using identical weights. Demonstrates that scaling language model principles to non-text domains produces surprisingly capable generalist behavior.
Flamingo: A Visual Language Model for Few-Shot Learning
DeepMind's Flamingo (2022) bridges a frozen vision encoder and a frozen large language model with cross-attention layers, enabling powerful few-shot vision-language capabilities without retraining either component. It set new few-shot records on image captioning and VQA benchmarks by treating visual inputs as just another type of context for an LM.
