<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Q-Learning on Ryan Orban</title><link>https://ryanorban.com/categories/q-learning/</link><description>Recent content in Q-Learning on Ryan Orban</description><generator>Hugo</generator><language>en-us</language><managingEditor>me@ryanorban.com (Ryan Orban)</managingEditor><webMaster>me@ryanorban.com (Ryan Orban)</webMaster><copyright>Ryan Orban</copyright><lastBuildDate>Wed, 27 Apr 2022 00:00:00 +0000</lastBuildDate><atom:link href="https://ryanorban.com/categories/q-learning/index.xml" rel="self" type="application/rss+xml"/><item><title>Reinforcement Learning: An Introduction</title><link>https://ryanorban.com/notes/reinforcement-learning-an-introduction/</link><pubDate>Wed, 27 Apr 2022 00:00:00 +0000</pubDate><author>me@ryanorban.com (Ryan Orban)</author><guid>https://ryanorban.com/notes/reinforcement-learning-an-introduction/</guid><description>&lt;h3 id="summary" class="scroll-mt-8 group"&gt;
 Summary
 
 &lt;a href="#summary"
 class="no-underline hidden opacity-50 hover:opacity-100 !text-inherit group-hover:inline-block"
 aria-hidden="true" title="Link to this heading" tabindex="-1"&gt;
 &lt;svg
 xmlns="http://www.w3.org/2000/svg"
 width="16"
 height="16"
 fill="none"
 stroke="currentColor"
 stroke-linecap="round"
 stroke-linejoin="round"
 stroke-width="2"
 class="lucide lucide-link w-4 h-4 block"
 viewBox="0 0 24 24"
&gt;
 &lt;path d="M10 13a5 5 0 0 0 7.54.54l3-3a5 5 0 0 0-7.07-7.07l-1.72 1.71" /&gt;
 &lt;path d="M14 11a5 5 0 0 0-7.54-.54l-3 3a5 5 0 0 0 7.07 7.07l1.71-1.71" /&gt;
&lt;/svg&gt;

 &lt;/a&gt;
 
&lt;/h3&gt;
&lt;p&gt;&lt;em&gt;Reinforcement Learning: An Introduction&lt;/em&gt; by Richard S. Sutton and Andrew G. Barto (MIT Press, 2nd ed., 2018) is the canonical textbook for the field. It defines the vocabulary, formalizes the core problems, and develops the algorithmic solutions that all subsequent RL work builds on. The second edition, available free online, extended coverage through deep reinforcement learning and policy gradient methods, bringing the book current with the AlphaGo-era explosion of the field.&lt;/p&gt;</description></item></channel></rss>