Master's Thesis
Single-partition adaptive Q-learning: algorithm and applications
— 2020
Key information
Authors:
Supervisors:
Published in
July 23, 2020
Abstract
Reinforcement learning (RL) is an area within machine learning that studies how agents can learn to perform their tasks without being explicitly told how to do so. An important concept in RL is sample efficiency: an algorithm is sample-efficient if it requires a low amount of samples to learn its task. Until recently, it was thought that model-based RL algorithms were more sample-efficient than model-free ones. This changed with recent developments in provably efficient model-free algorithms. One of the latest algorithms developed is adaptive Q-learning (AQL), an efficient model-free algorithm that handles continuous state and action spaces by adaptively partitioning them in a data-driven manner. By design, AQL learns time-variant policies. However, many problems (such as control of time-invariant systems) can be solved satisfactorily using time-invariant policies. This thesis introduces single-partition adaptive Q-learning (SPAQL), an improved version of AQL designed to learn time-invariant policies. SPAQL is evaluated empirically on four different problems, out of which two are control problems. SPAQL agents perform better than AQL ones, while at the same time learning simpler policies. For the control problems, SPAQL with terminal state (SPAQL-TS) is introduced, and, along with SPAQL, is compared to trust region policy optimization (TRPO), an RL algorithm known to perform well in control problems. In one of the control problems (CartPole), SPAQL and SPAQL-TS display a higher sample-efficiency than TRPO.
Publication details
Authors in the community:
Supervisors of this institution:
Fields of Science and Technology (FOS)
mechanical-engineering - Mechanical engineering
Publication language (ISO code)
por - Portuguese
Rights type:
Embargo lifted
Date available:
May 24, 2021
Institution name
Instituto Superior Técnico