Dissertação de Mestrado
Single-partition adaptive Q-learning: algorithm and applications
— 2020
Informações chave
Autores:
Orientadores:
Publicado em
July 23, 2020
Resumo
Reinforcement learning (RL) is an area within machine learning that studies how agents can learn to perform their tasks without being explicitly told how to do so. An important concept in RL is sample efficiency: an algorithm is sample-efficient if it requires a low amount of samples to learn its task. Until recently, it was thought that model-based RL algorithms were more sample-efficient than model-free ones. This changed with recent developments in provably efficient model-free algorithms. One of the latest algorithms developed is adaptive Q-learning (AQL), an efficient model-free algorithm that handles continuous state and action spaces by adaptively partitioning them in a data-driven manner. By design, AQL learns time-variant policies. However, many problems (such as control of time-invariant systems) can be solved satisfactorily using time-invariant policies. This thesis introduces single-partition adaptive Q-learning (SPAQL), an improved version of AQL designed to learn time-invariant policies. SPAQL is evaluated empirically on four different problems, out of which two are control problems. SPAQL agents perform better than AQL ones, while at the same time learning simpler policies. For the control problems, SPAQL with terminal state (SPAQL-TS) is introduced, and, along with SPAQL, is compared to trust region policy optimization (TRPO), an RL algorithm known to perform well in control problems. In one of the control problems (CartPole), SPAQL and SPAQL-TS display a higher sample-efficiency than TRPO.
Detalhes da publicação
Autores da comunidade :
Orientadores desta instituição:
Domínio Científico (FOS)
mechanical-engineering - Engenharia Mecânica
Idioma da publicação (código ISO)
por - Português
Acesso à publicação:
Embargo levantado
Data do fim do embargo:
May 24, 2021
Nome da instituição
Instituto Superior Técnico