arXiv:2007.00722 Abstract | arXiv Analytics

arXiv:2007.00722 [cs.LG]Abstract References Reviews Resources

Sequential Transfer in Reinforcement Learning with a Generative Model

Andrea Tirinzoni, Riccardo Poiani, Marcello Restelli

Published 2020-07-01Version 1

We are interested in how to design reinforcement learning agents that provably reduce the sample complexity for learning new tasks by transferring knowledge from previously-solved ones. The availability of solutions to related problems poses a fundamental trade-off: whether to seek policies that are expected to achieve high (yet sub-optimal) performance in the new task immediately or whether to seek information to quickly identify an optimal solution, potentially at the cost of poor initial behavior. In this work, we focus on the second objective when the agent has access to a generative model of state-action pairs. First, given a set of solved tasks containing an approximation of the target one, we design an algorithm that quickly identifies an accurate solution by seeking the state-action pairs that are most informative for this purpose. We derive PAC bounds on its sample complexity which clearly demonstrate the benefits of using this kind of prior knowledge. Then, we show how to learn these approximate tasks sequentially by reducing our transfer setting to a hidden Markov model and employing spectral methods to recover its parameters. Finally, we empirically verify our theoretical findings in simple simulated domains.

Comments: ICML 2020

Categories: cs.LG, stat.ML

Keywords: generative model, sequential transfer, sample complexity, state-action pairs, poor initial behavior

Related articles: Most relevant | Search more

arXiv:1206.6461 [cs.LG] (Published 2012-06-27)

On the Sample Complexity of Reinforcement Learning with a Generative Model

Mohammad Gheshlaghi Azar, Remi Munos, Bert Kappen

arXiv:2112.01506 [cs.LG] (Published 2021-12-02, updated 2022-05-14)

Sample Complexity of Robust Reinforcement Learning with a Generative Model

Kishan Panaganti, Dileep Kalathil

arXiv:2003.11399 [cs.LG] (Published 2020-03-25)

Discriminative Viewer Identification using Generative Models of Eye Gaze