arXiv:2210.08642 Abstract | arXiv Analytics

arXiv:2210.08642 [cs.LG]Abstract References Reviews Resources

Data-Efficient Pipeline for Offline Reinforcement Learning with Limited Data

Allen Nie, Yannis Flet-Berliac, Deon R. Jordan, William Steenbergen, Emma Brunskill

Published 2022-10-16Version 1

Offline reinforcement learning (RL) can be used to improve future performance by leveraging historical data. There exist many different algorithms for offline RL, and it is well recognized that these algorithms, and their hyperparameter settings, can lead to decision policies with substantially differing performance. This prompts the need for pipelines that allow practitioners to systematically perform algorithm-hyperparameter selection for their setting. Critically, in most real-world settings, this pipeline must only involve the use of historical data. Inspired by statistical model selection methods for supervised learning, we introduce a task- and method-agnostic pipeline for automatically training, comparing, selecting, and deploying the best policy when the provided dataset is limited in size. In particular, our work highlights the importance of performing multiple data splits to produce more reliable algorithm-hyperparameter selection. While this is a common approach in supervised learning, to our knowledge, this has not been discussed in detail in the offline RL setting. We show it can have substantial impacts when the dataset is small. Compared to alternate approaches, our proposed pipeline outputs higher-performing deployed policies from a broad range of offline policy learning algorithms and across various simulation domains in healthcare, education, and robotics. This work contributes toward the development of a general-purpose meta-algorithm for automatic algorithm-hyperparameter selection for offline RL.

Comments: 32 pages. To be published at NeurIPS 2022. Presented at RLDM 2022

Categories: cs.LG, cs.AI, cs.RO, stat.ML

Keywords: offline reinforcement learning, data-efficient pipeline, limited data, algorithm-hyperparameter selection, offline rl

Related articles: Most relevant | Search more

arXiv:2012.11547 [cs.LG] (Published 2020-12-21)

Offline Reinforcement Learning from Images with Latent Space Models

Rafael Rafailov, Tianhe Yu, Aravind Rajeswaran, Chelsea Finn

arXiv:2210.08323 [cs.LG] (Published 2022-10-15)

A Policy-Guided Imitation Approach for Offline Reinforcement Learning

Haoran Xu, Li Jiang, Jianxiong Li, Xianyuan Zhan

arXiv:2111.10919 [cs.LG] (Published 2021-11-21, updated 2022-08-30)

Offline Reinforcement Learning: Fundamental Barriers for Value Function Approximation