欢迎访问中国科学院大学学报,今天是
电子信息与计算机科学

不确定度增强下基于模型的探索-学习策略联合优化

  • 肖士湘 ,
  • 黄文振 ,
  • 焦建彬
展开
  • 1.中国科学院大学电子电气与通信工程学院,北京 100049
    2.清华大学信息科学技术学院,北京 100084

收稿日期: 2023-12-04

  修回日期: 2024-09-05

  网络出版日期: 2024-12-23

基金资助

中国科学院先导科技专项(XDA27010300)

Model-based exploration-learning joint optimization via uncertainty augmentation

  • Shixiang XIAO ,
  • Wenzhen HUANG ,
  • Jianbin JIAO
Expand
  • 1.School of Electronic,Electrical and Communication Engineering,University of Chinese Academy of Sciences,Beijing 100049,China
    2.School of Information Science and Technology,Tsinghua University,Beijing 100084,China

Received date: 2023-12-04

  Revised date: 2024-09-05

  Online published: 2024-12-23

摘要

在现有基于模型的强化学习方法中,智能体采用同一策略进行与真实环境和环境模型的交互,难以兼顾探索环境的效率和策略更新的稳定性。为解决此问题,提出一种不确定度增强下基于模型的探索-学习策略联合优化方法。该方法联合优化2个策略,分别是用于与真实环境交互的探索策略和与环境模型交互的学习策略。在优化探索策略的过程中,引入基于模型不确定度的内在激励以提升探索真实环境的效率,同时在优化学习策略时将模型不确定度作为约束条件,保证学习策略优化的稳定性。多个连续控制任务的实验结果表明,该方法相比于现有的代表性方法在渐进性能和样本效率上有明显的优势。

本文引用格式

肖士湘 , 黄文振 , 焦建彬 . 不确定度增强下基于模型的探索-学习策略联合优化[J]. 中国科学院大学学报, 2026 , 43(5) : 667 -676 . DOI: 10.7523/j.ucas.2024.072

Abstract

In existing model-based reinforcement learning methods, a single policy is adopted to interact with the real environment and the environment model, which makes it hard for the agent to balance the efficiency of exploring the environment and the stability of policy updating. To address this issue, this paper proposes a model-based explorer-learner joint optimization via uncertainty augmentation method (MELO-UA). MELO-UA simultaneously optimizes a pair of policies, namely the explorer policy for interacting with the real environment, and the learner policy for interacting with the environment model. During the optimization of the explorer, an implicit bonus based on model uncertainty is introduced to enhance the efficiency of exploring the real environment. At the same time, during the optimization of the learner, the model uncertainty is used as a constraint to ensure the stability of the policy optimization. Experimental results on multiple continuous control tasks show that the proposed method has significant advantages in both asymptotic performance and sample efficiency compared to state-of-the-art methods.

参考文献

[1] Schrittwieser J, Antonoglou I, Hubert T, et al. Mastering Atari, Go, chess and shogi by planning with a learned model[J]. Nature2020588(7839): 604-609. DOI: 10.1038/s41586-020-03051-4 .
[2] Silver D, Schrittwieser J, Simonyan K, et al. Mastering the game of Go without human knowledge[J]. Nature2017550(7676): 354-359. DOI: 10.1038/nature24270 .
[3] Schulman J, Wolski F, Dhariwal P, et al. Proximal policy optimization algorithms[EB/OL]. arXiv 2017:1707.06347. (2017-07-20)[2024-07-23]. .
[4] Mnih V, Kavukcuoglu K, Silver D, et al. Human-level control through deep reinforcement learning[J]. Nature2015518(7540): 529-533. DOI: 10.1038/nature14236 .
[5] Janner M, Fu J, Zhang M, et al. When to trust your model: model-based policy optimization[EB/OL]. arXiv 2019:1906.08253. (2019-06-19)[2024-07-23]. .
[6] Nagabandi A, Kahn G, Fearing R S, et al. Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuning[C]//2018 IEEE International Conference on Robotics and Automation (ICRA). Brisbane, QLD, Australia. IEEE, 2018: 7559-7566. DOI: 10.1109/ICRA.2018.8463189 .
[7] Higuera J C G, Meger D, Dudek G. Synthesizing neural network controllers with probabilistic model-based reinforcement learning[C]//2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, 2018: 2538-2544. DOI: 10.1109/IROS.2018.8594018 .
[8] Chua K, Calandra R, McAllister R, et al. Deep reinforcement learning in a handful of trials using probabilistic dynamics models[C]//Proceedings of the 32nd International Conference on Neural Information Processing Systems. December 3 - 8, 2018, Montréal, Canada. ACM, 2018: 4759-4770. DOI: 10.5555/3327345.3327385 .
[9] Feinberg V, Wan A, Stoica I, et al. Model-based value estimation for efficient model-free reinforcement learning[EB/OL]. arXiv 2018:1803.00101. (2018-02-28)[2024-07-23]. .
[10] Sutton R S. Integrated architectures for learning, planning, and reacting based on approximating dynamic programming[M]//Machine Learning Proceedings 1990. Amsterdam: Elsevier, 1990: 216-224. DOI: 10.1016/b978-1-55860-141-3.50030-4 .
[11] Luo Y, Xu H, Li Y, et al. Algorithmic Framework for model-based deep reinforcement learning with theoretical guarantees[C/OL]//International Conference on Learning Representations. 2019. (2018-12-21)[2024-07-23]. .
[12] Buckman J, Hafner D, Tucker G, et al. Sample-efficient reinforcement learning with stochastic ensemble value expansion[C]//Proceedings of the 32nd International Conference on Neural Information Processing Systems. December 3 - 8, 2018, Montréal, Canada. ACM, 2018: 8234-8244. DOI: 10.5555/3327757.3327916 .
[13] Deisenroth M P, Rasmussen C E. PILCO: a model-based and data-efficient approach to policy search[C]//Proceedings of the 28th International Conference on International Conference on Machine Learning. 28 June 2011, Bellevue, Washington, USA. ACM, 2011: 465-472. DOI: 10.5555/3104482.3104541 .
[14] Heess N, Wayne G, Silver D, et al. Learning continuous control policies by stochastic value gradients[EB/OL]. arXiv 2015:1510.09142. (2015-10-30)[2024-07-23]. .
[15] Hafner D, Lillicrap T, Ba J, et al. Dream to control: Learning behaviors by latent imagination[EB/OL]. arXiv 2019:1912.01603. (2019-12-03)[2024-07-24]. .
[16] Hafner D, Pasukonis J, Ba J, et al. Mastering diverse domains through world models[EB/OL]. arXiv 2023:2301.04104. (2023-01-10)[2024-07-23]. .
[17] Hafner D, Lillicrap T, Fischer I, et al. Learning latent dynamics for planning from pixels[EB/OL]. arXiv 2018:1811.04551. (2018-11-12)[2024-07-23]. .
[18] Kurutach T, Clavera I, Duan Y, et al. Model-ensemble trust-region policy optimization[EB/OL]. arXiv 2018:1802.10592. (2018-02-28)[2024-07-23]. .
[19] He W, Jiang Z. A Comprehensive Survey on Uncertainty Quantification for Deep Learning[EB/OL]. arXiv 2023:2302.13425. (2023-02-26)[2024-07-23]. .
[20] Psaros A F, Meng X H, Zou Z R, et al. Uncertainty quantification in scientific machine learning: methods, metrics, and comparisons[J]. Journal of Computational Physics2023477: 111902. DOI: 10.1016/j.jcp.2022.111902 .
[21] Lakshminarayanan B, Pritzel A, Blundell C. Simple and scalable predictive uncertainty estimation using deep Ensembles[EB/OL]. arXiv 2016:1612.01474. (2016-12-05)[2024-07-23]. .
[22] Wang Z, Jusup M, Shi L, et al. Exploiting a cognitive bias promotes cooperation in social dilemma experiments[J]. Nature Communications20189(1): 2954. DOI: 10.1038/s41467-018-05259-5 .
[23] Lu C, Ball P, Parker-Holder J, et al. Revisiting design choices in offline model-based reinforcement learning[EB/OL]. arXiv 2021:2110.04135. (2021-10-08)[2024-07-23]. .
[24] Pan F, He J, Tu D, et al. Trust the model when it is confident: masked model-based actor-critic[EB/OL]. arXiv 2020:2020.04893. (2020-10-10)[2024-07-23]. .
[25] Yu T H, Thomas G, Yu L T, et al. MOPO: model-based offline policy optimization[EB/OL]. arXiv 2020:2005.13239. (2020-05-27)[2024-07-23]. .
[26] Bechtle S, Lin Y X, Rai A, et al. Curious iLQR: resolving uncertainty in model-based RL[EB/OL]. arXiv 2019:1904.06786. (2019-04-15)[2024-07-23]. .
[27] Hao J Y, Yang T P, Tang H Y, et al. Exploration in deep reinforcement learning: from single-agent to multiagent domain[J]. IEEE Transactions on Neural Networks and Learning Systems2023. DOI: 10.1109/TNNLS.2023.3236361 .
[28] Abbasi-Yadkori Y, P??l D, Szepesvári C. Improved algorithms for linear stochastic bandits[C]//Proceedings of the 24th International Conference on Neural Information Processing Systems. December 12 - 15, 2011, Granada, Spain. ACM, 2011: 2312-2320. DOI: 10.5555/2986459.2986717&widgetKey=ux3-publicationContent-widget_f69d88a8-b404 -4aae-83 a9-9 acea 4426d78_1498 _en.
[29] Bai C J, Wang L X, Han L, et al. Principled exploration via optimistic bootstrapping and backward induction[EB/OL]. arXiv 2021:2105.06022. (2021-05-13)[2024-07-23]. .
[30] Zheng Y, Liu Y, Xie X F, et al. Automatic web testing using curiosity-driven reinforcement learning[C]//2021 IEEE/ACM 43rd International Conference on Software Engineering (ICSE). Madrid, ES. IEEE, 2021: 423-435. DOI: 10.1109/ICSE43902.2021.00048 .
[31] Ciosek K, Vuong Q, Loftin R, et al. Better exploration with optimistic actor-critic[EB/OL]. arXiv 2019:1910.12807. (2019-10-28)[2024-07-23]. .
[32] Haarnoja T, Zhou A, Abbeel P, et al. Soft actor-critic: off-policy maximum entropy deep reinforcement learning with a stochastic actor[EB/OL]. arXiv 2018:1801.01290. (2018-01-04)[2024-07-23]. .
[33] Burda Y, Edwards H, Storkey A, et al. Exploration by random network distillation[EB/OL]. arXiv 2018:1810.12894. (2018-10-30)[2024-07-23]. .
[34] Liu J Y, Wang Z, Zheng Y, et al. OVD-explorer: optimism should not be the sole pursuit of exploration in noisy environments[J]. Proceedings of the AAAI Conference on Artificial Intelligence202438(12): 13954-13962. DOI: 10.1609/aaai.v38i12.29303 .
[35] Bai C J, Wang L X, Han L, et al. Dynamic bottleneck for robust self-supervised exploration[EB/OL]. arXiv 2021:2110.10735. (2021-10-20)[2024-07-23]. .
[36] Oudeyer P Y, Kaplan F, Hafner V V. Intrinsic motivation systems for autonomous mental development[J]. IEEE Transactions on Evolutionary Computation200711(2): 265-286. DOI: 10.1109/TEVC.2006.890271 .
[37] Pathak D, Agrawal P, Efros A A, et al. Curiosity-driven exploration by self-supervised prediction[C/OL]//2017 IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW). 2017: 488-489. DOI:10.1109/CVPRW.2017.70 .
[38] Pathak D, Gandhi D, Gupta A. Self-supervised exploration via disagreement[EB/OL]. arXiv 2019:1906.04161. (2019-01-10)[2024-07-23]. .
[39] Yuan Y F, Hao J Y, Ni F, et al. EUCLID: towards efficient unsupervised reinforcement learning with multi-choice dynamics model[EB/OL]. arXiv 2022:2210.00498. (2022-10-02)[2024-07-23]. .
[40] Haarnoja T, Tang H R, Abbeel P, et al. Reinforcement learning with deep energy-based policies[EB/OL]. arXiv 2017:1702.08165. (2017-02-27)[2024-07-23]. .
[41] Pineda L, Amos B, Zhang A, et al. MBRL-lib: a modular library for model-based reinforcement learning[EB/OL]. arXiv 2021:2104.10159. (2021-04-20)[2024-07-23]. .
文章导航

/