Author: Yuan, Mingqi
Title: Intrinsically-motivated exploration and task-aware data exploitation : a path to accelerated reinforcement learning
Advisors: Li, Bo (COMP)
Degree: Ph.D.
Year: 2026
Department: Department of Computing
Pages: xxii, 167 pages : color illustrations
Language: English
Abstract: Learning through interaction with the environment is a lifelong behavior of human beings, which has significantly inspired the advancement of reinforcement learning (RL). As a general and powerful paradigm, RL has been widely adopted across the technology revolution, from large language models (LLMs) to cutting-edge semiconductor design. Despite its success, RL remains fundamentally constrained by severe sample inefficiency, in which training often requires millions of interactions and may still yield suboptimal policies or fail entirely. This limitation leads to substantial computational overhead and significantly restricts the applicability of RL in domains where data collection is costly or impractical.
As RL increasingly involves diverse domains, the need for more data-efficient algorithms becomes ever more critical. This thesis is dedicated to accelerating RL training by systematically enhancing exploration and exploitation, the two fundamental mechanisms that govern learning efficiency. In Chapter 4, we conduct a comprehensive study on the intrinsically-motivated exploration mechanism. More specifically, we propose computation-efficient intrinsic reward formulations that provide sustained exploration incentives, alleviating the biased-objective issue arising from intrinsic reward shaping, and establish a unified benchmark to support reproducible research and principled comparison in this area. In Chapter 5, we focus on task-aware data exploitation and develop a series of acceleration strategies, including adaptive exploitation scheduling, ultra-lightweight hyperparameter optimization, and state representation learning tailored for RL. These techniques collectively enhance training stability, sample efficiency, and generalization while maintaining low computational overhead.
Beyond algorithmic development, we demonstrate the practical impact of the proposed acceleration techniques on complex humanoid robot learning tasks. By applying the developed methods to whole-body control problems, we show remarkable improvements in learning speed, robustness, and real-world performance. In summary, this thesis makes significant contributions to the field of data-efficient RL, advancing the development of a more efficient and scalable RL paradigm to support future AI research across diverse domains.
Rights: All rights reserved
Access: open access

Files in This Item:
File Description SizeFormat 
9102.pdfFor All Users21.33 MBAdobe PDFView/Open


Copyright Undertaking

As a bona fide Library user, I declare that:

  1. I will abide by the rules and legal ordinances governing copyright regarding the use of the Database.
  2. I will use the Database for the purpose of my research or private study only and not for circulation or further reproduction or any other purpose.
  3. I agree to indemnify and hold the University harmless from and against any loss, damage, cost, liability or expenses arising from copyright infringement or unauthorized usage.

By downloading any item(s) listed above, you acknowledge that you have read and understood the copyright undertaking as stated above, and agree to be bound by all of its terms.

Show full item record

Please use this identifier to cite or link to this item: https://theses.lib.polyu.edu.hk/handle/200/14658