Full metadata record
DC FieldValueLanguage
dc.contributorDepartment of Computingen_US
dc.contributor.advisorLi, Bo (COMP)en_US
dc.creatorYuan, Mingqi-
dc.identifier.urihttps://theses.lib.polyu.edu.hk/handle/200/14658-
dc.languageEnglishen_US
dc.publisherHong Kong Polytechnic Universityen_US
dc.rightsAll rights reserveden_US
dc.titleIntrinsically-motivated exploration and task-aware data exploitation : a path to accelerated reinforcement learningen_US
dcterms.abstractLearning through interaction with the environment is a lifelong behavior of human beings, which has significantly inspired the advancement of reinforcement learning (RL). As a general and powerful paradigm, RL has been widely adopted across the technology revolution, from large language models (LLMs) to cutting-edge semiconductor design. Despite its success, RL remains fundamentally constrained by severe sample inefficiency, in which training often requires millions of interactions and may still yield suboptimal policies or fail entirely. This limitation leads to substantial computational overhead and significantly restricts the applicability of RL in domains where data collection is costly or impractical.en_US
dcterms.abstractAs RL increasingly involves diverse domains, the need for more data-efficient algorithms becomes ever more critical. This thesis is dedicated to accelerating RL training by systematically enhancing exploration and exploitation, the two fundamental mechanisms that govern learning efficiency. In Chapter 4, we conduct a comprehensive study on the intrinsically-motivated exploration mechanism. More specifically, we propose computation-efficient intrinsic reward formulations that provide sustained exploration incentives, alleviating the biased-objective issue arising from intrinsic reward shaping, and establish a unified benchmark to support reproducible research and principled comparison in this area. In Chapter 5, we focus on task-aware data exploitation and develop a series of acceleration strategies, including adaptive exploitation scheduling, ultra-lightweight hyperparameter optimization, and state representation learning tailored for RL. These techniques collectively enhance training stability, sample efficiency, and generalization while maintaining low computational overhead.en_US
dcterms.abstractBeyond algorithmic development, we demonstrate the practical impact of the proposed acceleration techniques on complex humanoid robot learning tasks. By applying the developed methods to whole-body control problems, we show remarkable improvements in learning speed, robustness, and real-world performance. In summary, this thesis makes significant contributions to the field of data-efficient RL, advancing the development of a more efficient and scalable RL paradigm to support future AI research across diverse domains.en_US
dcterms.extentxxii, 167 pages : color illustrationsen_US
dcterms.isPartOfPolyU Electronic Thesesen_US
dcterms.issued2026en_US
dcterms.educationalLevelPh.D.en_US
dcterms.educationalLevelAll Doctorateen_US
dcterms.accessRightsopen accessen_US

Files in This Item:
File Description SizeFormat 
9102.pdfFor All Users21.33 MBAdobe PDFView/Open


Copyright Undertaking

As a bona fide Library user, I declare that:

  1. I will abide by the rules and legal ordinances governing copyright regarding the use of the Database.
  2. I will use the Database for the purpose of my research or private study only and not for circulation or further reproduction or any other purpose.
  3. I agree to indemnify and hold the University harmless from and against any loss, damage, cost, liability or expenses arising from copyright infringement or unauthorized usage.

By downloading any item(s) listed above, you acknowledge that you have read and understood the copyright undertaking as stated above, and agree to be bound by all of its terms.

Show simple item record

Please use this identifier to cite or link to this item: https://theses.lib.polyu.edu.hk/handle/200/14658