Full metadata record
| DC Field | Value | Language |
|---|---|---|
| dc.contributor | Department of Applied Mathematics | en_US |
| dc.contributor.advisor | Li, Xun (AMA) | en_US |
| dc.creator | Zhang, Haoran | - |
| dc.identifier.uri | https://theses.lib.polyu.edu.hk/handle/200/14649 | - |
| dc.language | English | en_US |
| dc.publisher | Hong Kong Polytechnic University | en_US |
| dc.rights | All rights reserved | en_US |
| dc.title | Reinforcement learning for linear quadratic control under multiplicative noise : policy methods, scalability, and entropy regularization | en_US |
| dcterms.abstract | The discrete time Linear Quadratic (LQ) control problem has been a cornerstone of control theory, and it provides an optimization framework for regulating linear dynamical systems, with quadratic cost functions used for minimization. With a closed-form solution from the algebraic Riccati equation (ARE), the LQ problem offers unique tractability in optimal control problems. However, in many practical scenarios, the true system parameters and noise characteristics are unknown, rendering model-based solutions infeasible. Consequently, reinforcement learning (RL) has become a powerful data-driven approach for policy optimization in the LQ problem under parameter uncertainty. Recent advances have highlighted its versatility with applications in the field of dynamic portfolio management, financial derivative pricing, and biomedical engineering. This thesis is organized into three sections. It focuses on key aspects of RL for the discrete time LQ problem and hence highlights their relevance to financial decision-making problems. In addition, the thesis contributes to expanding the single-agent LQ problem to multi-agent settings to address challenges related to scalability and decentralized control structures. Throughout, the emphasis is on advancing theoretical insights and offering practical tools for data-driven control in financial and engineering applications. Examples include the mean-variance (MV) problem and the optimal liquidation problem. Each section offers theoretical insights and practical contributions to enhance policy optimization techniques under diverse contexts. | en_US |
| dcterms.abstract | Part I studies finite horizon LQ control with mixed noise through a perturbance-wise viewpoint that unifies the classical model, constrained control case, and a distributionally robust extension calibrated from an unknown noise distribution framework. By augmenting the affine controller into a linear feedback form, I derive an implementable policy gradient method that accommodates a nonzero noise mean estimated from samples. I establish global decrease and convergence guarantees under constant step sizes chosen as explicit low-order functions of problem parameters. Numerical studies in MV portfolio allocation and dynamic benchmark tracking validate stable convergence and illustrate sensitivity trade-offs across horizon length, trading frictions, and estimation windows. | en_US |
| dcterms.abstract | Part II focuses on an iterative policy optimization (IPO) framework for multi-agent LQ control problems, where the system is subject to multiplicative noise and governed by a Shannon entropy-regularized cost function. It demonstrates that IPO achieves a super-linear convergence rate in the neighborhood of the optimal policy. The applications of MV optimization and optimal liquidation problems underscore the utility of IPO in financial decision-making contexts. This chapter explores the understudied scenario in which multiplicative noise simultaneously influences both control and state variables. It fills a gap in the extensive policy iteration literature by analyzing how noise impacts model efficiency in a multi-agent case. The findings establish a theoretical basis for enhancing the effectiveness of RL algorithms in the multi-agent LQ problem. | en_US |
| dcterms.abstract | Part III presents an innovative, data-driven policy iteration approach for LQ problem under Tsallis-entropy regularization. This study utilizes Tsallis entropy to balance exploration and sparsity in control laws, contrasting with the prevailing emphasis on Shannon entropy in the previous literature. The methodology utilizes Q-functions estimated by least-squares methods with instrumental variables, offering a robust solution for problems associated with state/control-multiplicative noise. This study tackles significant deficiencies in the integration of Tsallis entropy with data-driven methods, thereby opening new directions for enhancing RL-based policy in stochastic control systems. | en_US |
| dcterms.extent | xxiv, 160 pages : color illustrations | en_US |
| dcterms.isPartOf | PolyU Electronic Theses | en_US |
| dcterms.issued | 2026 | en_US |
| dcterms.educationalLevel | Ph.D. | en_US |
| dcterms.educationalLevel | All Doctorate | en_US |
| dcterms.accessRights | open access | en_US |
Copyright Undertaking
As a bona fide Library user, I declare that:
- I will abide by the rules and legal ordinances governing copyright regarding the use of the Database.
- I will use the Database for the purpose of my research or private study only and not for circulation or further reproduction or any other purpose.
- I agree to indemnify and hold the University harmless from and against any loss, damage, cost, liability or expenses arising from copyright infringement or unauthorized usage.
By downloading any item(s) listed above, you acknowledge that you have read and understood the copyright undertaking as stated above, and agree to be bound by all of its terms.
Please use this identifier to cite or link to this item:
https://theses.lib.polyu.edu.hk/handle/200/14649

