Author: Feng, Yujie
Title: Towards adaptive and evolving large language models : continual learning and knowledge editing
Advisors: Wu, Xiao-ming (DSAI)
Degree: Ph.D.
Year: 2026
Department: Department of Data Science and Artificial Intelligence
Pages: xvi, 153 pages : color illustrations
Language: English
Abstract: Large language models (LLMs) have become central to a wide range of applications, but their ability to adapt to ever-evolving environments, including changing tasks, data, and user needs, remains a significant challenge. Knowledge adaptation in LLMs is crucial for maintaining their effectiveness in dynamic settings where continuous learning and knowledge updates are essential. Traditional methods, such as multi-task learning, attempt to integrate new knowledge with previously learned information by jointly training on both new and historical datasets. While these strategies can balance the integration of new and historical knowledge, they are computationally expensive, memory-intensive, and often inefficient.
This limitation constrains the adaptability of models in real-world, continually changing scenarios. In such settings, models must continuously acquire new knowledge while retaining previously learned information to maintain stability and performance. Moreover, it is crucial for models to update outdated knowledge to ensure consistency and reliability. To address these demands, this thesis examines strategies for optimizing knowledge adaptation in continual learning (CL) and model editing, with a focus on improving efficiency, reliability, and practical applicability.
Continual learning (CL) enables models to acquire new knowledge from a sequence of tasks while retaining previously learned information, addressing the challenge of catastrophic forgetting. We first propose Knowledge Identification and Fusion (KIF), which localizes the importance of model parameters for different tasks and consolidates both task-specific and task-shared knowledge by leveraging a model merging mechanism. We then introduce Recurrent KIF, which dynamically estimates parameter importance and iteratively integrates knowledge, overcoming the limitations of static importance estimation in earlier approaches. Finally, we propose Adaptive Iterative Model Merging (AIMMerging), which determines the optimal timing for model merging by leveraging learning and forgetting signals extracted from the training trajectory. Together, these progressive methods advance the continual learning field, providing scalable and effective solutions for continuous knowledge adaptation in large language models.
Model editing focuses on updating outdated or incorrect facts encoded in LLMs, ensuring that the new knowledge is incorporated without disrupting unrelated information. To address this, we propose a geometric knowledge editing method that distinguishes neurons associated with new knowledge updates from those linked to general knowledge perturbations by analyzing the geometric relationships of parameter changes induced by fine-tuning. By masking updates to general-knowledge-related neurons, we prevent negative impacts on generalization ability, while optimizing updates to new-knowledge-related neurons, thereby improving the precision, stability, and effectiveness of model editing.
To demonstrate the practical impact of our approaches, we evaluate them on dialogue state tracking (DST), a core task in task-oriented dialogue systems. The proposed continual learning techniques enhance the robustness and adaptability of LLM-driven DST models in cross-domain and evolving-task scenarios, highlighting the real-world value of knowledge adaptation.
In summary, this thesis develops and evaluates a set of methodologies for knowledge adaptation in large language models under practical constraints. By addressing key challenges such as catastrophic forgetting, efficient knowledge integration, and reliable model updates, the proposed approaches strengthen the robustness and applicability of LLMs in dynamic environments. These methods—spanning continual learning techniques that mitigate forgetting and editing strategies that ensure precise and consistent updates—offer targeted solutions to common limitations in real-world adaptation. Extensive experiments validate their effectiveness, showing consistent improvements in stability, generalization, and adaptability. Beyond empirical results, this thesis provides methodological insights that can guide future research and practical deployment of scalable and reliable LLM adaptation. Related contributions have been published or submitted to top NLP conferences.
Rights: All rights reserved
Access: open access

Files in This Item:
File Description SizeFormat 
9101.pdfFor All Users12.63 MBAdobe PDFView/Open


Copyright Undertaking

As a bona fide Library user, I declare that:

  1. I will abide by the rules and legal ordinances governing copyright regarding the use of the Database.
  2. I will use the Database for the purpose of my research or private study only and not for circulation or further reproduction or any other purpose.
  3. I agree to indemnify and hold the University harmless from and against any loss, damage, cost, liability or expenses arising from copyright infringement or unauthorized usage.

By downloading any item(s) listed above, you acknowledge that you have read and understood the copyright undertaking as stated above, and agree to be bound by all of its terms.

Show full item record

Please use this identifier to cite or link to this item: https://theses.lib.polyu.edu.hk/handle/200/14657