Full metadata record
DC FieldValueLanguage
dc.contributorDepartment of Computingen_US
dc.contributor.advisorZhang, Chen (COMP)en_US
dc.contributor.advisorLi, Qing (COMP)en_US
dc.contributor.advisorLi, Gang (EEE)en_US
dc.creatorZhang, Ting-
dc.identifier.urihttps://theses.lib.polyu.edu.hk/handle/200/14475-
dc.languageEnglishen_US
dc.publisherHong Kong Polytechnic Universityen_US
dc.rightsAll rights reserveden_US
dc.titleTowards accurate molecular property prediction : innovations in data cleansing, molecular descriptions, and model optimizationen_US
dcterms.abstractPredicting molecular properties with high accuracy is fundamental to accelerating progress in the discovery and optimization of functional materials. In particular, organic solar cells (OSCs) represent a rapidly advancing class of technologies for sustainable energy, characterized by low fabrication cost, mechanical adaptability, and suitability for scalable roll-to-roll manufacturing. Despite significant advancements, traditional experimental methodologies for identifying and optimizing OSC materials remain labor-intensive and time-consuming. Consequently, computational methods capable of accurately and efficiently predicting molecular properties have become increasingly essential. However, the effectiveness of these computational methods heavily depends on three interrelated aspects: the quality of the input data, the molecular descriptors utilized, and the capability of machine learning models to capture intricate molecular interactions. Recognizing this, the present thesis systematically addresses these three critical challenges through innovative strategies and novel methodologies aimed at improving predictive accuracy and efficiency in molecular property predictions, specifically targeting OSC materials.en_US
dcterms.abstractChapter 1 introduces the background, motivations, and objectives of this thesis, highlighting the significant potential yet persistent challenges in accurately predicting molecular and device-level properties of OSCs. It highlights the constraints of traditional computational approaches, inconsistencies in experimental data, and the limitations of current machine learning methods in effectively capturing complex molecular features and donor-acceptor interactions. To address these challenges, the thesis introduces advanced AI-driven frameworks that combine rigorous data preprocessing, novel molecular descriptors, and hierarchical transformer-based models. These innovations aim to enhance predictive accuracy, reduce dependence on extensive experimental validation, and speed up progress toward next-generation OSC materials with enhanced performance.en_US
dcterms.abstractChapter 2 addresses the foundational issue of input data quality, the thesis introduces a robust data cleansing framework designed to enhance the reliability of datasets used in computational modeling. In OSCs, non-fullerene acceptors (NFAs) have recently emerged as key materials for improving device efficiency. Accurate determination of their electronic properties, such as energy levels, is crucial but remains challenging, as experimental data frequently contain inconsistencies. Traditional computational techniques, such as density functional theory (DFT), commonly exhibit prediction errors ranging from 0.2 to 0.5 electron volts (eV). To overcome these limitations, this thesis proposes a systematic data cleansing methodology. This approach involves rigorous statistical cross-validation and the exclusion of anomalous experimental entries identified based on expert knowledge. Although the reliance on expert knowledge introduces potential subjectivity and limits scalability, this process significantly reduces input errors, thereby enhancing the reliability of subsequent computational predictions. Empirical results demonstrate a substantial improvement in predictive accuracy, reducing the error margin to approximately 0.06 eV, a considerable advancement over conventional computational methods. This contribution establishes a solid foundation for accurate downstream computational modeling.en_US
dcterms.abstractChapter 3 develops a sophisticated molecular representation method, which constitute the core innovation of this chapter. Effective molecular descriptors are essential for accurately capturing the intrinsic properties of molecules and predicting their behavior in complex chemical environments. Existing molecular descriptors predominantly emphasize atom-level features while often neglecting higher-order structural information, such as molecular subunits or chemical motifs, that critically influence molecular properties. Although some advanced representation methods have incorporated these subunit-level characteristics, they often require extensive, labor-intensive manual labeling and feature engineering. To address these drawbacks, the thesis introduces a novel, efficient, and fully automated molecular descriptor, Ring2Vec. Inspired by advancements in natural language processing (specifically the Word2Vec mechanism), Ring2Vec conceptualizes molecular rings as the basic linguistic units, comparable to words in natural language. Through unsupervised learning, Ring2Vec is capable of deriving information from molecules across atom-level and subunit-level hierarchies, constructing a comprehensive yet computationally efficient representation of molecules. Experiments validating Ring2Vec's performance reveal remarkable accuracy, notably achieving minimal energy-level prediction errors, significantly outperforming conventional approaches. Moreover, the generalizability of Ring2Vec suggests its potential utility beyond OSC applications, extending to broader domains within material science and molecular property prediction tasks.en_US
dcterms.abstractChapter 4 optimizes machine learning models specifically tailored to molecular property prediction tasks within OSCs. While graph-based representation learning methods have gained prominence in recent years, their application in the OSC domain has been under-explored, leading to suboptimal predictive performances. OSC molecules possess distinctive structural complexities, particularly intricate ring systems, donor-acceptor interactions, and hierarchical molecular features, which conventional graph neural network (GNN) methods frequently fail to fully capture. Recognizing these limitations, the thesis develops innovative machine learning frameworks specifically designed to address the unique structural characteristics of OSC molecules. In particular, this thesis introduces RingFormer, a hierarchical transformer model on graphs, tailored to combine atomic-scale and ring-based structural information, thereby capturing the multi-scale features essential for accurate property prediction in OSCs. By fusing local message-passing modules with long-range attention mechanisms, RingFormer is capable of capturing the hierarchical complexity of molecular architectures. Comprehensive empirical evaluations conducted on multiple curated OSC datasets demonstrate RingFormer's consistent superiority over state-of-the-art methods. Remarkably, RingFormer achieves a relative performance improvement exceeding 20% compared to competing approaches on benchmark datasets, thereby solidifying its effectiveness in accurately modeling complex molecular structures inherent in OSC materials.en_US
dcterms.abstractChapter 5 Building further on the concept of hierarchical modeling, the thesis culminates in the proposal of a sophisticated Hierarchical Graph Transformer with Multi-level Interactions (HGT-MI) framework, explicitly tailored to model intermolecular donor-acceptor interactions, which play a fundamental role in determining OSC device performance. Traditional computational models often overlook donor-acceptor interactions or account for only one component, leading to reduced predictive accuracy and reliability. In contrast, the HGT-MI model innovatively encodes donor-acceptor interactions through multi-scale, hierarchical representations at the atom-level, motif-level, and molecular-level. Specifically, HGT-MI leverages bi-directional multi-head cross-attention to systematically integrate donor and acceptor molecular embeddings, enabling rich and comprehensive information exchange between the two molecules. Extensive validation experiments demonstrate the substantial advantages of this approach, significantly outperforming existing models in predictive accuracy for key OSC metrics, notably power conversion efficiency (PCE). Importantly, the versatility and robustness of the HGT-MI framework highlight its broader applicability beyond OSCs. Thanks to its generalizable architecture, the framework can be applied beyond OSCs to domains like perovskite solar technologies, advanced battery chemistries, and organic light-emitting devices, underscoring its potential for broad material innovation.en_US
dcterms.abstractChapter 6 presents a comprehensive discussion and conclusion, synthesizing the key findings and contributions of this thesis. It traces the evolution of molecular property prediction from its roots in drug discovery to its emerging role in materials science, emphasizing the unique challenges of OSCs. The chapter integrates advances across four synergistic studies, data cleansing for noise-robust supervision, automated ring/subunit representations with Ring2Vec, hierarchical learning with RingFormer, and multi-level donor-acceptor interaction modeling with HGT-MI into a cohesive end-to-end framework for molecular and device-level prediction. Strengths and limitations are critically assessed, highlighting trade-offs between predictive accuracy, interpretability, and scalability, while also identifying open challenges such as data scarcity, morphological dynamics, and computational costs. Practical implications underscore the framework's potential to reduce experimental dependence, accelerate NFA and donor screening, and broaden the application of AI-driven methodologies across materials domains beyond OSCs. Finally, future directions are charted, including the integration of physics-informed constraints, explicit 3D conformational modeling, and collaborative data-sharing strategies, to further enhance the scalability, robustness, and translational impact of molecular property prediction frameworks.en_US
dcterms.extentxvi, 159 pages : color illustrationsen_US
dcterms.isPartOfPolyU Electronic Thesesen_US
dcterms.issued2026en_US
dcterms.educationalLevelPh.D.en_US
dcterms.educationalLevelAll Doctorateen_US
dcterms.accessRightsopen accessen_US

Files in This Item:
File Description SizeFormat 
8898.pdfFor All Users16.4 MBAdobe PDFView/Open


Copyright Undertaking

As a bona fide Library user, I declare that:

  1. I will abide by the rules and legal ordinances governing copyright regarding the use of the Database.
  2. I will use the Database for the purpose of my research or private study only and not for circulation or further reproduction or any other purpose.
  3. I agree to indemnify and hold the University harmless from and against any loss, damage, cost, liability or expenses arising from copyright infringement or unauthorized usage.

By downloading any item(s) listed above, you acknowledge that you have read and understood the copyright undertaking as stated above, and agree to be bound by all of its terms.

Show simple item record

Please use this identifier to cite or link to this item: https://theses.lib.polyu.edu.hk/handle/200/14475