Full metadata record
| DC Field | Value | Language |
|---|---|---|
| dc.contributor | Department of Computing | en_US |
| dc.contributor.advisor | Li, Ping (COMP) | en_US |
| dc.contributor.advisor | Li, Jing Amelia (COMP) | en_US |
| dc.creator | Duan, Xin | - |
| dc.identifier.uri | https://theses.lib.polyu.edu.hk/handle/200/14402 | - |
| dc.language | English | en_US |
| dc.publisher | Hong Kong Polytechnic University | en_US |
| dc.rights | All rights reserved | en_US |
| dc.title | Real-world dynamic shadow perception and shadow-aware colorization | en_US |
| dcterms.abstract | Visual perception is a fundamental component of machine intelligence, with learning-based methods outperforming traditional handcrafted techniques in understanding complex visual data. Despite these advances, many vision systems struggle with subtle yet critical cues such as shadows, which are often dismissed as noise or artifacts. Shadows encode rich contextual information about illumination, object geometry, and spatial layout, playing a vital role in accurate scene interpretation. This complexity is further amplified in dynamic real-world scenes where shadows evolve temporally alongside moving objects and varying lighting conditions. | en_US |
| dcterms.abstract | This thesis addresses the core research question: Can machines learn to robustly interpret shadows in complex, real-world scenes, and can shadow perception enhance high-level vision tasks such as image colorization? To address this, three core challenges are explored. First, robust video shadow detection is tackled through the development of a large-scale dataset and a tailored neural architecture designed for complex, dynamic environments. Second, an unsupervised, CLIP-driven framework is proposed to enable scalable shadow detection without the need for manual annotations. Third, the thesis explores the integration of shadow understanding into downstream vision tasks—specifically image colorization—demonstrating how shadow-aware models can improve semantic consistency and visual realism. Together, these contributions aim to advance machine perception of shadows as a critical component in video and image understanding. | en_US |
| dcterms.abstract | To overcome these challenges, we first introduce the Complex Video Shadow Dataset (CVSD) in Chapter 3, a large-scale, diverse dataset comprising 196 videos and over 19,000 frames with detailed shadow annotations in complex scenarios. Leveraging CVSD, we propose a novel two-stage, fully supervised shadow detection framework incorporating spatial and temporal adaptation modules to effectively model dynamic shadow patterns also in Chapter 3. To reduce annotation dependency, we develop CDCNet, a CLIP-driven Divide-and-Conquer network enabling fully unsupervised video shadow detection by exploiting semantic priors and motion cues in chapter 4. Finally, we present the first shadow-aware image colorization network that incorporates a Shadow-Aware Block, explicitly integrating shadow features to improve color accuracy and semantic fidelity in shadowed regions in Chapter 5. | en_US |
| dcterms.abstract | Our contributions demonstrate that precise shadow perception not only enhances the robustness of shadow detection in complex, real-world environments but also significantly improves the performance of high-level vision tasks. By bridging the gap between low-level shadow understanding and semantic image colorization, this work highlights the critical role of shadow reasoning in advancing broader visual perception capabilities. | en_US |
| dcterms.abstract | The main research contents can be summarized as follows: | en_US |
| dcterms.abstract | 1. Complex Video Shadow Dataset and Detection Framework: To facilitate the exploration of video shadow detection in real-world scenarios, we introduce the Complex Video Shadow Dataset (CVSD)—a large-scale dataset comprising 196 video clips across 149 scene categories, capturing diverse and challenging shadow patterns. The dataset includes 19,757 frames with complex shadow patterns, making it the first dataset for complex real-world video shadow detection. Building upon CVSD, we propose a two-stage training paradigm and a novel neural architecture specifically designed for dynamic shadow scenarios. By framing shadow detection as a conditioned feature adaptation problem, we introduce temporal- and spatial-adaptation blocks to effectively capture spatiotemporal context, significantly enhancing detection accuracy in complex, real-world environments. | en_US |
| dcterms.abstract | 2. Unsupervised Video Shadow Detection with CDCNet: We propose CDCNet, a CLIP-driven Divide-and-Conquer Network designed for unsupervised video shadow detection. The architecture incorporates two key modules—Spatial Divide and Temporal Conquer—which collaboratively capture spatial details and temporal coherence without requiring annotated supervision. By eliminating the need for pixel-level labels, CDCNet addresses the high cost and effort associated with manual dataset annotation. Extensive experiments demonstrate that CDCNet exhibits strong generalizability on unseen datasets, highlighting its effectiveness and scalability. | en_US |
| dcterms.abstract | 3. Shadow-Aware Image Colorization: We propose the first shadow-aware image colorization method, based on an extensive review of two decades of colorization techniques. Our approach employs a dual-branch network enhanced by a novel Shadow-Aware Block, which explicitly incorporates shadow-specific features into the colorization process. This enables more accurate and semantically consistent colorization in scenes with shadow conditions by effectively distinguishing between shadowed and non-shadowed regions. | en_US |
| dcterms.abstract | The results of our research have been published in or to be submitted to several leading conferences and journals in computer vision and computer graphics, including The Visual Computer (Journal), ECCV (Conference) and TVCG (Journal). | en_US |
| dcterms.extent | xix, 162 pages : color illustrations | en_US |
| dcterms.isPartOf | PolyU Electronic Theses | en_US |
| dcterms.issued | 2026 | en_US |
| dcterms.educationalLevel | Ph.D. | en_US |
| dcterms.educationalLevel | All Doctorate | en_US |
| dcterms.accessRights | open access | en_US |
Copyright Undertaking
As a bona fide Library user, I declare that:
- I will abide by the rules and legal ordinances governing copyright regarding the use of the Database.
- I will use the Database for the purpose of my research or private study only and not for circulation or further reproduction or any other purpose.
- I agree to indemnify and hold the University harmless from and against any loss, damage, cost, liability or expenses arising from copyright infringement or unauthorized usage.
By downloading any item(s) listed above, you acknowledge that you have read and understood the copyright undertaking as stated above, and agree to be bound by all of its terms.
Please use this identifier to cite or link to this item:
https://theses.lib.polyu.edu.hk/handle/200/14402

