Full metadata record
DC FieldValueLanguage
dc.contributorDepartment of Computingen_US
dc.contributor.advisorLi, Ping (COMP)en_US
dc.contributor.advisorLi, Jing Amelia (COMP)en_US
dc.creatorDuan, Xin-
dc.identifier.urihttps://theses.lib.polyu.edu.hk/handle/200/14402-
dc.languageEnglishen_US
dc.publisherHong Kong Polytechnic Universityen_US
dc.rightsAll rights reserveden_US
dc.titleReal-world dynamic shadow perception and shadow-aware colorizationen_US
dcterms.abstractVisual perception is a fundamental component of machine intelligence, with learning-based methods outperforming traditional handcrafted techniques in understanding complex visual data. Despite these advances, many vision systems struggle with subtle yet critical cues such as shadows, which are often dismissed as noise or artifacts. Shadows encode rich contextual information about illumination, object geometry, and spatial layout, playing a vital role in accurate scene interpretation. This complexity is further amplified in dynamic real-world scenes where shadows evolve temporally alongside moving objects and varying lighting conditions.en_US
dcterms.abstractThis thesis addresses the core research question: Can machines learn to robustly interpret shadows in complex, real-world scenes, and can shadow perception enhance high-level vision tasks such as image colorization? To address this, three core challenges are explored. First, robust video shadow detection is tackled through the development of a large-scale dataset and a tailored neural architecture designed for complex, dynamic environments. Second, an unsupervised, CLIP-driven framework is proposed to enable scalable shadow detection without the need for manual annotations. Third, the thesis explores the integration of shadow understanding into downstream vision tasks—specifically image colorization—demonstrating how shadow-aware models can improve semantic consistency and visual realism. Together, these contributions aim to advance machine perception of shadows as a critical component in video and image understanding.en_US
dcterms.abstractTo overcome these challenges, we first introduce the Complex Video Shadow Dataset (CVSD) in Chapter 3, a large-scale, diverse dataset comprising 196 videos and over 19,000 frames with detailed shadow annotations in complex scenarios. Leveraging CVSD, we propose a novel two-stage, fully supervised shadow detection framework incorporating spatial and temporal adaptation modules to effectively model dynamic shadow patterns also in Chapter 3. To reduce annotation dependency, we develop CDCNet, a CLIP-driven Divide-and-Conquer network enabling fully unsupervised video shadow detection by exploiting semantic priors and motion cues in chapter 4. Finally, we present the first shadow-aware image colorization network that incorporates a Shadow-Aware Block, explicitly integrating shadow features to improve color accuracy and semantic fidelity in shadowed regions in Chapter 5.en_US
dcterms.abstractOur contributions demonstrate that precise shadow perception not only enhances the robustness of shadow detection in complex, real-world environments but also significantly improves the performance of high-level vision tasks. By bridging the gap between low-level shadow understanding and semantic image colorization, this work highlights the critical role of shadow reasoning in advancing broader visual perception capabilities.en_US
dcterms.abstractThe main research contents can be summarized as follows:en_US
dcterms.abstract1. Complex Video Shadow Dataset and Detection Framework: To facilitate the exploration of video shadow detection in real-world scenarios, we introduce the Complex Video Shadow Dataset (CVSD)—a large-scale dataset comprising 196 video clips across 149 scene categories, capturing diverse and challenging shadow patterns. The dataset includes 19,757 frames with complex shadow patterns, making it the first dataset for complex real-world video shadow detection. Building upon CVSD, we propose a two-stage training paradigm and a novel neural architecture specifically designed for dynamic shadow scenarios. By framing shadow detection as a conditioned feature adaptation problem, we introduce temporal- and spatial-adaptation blocks to effectively capture spatiotemporal context, significantly enhancing detection accuracy in complex, real-world environments.en_US
dcterms.abstract2. Unsupervised Video Shadow Detection with CDCNet: We propose CDCNet, a CLIP-driven Divide-and-Conquer Network designed for unsupervised video shadow detection. The architecture incorporates two key modules—Spatial Divide and Temporal Conquer—which collaboratively capture spatial details and temporal coherence without requiring annotated supervision. By eliminating the need for pixel-level labels, CDCNet addresses the high cost and effort associated with manual dataset annotation. Extensive experiments demonstrate that CDCNet exhibits strong generalizability on unseen datasets, highlighting its effectiveness and scalability.en_US
dcterms.abstract3. Shadow-Aware Image Colorization: We propose the first shadow-aware image colorization method, based on an extensive review of two decades of colorization techniques. Our approach employs a dual-branch network enhanced by a novel Shadow-Aware Block, which explicitly incorporates shadow-specific features into the colorization process. This enables more accurate and semantically consistent colorization in scenes with shadow conditions by effectively distinguishing between shadowed and non-shadowed regions.en_US
dcterms.abstractThe results of our research have been published in or to be submitted to several leading conferences and journals in computer vision and computer graphics, including The Visual Computer (Journal), ECCV (Conference) and TVCG (Journal).en_US
dcterms.extentxix, 162 pages : color illustrationsen_US
dcterms.isPartOfPolyU Electronic Thesesen_US
dcterms.issued2026en_US
dcterms.educationalLevelPh.D.en_US
dcterms.educationalLevelAll Doctorateen_US
dcterms.accessRightsopen accessen_US

Files in This Item:
File Description SizeFormat 
8827.pdfFor All Users98.83 MBAdobe PDFView/Open


Copyright Undertaking

As a bona fide Library user, I declare that:

  1. I will abide by the rules and legal ordinances governing copyright regarding the use of the Database.
  2. I will use the Database for the purpose of my research or private study only and not for circulation or further reproduction or any other purpose.
  3. I agree to indemnify and hold the University harmless from and against any loss, damage, cost, liability or expenses arising from copyright infringement or unauthorized usage.

By downloading any item(s) listed above, you acknowledge that you have read and understood the copyright undertaking as stated above, and agree to be bound by all of its terms.

Show simple item record

Please use this identifier to cite or link to this item: https://theses.lib.polyu.edu.hk/handle/200/14402