Author: Li, Haotian
Title: A study on RGB-thermal semantic segmentation for autonomous driving
Advisors: Sun, Yuxiang (ME)
Chu, Kar Hang Henry (ME)
Degree: Ph.D.
Year: 2026
Department: Department of Mechanical Engineering
Pages: xxiv, 128 pages : color illustrations
Language: English
Abstract: Semantic scene understanding is a fundamental capability of autonomous driving systems, providing crucial perception information for downstream tasks such as localization, navigation, and decision-making. In complex lighting scenarios such as nighttime and strong glare, semantic segmentation methods relying solely on RGB images exhibit performance degradation. Thermal images, being less affected by changes in ambient lighting, can supplement RGB images. Therefore, RGB-Thermal (RGB-T) semantic segmentation methods have received widespread attention in recent years. Although current methods have made significant progress in segmentation capability, several key issues remain unresolved, such as: lack of interpretability of multi-modal fusion strategies, lack of temporal consistency in segmentation results, limited training data, and uneven spatial distribution of segmentation performance. These issues limit the application potential of RGB-T semantic segmentation in autonomous driving.
To address these challenges, the main contributions of this study are as follows:
First, addressing the insufficient interpretability of multi-modal fusion, this dissertation analyzes the inherent characteristics of RGB and thermal images under different lighting conditions and proposes an Illumination-Guided Fusion Network (IGFNet). This network generates a weight mask for each pixel through an Illumination Estimation Module (IEM), quantifying the importance of RGB information at that location. Subsequently, the Illumination-Guided Cross-Modal Rectification Module (IGCM-RM) adaptively rectifies and fuses RGB and thermal features, thereby more effectively utilizing multi-modal information.
Second, to alleviate the problem of inconsistent segmentation results between consecutive frames, this dissertation proposes a temporal consistency enhancement framework. Existing methods mostly focus on single-frame segmentation performance, neglecting temporal consistency, which may lead to hesitation in autonomous driving decisions. To alleviate this, this dissertation designs a Virtual View Image Generation (VVIG) module, which uses a virtual view transformation (VVT) matrix to synthesize the next frame image, and proposes two consistency loss functions to constrain the segmentation results of consecutive frames. Simultaneously, Consistent Accuracy (CA) is introduced as an evaluation metric, balancing segmentation accuracy and temporal consistency. Experiments demonstrate that this method improves temporal consistency while maintaining high segmentation accuracy.
Third, this dissertation proposes SyntheticSeg, a synthetic data augmentation method, to alleviate the problems of scarce training data and class imbalance. First, a large-scale RGB-T synthetic dataset is constructed using a layout-to-image generative model for joint training with real data. Furthermore, a sampling mechanism considering the scarcity of semantic layout and segmentation difficulty is designed to selectively sample from synthetic data, thereby alleviating class imbalance. Experiments show that SyntheticSeg improves performance by 2.2% compared to baseline methods, with particularly significant results on rare and difficult-to-segment classes such as guardrail and car stop.
Fourth, this dissertation is the first to systematically study the spatial performance imbalance problem in RGB-T semantic segmentation. The study found that in autonomous driving scenarios, the segmentation performance of the image central regions is often weaker than that of the edge regions. The finding challenges the conventional wisdom that "more training samples lead to better performance" and points out that target complexity and image quality are the main causes of spatial bias. Based on this, a Gaussian-guided Regional Balancing Masking (GRBM) method is proposed, which enhances the model's utilization of the thermal features of the central region through masking; simultaneously, a spatially weighted loss function is designed to further improve model performance. Experimental results show that this method effectively alleviates spatial bias and improves overall segmentation accuracy.
Rights: All rights reserved
Access: open access

Files in This Item:
File Description SizeFormat 
9033.pdfFor All Users22.97 MBAdobe PDFView/Open


Copyright Undertaking

As a bona fide Library user, I declare that:

  1. I will abide by the rules and legal ordinances governing copyright regarding the use of the Database.
  2. I will use the Database for the purpose of my research or private study only and not for circulation or further reproduction or any other purpose.
  3. I agree to indemnify and hold the University harmless from and against any loss, damage, cost, liability or expenses arising from copyright infringement or unauthorized usage.

By downloading any item(s) listed above, you acknowledge that you have read and understood the copyright undertaking as stated above, and agree to be bound by all of its terms.

Show full item record

Please use this identifier to cite or link to this item: https://theses.lib.polyu.edu.hk/handle/200/14613