Author: Ling, Tao
Title: Efficient visual learning and reasoning system with spurious and redundant data on edge computing
Advisors: Wang, Dan (COMP)
Degree: Ph.D.
Year: 2026
Department: Department of Computing
Pages: xvi, 148 pages : color illustrations
Language: English
Abstract: The proliferation of deep visual learning has spurred the development of sophisticated applications deployed on resource-constrained edge devices, often operating in collaboration with the cloud. The practical efficiency and robustness of these edge-cloud visual systems, however, are fundamentally challenged by the inherent complexity of real-world data. This complexity manifests as performance-degrading artifacts at different stages of the machine learning lifecycle: during collaborative training, spurious data artifacts can corrupt the learning process, while during inference, informational redundancy in visual data can create debilitating latency. This thesis presents a unified approach to designing efficient and robust edge visual systems by systematically tackling these two critical data challenges. We make the following original contributions.
First, we address the impact of spurious data artifacts on Federated Learning (FL), a cornerstone of privacy-preserving edge-based training. Client-side watermarks, while protecting data ownership, introduce spurious correlations that cause shortcut learning and degrade model accuracy. To counter this, we develop a multi-layered, system-aware framework. At the algorithmic level, we introduce Federated Morozov Regularization, a technique that adaptively mitigates the influence of watermarks. At the system level, we design LotusFL, which co-optimizes client selection and batch sizing by intelligently navigating the misleading utility signals caused by these artifacts. Finally, we integrate these solutions into LotusFA, a decoupled Federated Analytics system that efficiently performs watermark analysis as a separate, lightweight task. This provides the necessary intelligence to the FL system without encumbering the training loop, thereby maximizing both robustness and system efficiency. Our solutions, validated on a 40-device edge testbed, demonstrably improve model accuracy and reduce time-to-accuracy.
Secondly, we tackle the challenge of informational redundancy that plagues large-scale visual reasoning at the edge. Advanced Vision-Language Models (VLMs) employing Chain-of-Thought (CoT) reasoning suffer from prohibitive latency when processing extended or high-resolution visual data, such as long videos. To overcome this efficiency bottleneck, we design Compressed Scene Graph-enabled CoT (CSG-CoT), a novel paradigm that tames data redundancy. Inspired by video codec principles, our method compresses vast, redundant visual inputs into a compact scene graph (SG) composed of semantically-rich Key-SGs and lightweight Delta-SGs. By feeding this compressed representation to the VLM, we drastically reduce the input token length for the reasoning process. Experiments show that this approach slashes end-to-end latency by over 55-62%, enabling near real-time, accurate visual reasoning on edge-cloud platforms.
Thirdly, building upon this principle of redundancy reduction, we address the acute challenge of spatiotemporal redundancy in time-sensitive, mission-critical edge applications. Specifically, we focus on autonomous UAV path planning, where conventional VLM-CoT methods suffer from extreme latency by redundantly processing continuous and overlapping visual streams. To this end, we introduce Spatial Scene Graph-enabled CoT (SSG-CoT), a framework that abstracts the dynamic environment into a persistent yet lightweight spatial scene graph. We design a novel dual-pathway architecture where an efficient perception model handles routine updates to the graph's spatial relationships, reserving the powerful but computationally expensive VLM for high-level reasoning on key frames and critical decisions. This decoupling of continuous perception from complex reasoning is key to efficiency. Real-world deployments demonstrate that SSG-CoT reduces average mission completion time by over 45% compared to conventional methods while maintaining a high mission success rate.
In summary, this thesis advances the state-of-the-art in edge-cloud visual learning by providing a cohesive set of solutions to the critical problems of data complexity. By systematically mitigating spurious artifacts in federated training and taming informational redundancy in both large-scale static analysis and dynamic, real-time applications, our work enables the development of AI systems that are not only more accurate and robust but also significantly more efficient for practical deployment. All proposed methods are implemented and rigorously evaluated in real-world system contexts, demonstrating their effectiveness and paving the way for the next generation of intelligent edge applications.
Rights: All rights reserved
Access: open access

Files in This Item:
File Description SizeFormat 
8829.pdfFor All Users5.91 MBAdobe PDFView/Open


Copyright Undertaking

As a bona fide Library user, I declare that:

  1. I will abide by the rules and legal ordinances governing copyright regarding the use of the Database.
  2. I will use the Database for the purpose of my research or private study only and not for circulation or further reproduction or any other purpose.
  3. I agree to indemnify and hold the University harmless from and against any loss, damage, cost, liability or expenses arising from copyright infringement or unauthorized usage.

By downloading any item(s) listed above, you acknowledge that you have read and understood the copyright undertaking as stated above, and agree to be bound by all of its terms.

Show full item record

Please use this identifier to cite or link to this item: https://theses.lib.polyu.edu.hk/handle/200/14404