| Author: | Zhang, Zihui |
| Title: | Learning 3D scene semantics without human annotations |
| Advisors: | Yang, Bo (COMP) Li, Bo (COMP) Xiao, Bin (COMP) |
| Degree: | Ph.D. |
| Year: | 2026 |
| Department: | Department of Computing |
| Pages: | xiv, 121 pages : color illustrations |
| Language: | English |
| Abstract: | Human beings live in a 3D physical world and have evolved efficient visual processing mechanisms to understand 3D scenes. In contrast, machines face significant difficulties in inferring semantic information from 3D visual inputs. The current 3D scene understanding models often rely on human annotations, which are costly, time-consuming, and prone to bias, highlighting the need for unsupervised approaches. This thesis is dedicated to developing intelligent systems that can comprehend the semantics of 3D real-world environments without manual annotations. To this end, we derive semantic priors from spatial geometry, enabling them to be discriminative, generalizable, and aware of objectness. These priors are integrated into a learning framework alongside corresponding identification strategies, forming a cohesive pipeline for unsupervised 3D semantic understanding. In Chapter 3, an unsupervised method is proposed to automatically discover multiple 3D semantics via a superpoint progressively growing strategy during training, enabling effective learning of meaningful semantic elements. Experiments on various indoor and outdoor datasets validate its effectiveness. In Chapter 4 also pays attention to learn semantics from complex 3D scenes without human labels. By leveraging semantic priors transferred from self-supervised 2D features to 3D and a novel global grouping strategy for superpoints, accurate semantic pseudo-labels are obtained in the frequency domain. This chapter achieves state-of-the-art performances, expanding the boundaries of unsupervised learning and reducing annotation costs. In Chapter 5, a two-stage pipeline is proposed to learn objectness from 3D scenes. The object prior is first encoded within a generative model, where a Reinforcement Learning-based embodied agent is employed for object discovery. Experiments in this chapter demonstrates excellence in single and multi-category object segmentation across various benchmarks. Overall, this thesis makes significant contributions to the field of unsupervised 3D scene semantic learning, advancing the development of 3D scene understanding from the supervised learning stage into unsupervised learning fashion. By breaking through the limitations of heavy reliance on labeled data, it realizes effective semantic parsing of complex 3D scenes, greatly reducing the cost of practical application. In doing so, it lays a solid theoretical and technical foundation for the widespread application of 3D scene understanding in fields such as autonomous driving, augmented reality, and robotics, where labeled data is scarce or expensive to obtain, thus promoting the further evolution of intelligent perception systems. |
| Rights: | All rights reserved |
| Access: | open access |
Copyright Undertaking
As a bona fide Library user, I declare that:
- I will abide by the rules and legal ordinances governing copyright regarding the use of the Database.
- I will use the Database for the purpose of my research or private study only and not for circulation or further reproduction or any other purpose.
- I agree to indemnify and hold the University harmless from and against any loss, damage, cost, liability or expenses arising from copyright infringement or unauthorized usage.
By downloading any item(s) listed above, you acknowledge that you have read and understood the copyright undertaking as stated above, and agree to be bound by all of its terms.
Please use this identifier to cite or link to this item:
https://theses.lib.polyu.edu.hk/handle/200/14418

