Author: Song, Ziyang
Title: Segmentation and reconstruction of 3D objects from dynamic scenes
Advisors: Yang, Bo (COMP)
Degree: Ph.D.
Year: 2026
Department: Department of Computing
Pages: xvi, 119 pages : color illustrations
Language: English
Abstract: Identifying multiple objects and recovering their 3D geometries from complex scenes are fundamental tasks in computer vision, which can benefit the development of intelligent machines such as robots and automobiles, as well as the 3D asset collection for filming, gaming, and AR/VR. A plethora of works have devoted to the 3D object segmentation and reconstruction from static scenes, ignoring the common dynamic scenes with objects moving in our real world. Compared to static scenes, dynamic scenes present different cues along with challenges for the segmentation and reconstruction of 3D objects. This thesis makes three core contributions to leverage the cues and overcome the challenges presented in dynamic scenes, addressing a spectrum of tasks ranging from segmentation and reconstruction of rigid objects to reconstruction of complex deformable objects.
In Chapter 3, we start by the segmentation of multiple rigid objects from 3D point clouds. While existing methods usually require a large amount of human annotations for full supervision, the object motions in point cloud sequences provide additional cues to identify objects. With this observation, we propose the first unsupervised method, called OGC, to simultaneously identify multiple 3D objects in a single forward pass, without needing any type of human annotations in training networks. The key to our approach is to fully leverage the dynamic motion patterns over sequential point clouds as supervision signals to automatically discover rigid objects.
In Chapter 4, we explore the reconstruction of dynamic scenes with multiple rigid objects from monocular videos. Existing works formulate it into finding a single most plausible solution by adding various constraints such as depth priors and strong geometry constraints, ignoring the fact that there could be infinitely many 3D scene representations corresponding to a single dynamic video, owing to the interplay between object dynamics and camera motions. In contrast, we learn all plausible 3D scene configurations that match the input video, instead of just inferring a specific one. The key to our approach is a simple yet innovative object scale network together with a joint optimization module to learn an accurate scale range for every dynamic 3D rigid object. This allows us to sample as many faithful 3D scene configurations as possible during inference.
In Chapter 5, we extend our study to the reconstruction of deformable objects from monocular videos. This is also a hard problem due to the lack of multi-view observations at each time step. Although we may observe different sides of objects along the monocular video due to object motions, there are always some object parts occluded at each time step. Inferring the motion states of occluded object parts is the key for full shape reconstruction and free view synthesis of deformable objects just from monocular videos. To achieve this goal, a dynamic fusion framework is built to stably evolve the 3D positions of occluded object parts from the moments they are observed, paired with a set of dynamically updated control nodes to interpolate smooth and compact motions for the whole object.
Overall, this thesis develops a series of novel algorithms to tackle the segmentation and reconstruction of 3D objects for dynamic scenes, arguably pushing the boundary of 3D computer vision and machine learning.
Rights: All rights reserved
Access: open access

Files in This Item:
File Description SizeFormat 
8831.pdfFor All Users41.59 MBAdobe PDFView/Open


Copyright Undertaking

As a bona fide Library user, I declare that:

  1. I will abide by the rules and legal ordinances governing copyright regarding the use of the Database.
  2. I will use the Database for the purpose of my research or private study only and not for circulation or further reproduction or any other purpose.
  3. I agree to indemnify and hold the University harmless from and against any loss, damage, cost, liability or expenses arising from copyright infringement or unauthorized usage.

By downloading any item(s) listed above, you acknowledge that you have read and understood the copyright undertaking as stated above, and agree to be bound by all of its terms.

Show full item record

Please use this identifier to cite or link to this item: https://theses.lib.polyu.edu.hk/handle/200/14406