| Author: | Chen, Zuyao |
| Title: | Towards generalized scene graph generation |
| Advisors: | Chen, Changwen (COMP) |
| Degree: | Ph.D. |
| Year: | 2026 |
| Department: | Department of Computing |
| Pages: | xxii, 153 pages : color illustrations |
| Language: | English |
| Abstract: | Scene Graph Generation (SGG) is a crucial task in computer vision that seeks to represent images as structured graphs, capturing objects and their semantic relationships. While existing SGG methods have made remarkable progress, they often suffer from three key limitations: closed set assumptions, reliance on costly manual annotations, and a lack of alignment between training objectives and structured output formats. This thesis aims to overcome these limitations by exploring generalized SGG across four dimensions: open-vocabulary recognition, language-based supervision, reinforcement learning for structured prediction, and scene graph-guided image generation. We first propose a unified transformer-based framework, OvSGTR, for Fully Open-Vocabulary SGG, enabling the recognition of novel objects and relations through visual-language alignment and retention. We then introduce a weakly-supervised paradigm, GPT4SGG, that reverses the traditional weakly-SGG pipeline by leveraging large language models (LLMs) to synthesize scene graphs from region and holistic descriptions. To further align model optimization with structured evaluation, we develop a reinforcement learning framework R1-SGG , which employs rule-based rewards for fine-grained and format-consistent scene graph output. Finally, we explore the inverse task of using scene graphs as controllable interfaces for image generation, and present a novel benchmark, Scene-Bench, with a new metric, SGScore, for evaluating object and relation fidelity in generated images. Extensive experiments conducted on widely used public datasets such as Visual Genome, GQA, and PSG validate the effectiveness and generalizability of the proposed approaches across a diverse range of evaluation scenarios. These include traditional closed-set settings, where object and relation categories are predefined; open-vocabulary and zero-shot setups, which test the model's ability to recognize and reason about unseen concepts; and generative tasks, such as image synthesis guided by structured scene representations. The proposed methods exhibit consistent improvements in structural accuracy, semantic coverage, and compositional generalization under these varied conditions. Collectively, the models, algorithms, and benchmarks introduced in this thesis contribute to the development of scalable, interpretable, and robust scene understanding systems that are capable of operating effectively in open-world environments, where visual content is complex, diverse, and continually evolving. |
| Rights: | All rights reserved |
| Access: | open access |
Copyright Undertaking
As a bona fide Library user, I declare that:
- I will abide by the rules and legal ordinances governing copyright regarding the use of the Database.
- I will use the Database for the purpose of my research or private study only and not for circulation or further reproduction or any other purpose.
- I agree to indemnify and hold the University harmless from and against any loss, damage, cost, liability or expenses arising from copyright infringement or unauthorized usage.
By downloading any item(s) listed above, you acknowledge that you have read and understood the copyright undertaking as stated above, and agree to be bound by all of its terms.
Please use this identifier to cite or link to this item:
https://theses.lib.polyu.edu.hk/handle/200/14652

