| Author: | Liu, Ye |
| Title: | Towards interactive multi-modal content understanding |
| Advisors: | Chen, Chang Wen (COMP) |
| Degree: | Ph.D. |
| Year: | 2026 |
| Department: | Department of Computing |
| Pages: | xxviii, 204 pages : color illustrations |
| Language: | English |
| Abstract: | Human cognition process is inherently interactive, involving a dynamic interplay between perception and reasoning. In contrast, contemporary multi-modal AI systems typically operate under a single-pass processing paradigm, where perception and reasoning are performed sequentially and passively. While this approach has proven effective for constrained, well-defined tasks, it exhibits significant limitations in both efficiency and generalizability when applied to complex, open-ended scenarios. In such cases, this content-as-input paradigm struggles due to fundamental constraints in the single-pass pipeline, which hampers the comprehensive encoding of information-dense content for diverse downstream objectives. To address this challenge, we propose a new paradigm called interactive multimodal content understanding, in which multi-modal content is treated not as static input but as an interactive environment. Given a task instruction (e.g., a question about the content), the AI agent actively explores this environment, iteratively engaging in perception and reasoning to derive the final answer. This paradigm draws inspiration from human cognition, facilitating deeper exploration and more nuanced analysis during the reasoning process. This thesis begins with a systematic discussion on the proposed paradigm, followed by in-depth analysis of the technical challenges it presents and our strategies to address them. Specifically, Part I investigates the essential capabilities of video temporal perception, including visual-audio-text integration, effective video temporal grounding, and techniques for transferring these capabilities from images to videos. Part II focuses on combining fine-grained perception with open-ended reasoning through large language models, with particular attention to temporal and spatiotemporal reasoning. Finally, Part III synthesizes the previously discussed technologies, culminating in the development of an agentic, interactive multi-modal content understanding system capable of planning, grounding, verification, and question answering to tackle complex scenarios. Concluding remarks and future research directions are summarized at the end. |
| Rights: | All rights reserved |
| Access: | open access |
Copyright Undertaking
As a bona fide Library user, I declare that:
- I will abide by the rules and legal ordinances governing copyright regarding the use of the Database.
- I will use the Database for the purpose of my research or private study only and not for circulation or further reproduction or any other purpose.
- I agree to indemnify and hold the University harmless from and against any loss, damage, cost, liability or expenses arising from copyright infringement or unauthorized usage.
By downloading any item(s) listed above, you acknowledge that you have read and understood the copyright undertaking as stated above, and agree to be bound by all of its terms.
Please use this identifier to cite or link to this item:
https://theses.lib.polyu.edu.hk/handle/200/14405

