| Author: | Wang, Kangzhong |
| Title: | Towards multimodal analysis of human interaction patterns in group context |
| Advisors: | Ngai, Grace (COMP) Leong, Hong Va (COMP) |
| Degree: | Ph.D. |
| Year: | 2026 |
| Department: | Department of Computing |
| Pages: | xviii, 155 pages : color illustrations |
| Language: | English |
| Abstract: | Understanding how humans communicate, respond, and learn in social and educational settings is central to developing human-aware artificial intelligence. This thesis investigates how multimodal behavioural signals—visual, acoustic, and linguistic—can be modelled to interpret key aspects of human interaction, including responsiveness, agreement, and learning outcomes. Through three consecutive studies, it establishes a computational framework for analysing interactional behaviour across conversational and learning contexts. The first study focuses on backchannel detection in group conversations, where listeners' subtle nonverbal cues such as head nods and facial movements indicate attentiveness and understanding. A temporal visual attention framework is proposed to capture both micro- and macro-level temporal dependencies across behavioural channels. Experiments on the MPIIGroupInteraction and CCDb datasets demonstrate that the framework effectively recognises diverse backchannel behaviours and outperforms contemporary baselines, highlighting the importance of temporal reasoning and motion-based feature representation. Building on this foundation, the second study extends the analysis from behavioural events to attitudinal meaning by estimating agreement intensity from multimodal backchannel segments. Audio and visual modalities are integrated through adaptive temporal fusion to model how listeners express varying degrees of agreement. Results show that combining visual motion cues with prosodic and spectral features could help improve predictive accuracy. Incorporating speaker diarisation further enhances role-specific modelling, revealing that agreement is often conveyed implicitly through coordinated facial and vocal cues. The final study applies multimodal modelling to an educational context by predicting students' self-reported learning gains in online service-learning programmes. Using audio, video, and transcript features from synchronous sessions, the analysis identifies behavioural indicators associated with different dimensions of learning outcomes—intellectual, social, civic, and intrapersonal. Screen-content features that capture the semantic richness and structure of instructional materials are found to best predict intellectual learning, while visual participation cues—such as camera usage and body posture—correlate more strongly with social and civic learning. These findings indicate that both communicative behaviour and visible participation contribute to effective collaborative learning. Overall, the three studies establish a coherent research trajectory from low-level behavioural detection to high-level outcome prediction. Methodologically, the thesis contributes interpretable frameworks for temporal attention and multimodal fusion that generalise across conversational and learning domains. Conceptually, it advances the computational understanding of how responsiveness, agreement, and learning outcomes can be inferred from multimodal behaviour, supporting the development of human-aware AI systems capable of perceiving and facilitating meaningful interaction. |
| Rights: | All rights reserved |
| Access: | open access |
Copyright Undertaking
As a bona fide Library user, I declare that:
- I will abide by the rules and legal ordinances governing copyright regarding the use of the Database.
- I will use the Database for the purpose of my research or private study only and not for circulation or further reproduction or any other purpose.
- I agree to indemnify and hold the University harmless from and against any loss, damage, cost, liability or expenses arising from copyright infringement or unauthorized usage.
By downloading any item(s) listed above, you acknowledge that you have read and understood the copyright undertaking as stated above, and agree to be bound by all of its terms.
Please use this identifier to cite or link to this item:
https://theses.lib.polyu.edu.hk/handle/200/14707

