Full metadata record
DC FieldValueLanguage
dc.contributorDepartment of Industrial and Systems Engineeringen_US
dc.contributor.advisorZheng, Pai (ISE)en_US
dc.creatorWang, Tian-
dc.identifier.urihttps://theses.lib.polyu.edu.hk/handle/200/14676-
dc.languageEnglishen_US
dc.publisherHong Kong Polytechnic Universityen_US
dc.rightsAll rights reserveden_US
dc.titleTowards human-centric smart manufacturing : a systematic vision-language model-based methodology for human-robot interactive, task planning, navigation, and manipulationen_US
dcterms.abstractRooted in the principles of Industry 5.0, Human-centric Smart Manufacturing (HSM) prioritizes human welfare within manufacturing processes, leveraging Artificial Intelligence (AI) to simultaneously improve operational efficiency and employee well-being. Under this paradigm, Human-Robot Interaction (HRI) technologies are central because they unify human problem-solving and decision-making strengths with the precision and efficiency of robotic systems, resulting in adaptive, intelligent, and superior manufacturing processes.en_US
dcterms.abstractNevertheless, existing HRI works commonly face several challenges: 1) Data scarcity for specific tasks - limited availability of data for particular tasks can hinder the development of efficient models or systems; 2) Requirement for re-training for different scenarios - many existing approaches necessitate re-training or fine-tuning models for different scenarios, which can be time-consuming and impractical; 3) Predominantly pre-defined manners - most interaction approaches are predetermined, offering limited adaptability and failing to support open-ended or natural exchanges between humans and robots..en_US
dcterms.abstractRecently, Vision-Language Models (VLMs) and Large Language Models (LLMs) have attracted significant attention in the field of HRI, and they hold strong potential as effective solutions to address the current limitations faced in HRI. Firstly, as pretrained on vast repositories of world knowledge, VLMs and LLMs possess the capacity to generalize across diverse scenarios and tasks, thereby addressing the limitations of previous models that suffered from insufficient datasets and required frequent retraining, thus rendering HRI more adaptable and effective across multiple domains. Secondly, VLMs strengthen robotic intelligence by fusing visual and linguistic inputs with unified reasoning, thereby mirroring the multimodal perception characteristic of humans. Thirdly, VLMs and LLMs enhance user experience by enabling flexible interaction and communication, thereby making HRI more natural, efficient, and intuitive.en_US
dcterms.abstractIn this case, this research aims to develop a VLM-based HRI methodology that makes it possible for robots to engage with humans naturally and systematically follow their commands to collaborate on accomplishing complex tasks—such as supplementing assembly materials, repairing malfunctioning machines, and detecting potential safety issues—toward future human-centric smart manufacturing. To realize this goal, three fundamental robotic tasks warrant exploration: robot system task planning, mobile robot navigation, and robot arm manipulation. Accordingly, this research proposes a systematic vision-language model-based methodology for human-robot interactive task planning, navigation, and manipulation in manufacturing.en_US
dcterms.abstractChapter 2 provides a survey of vision-language techniques applied to HRI within human-centered smart manufacturing, and concludes by outlining unresolved challenges in task planning, navigation, and manipulation.en_US
dcterms.abstractIn Chapter 3, the background and research challenge of vision and language-based task planning are introduced, and a coarse-to-fine vision-grounded Chain-of-Thought (CoT) reasoning approach is proposed to address the unforeseen challenge faced by smart manufacturing systems. By integrating multi-granularity visual parsing with a coarse-to-fine CoT framework, our method jointly leverages global scene context and local visual perspectives to enable interpretable, zero-shot reasoning, while generating clear and precise robot-executable task plans. The evaluation results indicate that our framework attains remarkable success and generalization in challenging as well as novel scenarios, emphasizing its reliability and applicability to real-world smart manufacturing environments.en_US
dcterms.abstractIn Chapter 4, the background and research gaps of vision and language navigation is introduced. Then proposes a human-guided mobile robot navigation method for unstructured manufacturing environment. Finally, we use some experiments to evaluate these methods. Learning-based VLN methods heavily rely on data and are constrained by the availability of a fixed number of scenes for training. This limitation prevents them from effectively addressing the challenges of realistic application scenarios, where objects may exist outside the pre-defined domains and language instructions can be more diverse. Additionally, although the current method shows promising performance in simulated environments, it faces challenges when applied in real-world scenarios due to sensor noise. This noise can lead to the loss of geometric details in the reconstructed maps, resulting in inaccurate segmentation of the navigation target. To tackle these issues, we integrate classical 3D reconstruction techniques with zero-shot VLM, aiming to construct an open-vocabulary 3D semantic mapping framework. Furthermore, an LLM is utilized to interpret natural language commands and produce control code for robots, thereby facilitating navigation supported by vision and language.en_US
dcterms.abstractIn Chapter 5, the background and research challenge of vision and language-based robot manipulation is firstly introduced, and then a foundational model-based framework is proposed for 3D Hierarchical Affordance Reasoning, concentrating on high-precision tool operations in factory environments. Within this framework, two components are proposed: 3D Joint Tool-Part Affordance Grounding enables open-vocabulary 3D affordance detection, jointly grounding tool and part interactions; task-aware Trajectory Generation leverages an LLM alongside motion affordance priors to generate task-specific motion trajectories, enabling robots to perform complex tool-use operations with high precision and adaptability. Extensive experiments show that the proposed hierarchical affordance reasoning framework demonstrates substantial improvements over current leading approaches. Experiments carried out on actual manufacturing systems confirm the effectiveness and stability of the proposed approach, highlighting its suitability for real-world industrial settings that require strong dexterity, logical reasoning, and generalization capabilities.en_US
dcterms.abstractIt is hoped that this work can create a multi-robot intelligent system capable of achieving human-centric natural visual-language interaction. Under human guidance, this system will be able to accomplish various tasks in complex factory environments in future human-robot symbiotic manufacturing paradigms.en_US
dcterms.isPartOfPolyU Electronic Thesesen_US
dcterms.issued2026en_US
dcterms.educationalLevelPh.D.en_US
dcterms.educationalLevelAll Doctorateen_US
dcterms.accessRightsopen accessen_US

Files in This Item:
File Description SizeFormat 
9120.pdfFor All Users10.63 MBAdobe PDFView/Open


Copyright Undertaking

As a bona fide Library user, I declare that:

  1. I will abide by the rules and legal ordinances governing copyright regarding the use of the Database.
  2. I will use the Database for the purpose of my research or private study only and not for circulation or further reproduction or any other purpose.
  3. I agree to indemnify and hold the University harmless from and against any loss, damage, cost, liability or expenses arising from copyright infringement or unauthorized usage.

By downloading any item(s) listed above, you acknowledge that you have read and understood the copyright undertaking as stated above, and agree to be bound by all of its terms.

Show simple item record

Please use this identifier to cite or link to this item: https://theses.lib.polyu.edu.hk/handle/200/14676