Active Computing & AI Psychology & Behaviour
Research on Embodied Perception and Intelligent Interaction Methods Driven by Multimodal Large Models
Summary
Original abstract (not yet simplified)With the rapid development of artificial intelligence technology, embodied intelligence has shown tremendous potential in fields such as home robotics, medical robots, and industrial design. This research focuses on open-world embodied perception and intelligent interaction methods driven by multimodal large models, aiming to develop advanced models that achieve accurate 3D visual grounding and intelligent reasoning, thereby enhancing the perceptual and...
View original technical description
With the rapid development of artificial intelligence technology, embodied intelligence has shown tremendous potential in fields such as home robotics, medical robots, and industrial design. This research focuses on open-world embodied perception and intelligent interaction methods driven by multimodal large models, aiming to develop advanced models that achieve accurate 3D visual grounding and intelligent reasoning, thereby enhancing the perceptual and interactive capabilities of robots in complex environments. The study will explore how to utilize multimodal large models to process and integrate various sensory inputs, enabling robots to accurately perceive 3D environments, understand context, and make intelligent decisions. Additionally, this research will strive to improve robots' natural interaction capabilities with humans and the environment, including language understanding and situational adaptation. This research not only holds significant academic value and profound implications for advancing artificial intelligence and robotics technologies, but also provides innovative solutions to challenges in practical applications, contributing importantly to improving production efficiency, enhancing medical services, and elevating quality of life.
Related Research
Grants with similar aims, by meaning.
Enhanced AI Perception through Unified Joint Embedding of Multimodal Sensory Data
Multimodal Intention Prediction in Assistive Robotics
Multimodal active sensing and perception for human-robot collaborative.
Towards Generalised Multi-Robot Collaboration using VLM-Driven Semantic-Affordance Maps
Interactive Perception-Action-Learning for Modelling Objects
Original classification
HORIZONPlain English summaries and category classifications on this site are generated by AI and may not perfectly reflect the original research. Is something wrong? Let us know