Active Education & Skills Arts, Culture & Design

Learning from Collaborative Storytelling: Enhancing Visual Scene Understanding through Human-Robot Interaction

In plain English

AI plain-English summary

A robot and a human sit down together to watch a video, then take turns filling in the gaps in what the robot understands about the scene. This matters because robots today struggle to grasp unfamiliar environments when they cannot rely on perfect, pre-programmed data. Traditional visual systems fail in messy, real-world settings—cluttered kitchens, busy hospital wards, or a visually impaired person’s living room. The project tackles this by treating scene understanding as a collaborative storytelling exercise: the robot spots objects and actions, the human provides context and narrative, and together they build a richer, shared picture of the environment. If the approach works, robots could become far more useful in assistive care, education, and medical support—for example, guiding someone with sight loss through a cluttered room or helping a patient perform rehabilitation exercises correctly. The research is not yet ready for deployment; it is fundamental work on how machines learn from human interaction. But similar co-learning approaches have previously unlocked breakthroughs in language models and computer vision, suggesting that teaching robots to “talk through” what they see could quietly reshape how machines navigate the messy, unpredictable world people inhabit every day.

View original technical description
The research project aims to revolutionise how robots comprehend and interact with the physical world by combining advanced techniques in video analysis, knowledge acquisition, and collaborative storytelling. In recent years, there has been a growing emphasis on connecting language and the physical environment to improve robots' understanding of human activities. However, providing robots with complete knowledge of a new environment remains a significant challenge. This project seeks to address this challenge through a collaborative storytelling approach. To tackle this challenge, the project is developing a framework wherein robots and humans collaborate in a storytelling task. Within this framework, robots and humans work together to analyze a video depicting an environment, jointly filling in the narrative gaps initiated by the robot. This approach integrates sophisticated visual scene analysis and knowledge acquisition through conversational storytelling, ultimately providing robots with a thorough understanding of unfamiliar surroundings. The primary aim of this project is to achieve a significant breakthrough by progressively enhancing robots' interpretation of visual scenes through advanced techniques such as video analysis, knowledge graph construction, and human-robot cooperation. The ultimate objective is to empower robots with a detailed comprehension of scenes, opening up possibilities for applications in assistive and care robotics, education, and medical support, and addressing challenges like aiding the visually impaired and facilitating physical rehabilitation training. With a particular focus on improving human-robot interaction in domestic settings and the robot's ability to comprehend new environments and activities, this project addresses the limitations of traditional visual understanding approaches in dynamic real-world scenarios where obtaining perfect input is not guaranteed. By embracing a co-learning approach where humans and robots collaborate as storytellers, the robot gains fresh insights into the visual environment, anchoring human narratives in existing visual data, expanding its visual knowledge over time, and contributing to the narrative using previously unmentioned visual elements. This innovative approach combines video analysis and dialogue to tackle the intricate challenge of nuanced scene comprehension from video data, promising significant advancements in human-robot interaction and visual understanding.

View the original record at the funder ↗

Researchers

Yanchao Yu (Principal Investigator)

Related Research

Grants with similar aims, by meaning.

AI-Driven Conversational Storytelling with Humans in Natural Language
LISI - Learning to Imitate Nonverbal Communication Dynamics for Human-Robot Social Interaction
Child-Robot Communication and Collaboration: Edutainment, Behavioural Modelling and Cognitive Development in Typically Developing and Autistic Spectrum Children
Creating a Powerful Navigation Tool for image and video datasets
CiViL: Common-sense- and Visually-enhanced natural Language generation

Original classification

Research and Innovation

Plain English summaries and category classifications on this site are generated by AI and may not perfectly reflect the original research.