A computer is learning to watch a tennis match and describe what it sees, using both the video footage and the commentary to build its own internal grammar of the game. This matters because current AI struggles to understand sequences of real-world events the way humans do. Humans learn by linking what they see with what they hear, building mental rules about how events unfold. ACASVA aims to give machines that same ability by studying how visual and linguistic grammars interact. The researchers will analyse tennis footage and commentary separately, then combine them so that each mode of perception improves the other. They will also study how human gaze and rule-inference shape visual learning, feeding those insights back into the machine. If successful, the system could automatically annotate sports broadcasts, generating real-time captions and highlights without human editors. The broadcasting and online video search industries would benefit directly. More broadly, the fundamental science here—understanding how sight and sound combine to build event grammars—could eventually apply to any domain where machines need to interpret complex, rule-bound sequences of actions, from surveillance to manufacturing quality control.
View original technical description
The development of a machine that can autonomously understand and interpret patterns of real-world events remains a challenging goal in AI. Humans are able to achieve this by developing sophisticated internal representational structures for object and events and the grammars that connect them. ACASVA aims to investigate the interaction between visual and linguistic grammars in learning by developing grammars in a scenario where the number of different events is constrained, by a set of rules, to be small: a sport. We will analyse video footage of a game (e.g. tennis) and use computer vision techniques to progressively understand it as a sequence of (possibly overlapping) events, and build a grammar of events. We will do a similar audio/linguistic analysis on the commentary on the game. Both of these grammars will be used to build a representational structure for understanding the game. Visual representations are additionally constrained by the inference of game rules so that object-classification mechanisms are preferentially tuned to game-relevant entities like 'player' rather than game-irrelevant entities like 'crowd-member'. We will also investigate how the two modes, sight and sound, can influence each other in the learning process; interpretation of the video is affected by the linguistic grammar and vice versa. Furthermore, this coupling of modes will lead to improved recognition of both audio and video events when the grammars from the video modes are used to influence the audio recognition, and vice versa. The psychological component of the ACASVA correspondingly attempts to learn how these capabilities are developed in humans; how visual grammars are organized and employed in the learning problem, how these grammars are modified by prior linguistic knowledge of the domain, how visual grammars map onto linguistic grammars, and how game rule-inferences influence lower-level visual learning (determined via gaze-behaviour). These results will feedback into the machine-learning problem and vice versa, as well as providing a performance benchmark for the system.Potential beneficiaries of ACASVA (in addition to the knowledge beneficiaries within the fields of science and engineering) include the broadcasting and on-line video search industries.
Plain English summaries and category classifications on this site are generated by AI and may not perfectly reflect the original research.
Is something wrong? Let us know