Completed Computing & AI Arts, Culture & Design

Seebibyte: Visual Search for the Era of Big Data

In plain English

AI plain-English summary

A computer vision system will learn to search through millions of hours of video and billions of images as easily as Google searches the web. Today, searching visual data is slow and largely manual—a BBC archive of decades of programmes, for instance, cannot be searched for a specific scene or character. The same problem plagues medical imaging: radiologists and researchers still review ultrasound or X-ray scans frame by frame. This programme aims to automate that process. It will develop algorithms that can recognise objects, actions, and spatial layouts in video, and count or delineate structures in images, using deep learning that requires very little human supervision. The second theme extends these methods beyond standard camera footage to medical sensors like ultrasound and X-ray, which have different noise and dimensionality. If successful, the work could transform how the BBC’s iPlayer lets viewers skip to a specific hug between characters, enable cheaper imaging diagnostics in the NHS, and accelerate drug discovery by automatically analysing microscopy images. Open-source software and datasets will be released to spread the tools across academia and industry.

View original technical description
The Programme is organised into two themes. Research theme one will develop new computer vision algorithms to enable efficient search and description of vast image and video datasets - for example of the entire video archive of the BBC. Our vision is that anything visual should be searchable for, in the manner of a Google search of the web: by specifying a query, and having results returned immediately, irrespective of the size of the data. Such enabling capabilities will have widespread application both for general image/video search - consider how Google's web search has opened up new areas - and also for designing customized solutions for searching. A second aspect of theme 1 is to automatically extract detailed descriptions of the visual content. The aim here is to achieve human like performance and beyond, for example in recognizing configurations of parts and spatial layout, counting and delineating objects, or recognizing human actions and inter-actions in videos, significantly superseding the current limitations of computer vision systems, and enabling new and far reaching applications. The new algorithms will learn automatically, building on recent breakthroughs in large scale discriminative and deep machine learning. They will be capable of weakly-supervised learning, for example from images and videos downloaded from the internet, and require very little human supervision. The second theme addresses transfer and translation. This also has two aspects. The first is to apply the new computer vision methodologies to `non-natural' sensors and devices, such as ultrasound imaging and X-ray, which have different characteristics (noise, dimension, invariances) to the standard RGB channels of data captured by `natural' cameras (iphones, TV cameras). The second aspect of this theme is to seek impact in a variety of other disciplines and industry which today greatly under-utilise the power of the latest computer vision ideas. We will target these disciplines to enable them to leapfrog the divide between what they use (or do not use) today which is dominated by manual review and highly interactive analysis frame-by-frame, to a new era where automated efficient sorting, detection and mensuration of very large datasets becomes the norm. In short, our goal is to ensure that the newly developed methods are used by academic researchers in other areas, and turned into products for societal and economic benefit. To this end open source software, datasets, and demonstrators will be disseminated on the project website. The ubiquity of digital imaging means that every UK citizen may potentially benefit from the Programme research in different ways. One example is an enhanced iplayer that can search for where particular characters appear in a programme, or intelligently fast forward to the next `hugging' sequence. A second is wider deployment of lower cost imaging solutions in healthcare delivery. A third, also motivated by healthcare, is through the employment of new machine learning methods for validating targets for drug discovery based on microscopy images

View the original record at the funder ↗

Researchers

Alison Noble (Co-Investigator)Andrea Vedaldi (Co-Investigator)Andrew Zisserman (Principal Investigator)Jens Rittscher (Co-Investigator)Philip Torr (Co-Investigator)

Related Research

Grants with similar aims, by meaning.

Deep Learning-based Computer Vision Methods
Visual AI: An Open World Interpretable Visual Transformer
Visual Sense. Tagging visual data with semantic descriptions
Detailed and Deep Image Understanding
Multimodal Video Search by Examples (MVSE)

Original classification

Research Grant

Plain English summaries and category classifications on this site are generated by AI and may not perfectly reflect the original research.