Completed Arts, Culture & Design Computing & AI

Machine Listening using Sparse Representations

In plain English

AI plain-English summary

A computer that listens like a human—picking a single voice out of a crowd, recognising a dog bark or a breaking window—remains stubbornly out of reach. Current audio-processing systems are brittle: speech recognition works well in quiet rooms but fails in noisy, real-world scenes, and the techniques used cannot be transferred to other sounds. This Fellowship aims to build a general-purpose machine listening system by applying a mathematical approach called sparse representations, where a sound is described using only a few key components chosen from a vast library of possibilities. The researcher will also collaborate with vision scientists and biologists to uncover organising principles common to all sensory systems. If successful, the work could transform how machines handle everyday audio. Practical impacts include more intelligent hearing aids and cochlear implants that filter out background noise, automated incident detection on railways and roads, and new tools for searching audio and video archives. The research is largely fundamental—it seeks to understand how to represent sound efficiently—but the potential applications span health, public safety, and the creative industries.

View original technical description
My aim for this Fellowship is to undertake a concerted programme of research in machine listening, the automatic analysis and understanding of sounds from the world around us. Through this research, and in collaboration with other international researchers, I aim to establish machine listening as a key enabling technology to improve our ability to interact with the world, leading to advances in many areas such as health, security and the creative industries.Human listeners have many capabilities a machine listening system should ideally have: to recognize a wide range of sounds; to segregate one sound source from a mixture of many sound sources; to judge complex attributes of sound such as rhythm and timbre (sound quality). Most human listeners take these abilities for granted, yet it has proved extremely difficult for conventional audio signal processing methods to tackle many of these tasks. Even currently successful tasks, such as automatic speech recognition, have typically led to very specialized techniques which cannot easily be applied to other domains. I propose to introduce new methods for machine listening of general audio scenes.As part of this work, I also will develop new interdisciplinary collaborations with both the machine vision and biological sensory research communities toinvestigate and develop general organizational principles for machine listening. One such principle that currently looks very promising is that of sparse representations. New theoretical advances and practical applications mean that sparse representations has recently emerged as a new and powerful analysis method, based on the principle that observations should be represented by only a few items chosen from a large number of possible items. This approach now has great potential for analysis and measurement of audio as well as other sensory signals. I also plan to use sparse representations to explore new biologically-inspired machine listening methods, and in turn to improve our understandingof biological hearing systems.Success in this research will open the way for new devices and systems able to process, identify and respond to a wide range of sounds, with diverse applications including: audio searching for the music and video industry; advances in hearing aids and cochlear implants; and incident detection for improved public safety on stations, roads and airports.

View the original record at the funder ↗

Researchers

Mark Plumbley (Principal Investigator)

Related Research

Grants with similar aims, by meaning.

Structured machine listening for soundscapes with multiple birds
Unifying audio signal processing and machine learning: a fundamental framework for machine hearing
AI for Sound
Environment and Listener Optimised Speech Processing for Hearing Enhancement in Real Situations (ELO-SPHERES)
MIMIC: Musically Intelligent Machines Interacting Creatively

Original classification

Fellowship

Plain English summaries and category classifications on this site are generated by AI and may not perfectly reflect the original research.