Listening to a live orchestra or a football match at home currently requires a specific, carefully arranged set of speakers to create the illusion of being there. This project aims to free 3D sound from that controlled environment. The core problem is that existing audio formats—stereo or even cinema-grade 5.1 and 22.2 systems—are “channel-based,” meaning they demand a fixed number of speakers in precise positions. That setup is impractical for most living rooms, let alone for listening on the move. S3A will replace that rigid approach with an “object-based” method. Instead of mixing sound into fixed channels, it will treat each sound source—a singer’s voice, a crowd’s roar, a kicked football—as a separate object with its own spatial properties. The system will then adapt the reproduction in real time to whatever loudspeakers are available, the room’s acoustics, and where the listener is sitting. This requires major advances in how we perceive spatial audio and in the signal processing that translates that perception into sound. If successful, this could transform home entertainment and mobile listening, delivering genuinely immersive experiences without specialised hardware. It also has implications for broadcast, internet streaming, and digital devices, making 3D audio a practical, platform-independent feature rather than a luxury confined to cinemas.
View original technical description
3D sound can offer listeners the experience of "being there" at a live event, such as the Proms or Olympic 100m, but currently requires highly controlled listening spaces and loudspeaker setups. The goal of S3A is to realise practical 3D audio for the general public to enable immersive experiences at home or on the move. Virtually the whole of the UK population consume audio. S3A aims to unlock the creative potential of 3D sound and deliver to listeners a step change in immersive experiences. This requires a radical new listener centred approach to audio enabling 3D sound production to dynamically adapt to the listeners' environment. Achieving immersive audio experiences in uncontrolled living spaces presents a significant research challenge. This requires major advances in our understanding of the perception of spatial audio together with new representations of audio and the signal processing that allows content creation and perceptually accurate reproduction. Existing audio production formats (stereo, 5.1) and those proposed for future cinema spatial audio (24,128) are channel-based requiring specific controlled loudspeaker arrangements that are simply not practical for the majority of home listeners. S3A will pioneer a novel object-based methodology for audio signal processing that allows flexible production and reproduction in real spaces. The reproduction will be adaptive to loudspeaker configuration, room acoustics and listener locations. The fields of audio and visual 3D scene understanding will be brought together to identify and model audio-visual objects in complex real scenes. Audio-visual objects are sound sources or events with known spatial properties of shape and location over time, e.g. a football being kicked, a musical instrument being played or the crowd chanting at a football match. Object based representation will transform audio production from existing channel based signal mixing (stereo, 5.1, 22.2) to spatial control of isolated sound sources and events. This will realise the creative potential of 3D sound enabling intelligent user-centred content production, transmission and reproduction of 3D audio content in platform independent formats. Object-based audio will allow flexible delivery (broadcast, IP and mobile) and adaptive reproduction of 3D sound to existing and new digital devices.
Plain English summaries and category classifications on this site are generated by AI and may not perfectly reflect the original research.
Is something wrong? Let us know