Speech technology still sounds robotic because it cannot adapt to new speakers, noisy rooms, or changing contexts the way a human listener does. Current speech recognition and synthesis systems work well only in controlled conditions. A voice assistant that understands you in a quiet office may fail in a busy café. Text-to-speech can be perfectly intelligible yet sound flat and unnatural. The problem is that these systems lack the ability to learn about a speaker, the environment, or the situation as they go. Previous research has tackled these issues in isolation, leading to patchy progress. This programme aims to unify speech recognition and synthesis into a single theoretical framework that learns and adapts from real-world data without human supervision. If successful, the technology could approach human-like adaptability—handling unfamiliar speakers, new acoustic environments, and expressive communication. The impact would be felt in clinical settings, such as communication aids for people with motor neurone disease, as well as in everyday voice interfaces that feel genuinely natural. The project includes a user group with partners from industry and the NHS to test the systems on real tasks.
View original technical description
Humans are highly adaptable, and speech is our natural medium for informal communication. When communicating, we continuously adjust to other people, to the situation, and to the environment, using previously acquired knowledge to make this adaptation seem almost instantaneous. Humans generalise, enabling efficient communication in unfamiliar situations and rapid adaptation to new speakers or listeners. Current speech technology works well for certain controlled tasks and domains, but is far from natural, a consequence of its limited ability to acquire knowledge about people or situations, to adapt, and to generalise. This accounts for the uneasy public reaction to speech-driven systems. For example, text-to-speech synthesis can be as intelligible as human speech, but lacks expression and is not perceived as natural. Similarly, the accuracy of speech recognition systems can collapse if the acoustic environment or task domain changes, conditions which a human listener would handle easily. Research approaches to these problems have hitherto been piecemeal and as a result progress has been patchy. In contrast NST will focus on the integrated theoretical development of new joint models for speech recognition and synthesis. These models will allow us to incorporate knowledge about the speakers, the environment, the communication context and awareness of the task, and will learn and adapt from real world data in an online, unsupervised manner. This theoretical unification is already underway within the NST labs and, combined with our record of turning theory into practical state-of-the-art applications, will enable us to bring a naturalness to speech technology that is not currently attainable.The NST programme will yield technology which (1) approaches human adaptability to new communication situations, (2) is capable of personalised communication, and (3) takes account of speaker intention and expressiveness in speech recognition and synthesis. This is an ambitious vision. Its success will be measured in terms of how the theoretical development reshapes the field over the next decade, the takeup of the software systems that we shall develop, and through the impact of our exemplar interactive applications.We shall establish a strong User Group to maximise the impact of the project, with a members concerned with clinical applications, as well as more general speech technology. Members of the User Group include Toshiba, EADS Innovation Works, Cisco, Barnsley Hospital NHS Foundation Trust, and the Euan MacDonald Centre for MND Research. An important interaction with the User Group will be validating our systems on their data and tasks, discussed at an annual user workshop.
Plain English summaries and category classifications on this site are generated by AI and may not perfectly reflect the original research.
Is something wrong? Let us know