Completed Computing & AI Materials & Manufacturing

Continuous on-line adaptation in many-core systems: From graceful degradation to graceful amelioration

In plain English

AI plain-English summary

Computer chips with hundreds or thousands of processors are on the horizon, but they currently cannot manage their own heat, power use, or inevitable component failures without human intervention. This project tackles a fundamental problem: as chips pack in more cores, they become too complex to manage centrally, and existing architectures cannot adapt on their own to faults or energy spikes. The research proposes two self-regulating mechanisms. *Graceful degradation* would let a many-core chip automatically slow down to avoid overheating or to cope with a failed processor. *Graceful amelioration* would let the same chip continuously learn better ways to run software, improving its own performance and energy efficiency over time without external updates. If successful, this work could enable the commercial production of many-core processors—the kind needed to sustain future computing demands in data centres, energy grids, and communications infrastructure. It would also address a growing practical bottleneck: the rising cost and complexity of designing chips that must be perfect from the start. The project is applied fundamental science, directly targeting the reliability and efficiency limits that currently block the next generation of parallel computing.

View original technical description
Until recently, the ever-increasing demand of computing power has been met on one hand by increasing the operating frequency of processors and on the other by designing more and more complex processors capable of executing more than one instruction at the same time. However, both these approaches seem to be reaching (or possibly have already reached) their practical limits, mainly due to issues related to design complexity and cost-effectiveness. The current trend in computer design seems to favour a shift to systems where computational power is achieved not by a single very fast and very complex processor, but through the parallel operation of several on-chip processors, each executing a single thread. This kind of approach is implemented commercially today through multi-core processors and in research through the Network On Chip (NoC) or the Chip Multi-Processors (CMP) paradigms. The natural evolution of these approaches sees the number of cores increasing constantly and it is generally accepted that the next few decades will witness the introduction of many-core systems, that is, systems that integrate hundreds or thousands of cores. This shift introduces problems common to all massively parallel systems, ranging from the design of applications that can exploit large numbers of processors to technological challenges related to the implementation of such cores in silicon substrates that are increasingly error-prone, due to their size and to the increasing sensitivity to faults of next-generation technologies, and to the dissipation of heat generated by the computational activity in the cores. Current architectures are not suitable for this kind of systems and there is a strong need to devise novel mechanisms and technologies that will allow the development of many-core systems and eventually their commercialization as consumer products. Imagine then a many-core system with thousands or millions of processors that gets better and better with time at executing an application, "gracefully" providing optimal power usage while maximizing performance levels and tolerating component failures. The proposed project aims at investigating how such mechanisms can represent crucial enabling technologies for many-core systems. Specifically, this project focuses on how to overcome three critical issues related to the implementation of many-core systems: reliability, energy efficiency, and on-line optimisation. The need for reliability is an accepted challenge for many-core systems, considering the large number of components and the increasing likelihood of faults of next-generation technologies, as is the requirement to reduce the heat dissipation related to energy consumption. On the other hand, on-line optimisation, that is, the ability of the system to improve over time without the need for external intervention (including becoming better at reliability and energy efficiency), is a mechanism that could be vital to enable the implementation of these properties in systems that cannot be managed centrally due to the vast number of cores involved. The proposed approach is centred around two basic processes: Graceful degradation implies that the system will be able to cope with faults (permanent or temporary) or potentially damaging power consumption peaks by lowering its performance. Graceful amelioration implies that the system will constantly seek for alternative, better ways to execute an application.

View the original record at the funder ↗

Researchers

Andy Tyrrell (Co-Investigator)Bashir M. Al-Hashimi (Co-Investigator)Geoff Merrett (Co-Investigator)Gianluca Tempesti (Principal Investigator)Martin Trefzer (Co-Investigator)Stephen Furber (Co-Investigator)

Related Research

Grants with similar aims, by meaning.

M3: Managing Many-Cores for the Masses
Compiling for Energy Efficiency in Multicore Memory Hierarchies
DOME: Delaying and Overcoming Microprocessor Errors
Concurrent Proof of Concept
Compiler and Architectural Support for Ultra-Fine-Grained Parallelism

Original classification

Research Grant

Plain English summaries and category classifications on this site are generated by AI and may not perfectly reflect the original research.