A machine learns to make a sequence of decisions by trial and error, without a human telling it what to do at each step. This is deep reinforcement learning (DRL), and it already powers systems that play games, control robots, and optimise industrial processes. But most real-world settings involve multiple decision-makers sharing resources—think of warehouse robots coordinating to pack orders, or autonomous vehicles navigating the same junction. Current DRL methods struggle when many agents must cooperate under realistic constraints, such as partial information and limited communication. This research programme extends DRL to multi-agent systems (MADRL), enabling groups of autonomous agents to learn collaborative behaviours without human intervention. If successful, it could transform warehouse management, assembly lines, and clinical decision-support systems—anywhere that multiple intelligent systems must work together efficiently. The work is applied and industry-facing, with clear commercial interest. It does not aim for artificial general intelligence, but for practical, task-oriented intelligence that could quietly improve the logistics and manufacturing systems that keep supply chains running.
View original technical description
Despite being far from having reached 'artificial general intelligence' - the broad and deep capability for a machine to comprehend our surroundings - progress has been made in the last few years towards a more specialised AI: the ability to effectively address well-defined, specific goals in a given environment, which is the kind of task-oriented intelligence that is part of many human jobs. Much of this progress has been enabled by deep reinforcement learning (DRL), one of the most promising and fast-growing areas within machine learning. In DRL, an autonomous decision maker - the "agent" - learns how to make optimal decisions that will eventually lead to reaching a final goal. DRL holds the promise of enabling autonomous systems to learn large repertoires of collaborative and adaptive behavioural skills without human intervention, with application in a range of settings from simple games to industrial process automation to modelling human learning and cognition. Many real-world applications are characterised by the interplay of multiple decision-makers that operate in the same shared-resources environment and need to accomplish goals cooperatively. For instance, some of the most advanced industrial multi-agent systems in the world today are assembly lines and warehouse management systems. Whether the agents are robots, autonomous vehicles or clinical decision-makers, there is a strong desire for and increasing commercial interest in these systems: they are attractive because they can operate on their own in the world, alongside humans, under realistic constraints (e.g. guided by only partial information and with limited communication bandwidth). This research programme will extend the DRL methodology to systems comprising of many interacting agents that must cooperatively achieve a common goal: multi-agent DRL, or MADRL.
Plain English summaries and category classifications on this site are generated by AI and may not perfectly reflect the original research.
Is something wrong? Let us know