Completed Computing & AI Physics & Astronomy

GW4 Tier 2 HPC Centre for Advanced Architectures

In plain English

AI plain-English summary

A consortium of four UK universities is building one of the world's first supercomputers powered by ARM64 server chips, a radical departure from the Intel and AMD processors that dominate high-performance computing (HPC). The problem is that most scientific supercomputers are designed to maximise raw calculation speed (peak FLOP/s), but many real-world codes—weather simulations, climate models, fluid dynamics—are bottlenecked by how fast data can be moved between memory and the processor, not by how fast the processor can crunch numbers. This machine, based on Broadcom's Vulcan chip, deliberately trades peak FLOP/s for much greater memory bandwidth. The UK's HPC community currently has no way to test whether this trade-off actually improves performance for real scientific workloads. If the approach works, it could reshape how the UK procures its national supercomputing infrastructure—from Tier 1 national facilities down to university-scale Tier 3 systems. The immediate impact is on research efficiency: faster simulations of climate, materials, and drug interactions. The consortium will also port and optimise scientific codes for the new architecture, ensuring rigorous comparisons are possible. The system will run existing codes "out of the box" using standard parallel programming languages, so demand is expected to be strong. Operating costs are covered by the consortium and partners.

View original technical description
This proposal by a consortium of the GW4 Alliance of Bristol, Bath, Cardiff and Exeter, in partnership with Cray and the Met Office, is to provide a national 64-bit ARM-based HPC service. The system will be one of the world's first to be based on Broadcom's Vulcan server-class chip. Details of this device are still under NDA, but the Vulcan CPU is generating excitement because it trades off much greater provision of memory bandwidth for less emphasis on peak FLOP/s, the former being more important for most scientific codes. Providing access to such a machine as a national service should therefore enable the UK's HPC community to quantify the benefit of memory bandwidth focused CPUs, thus informing future system procurements from Tier 1 to Tier 3. If this greater focus on memory bandwidth does, as expected, result in greater performance and science throughput, then ARM64-based machines, such as the Cray XC Scout system that we are proposing in this bid, will be genuine contenders for Tier 1 and Tier 3 production systems from 2017. In addition to our goal of providing one of the world's first ARM64 production HPC systems, this proposal will also provide a service to enable algorithm development and the porting and optimisation of scientific codes in readiness for ARM64 machines. This algorithm and software effort is a crucial part of any architectural evaluation, as rigorous architecture-to-architecture comparisons are only possible when optimisation levels across the architectures are similar. There is already tremendous interest in evaluating ARM64 within the HPC community, with multiple ARM-based HPC projects underway around the world. Our proposed machine will be able to run most existing codes "out of the box", supporting the most common parallel programming languages, including OpenMP and MPI. Thus most users should be able to begin to evaluate the service with minimal effort, and so we expect demand to be strong. The system will be run as a national facility, with open calls for computing time allocated via a lightweight resource allocation process. A top-level Consortium Management Board will determine the policy for resource allocation between the different application areas as well as fundamental computational science research into next generation parallel algorithms. Operating expenses will be covered by the consortium and its partners. Systems administrator and power costs will be split across the partners, while a group of expert research software engineers will help the community develop new algorithms, port codes and rigorously evaluate this important new architecture.

View the original record at the funder ↗

Researchers

Beth Wingate (Co-Investigator)James Davenport (Co-Investigator)Jonathan Fieldsend (Co-Investigator)Mark Parsons (Co-Investigator)Ozgur Akman (Co-Investigator)Paul Calleja (Co-Investigator)Roger Whitaker (Co-Investigator)Simon McIntosh-Smith (Principal Investigator)

Related Research

Grants with similar aims, by meaning.

GW4 Tier-2 HPC Centre for Advanced Architectures
The GW4 Isambard Tier-2 service for advanced computer architectures
Enabling PETSc on the Cerebras Wafer Scale Engine
Baskerville: a national accelerated compute resource
Excalibur H&ES - ARM-GPU testbed & ARM Forge Licence

Original classification

Research Grant

Plain English summaries and category classifications on this site are generated by AI and may not perfectly reflect the original research.