Submit your papersSubmit Now
For Enquiries: [email protected]
IIARD LogoIIARD

From Cray to Exascale: A Critical Survey of Supercomputer Architecture Evolution, Heterogeneity, and Interconnect Bottlenecks

Oraye Godspower1, V.I.E Anireh1

Abstract

The evolution of supercomputer architecture has undergone several transformative phases since the 1960s, yet existing surveys have not adequately captured the engineering trade-offs that defined each generation. This paper presents a critical survey of supercomputer architecture from early vector machines to contemporary exascale systems, with particular emphasis on three persistent challenges: memory hierarchy design, interconnection network scalability, and thermal management. Drawing on foundational work by Cray and subsequent massively parallel systems, we examine how the transition from shared memory to distributed architectures enabled processor counts to grow from dozens to millions. The paper analyses recent exascale systems including Frontier, Fugaku, and LUMI, alongside emerging accelerator technologies such as GPU-based nodes, wafer-scale integration, and specialized interconnects including Slingshot, InfiniBand, and Omni-Path. Our survey identifies that while peak performance has grown exponentially, the gap between theoretical and sustained performance remains significant, largely due to interconnect latency and memory bandwidth limitations. We further evaluate contemporary cooling approaches from immersion to direct- to-chip liquid systems, noting that power density has emerged as the primary constraint on further scaling. Building upon earlier work by Oraye and Anireh (2022), this survey provides updated taxonomies and performance metrics that inform future exascale and post-exascale designs.

Keywords

exascale computingsupercomputer architectureheterogeneous computinghigh- speed interconnectsthermal managementTOP500

References

[1] Frisch, M. J. (1972). Remarks on algorithm 352 [S22], algorithm 385 [S13], algorithm 392. Communications of the ACM, 15(12), 1074. [2] Atchley, S., Zimmer, C., Lange, J., Bernholdt, D., Melesse Vergara, V., Beck, T., ... & Yeung, P. K. (2023, November). Frontier: exploring exascale. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (pp. 1-16). [3] Hill, M. D., Jouppi, N. P., & Sohi, G. (2000). Readings in computer architecture (pp. 40- 49). Morgan Kaufmann. [4] Hoffman, A. R. (1989). Supercomputers: directions in technology and applications (pp. 35- 47). National Academy Press. [5] Khan, A., Sim, H., Vazhkudai, S. S., Butt, A. R., & Kim, Y. (2021, January). An analysis of system balance and architectural trends based on top500 supercomputers. In The International Conference on High Performance Computing in Asia-Pacific Region (pp. 11-22). [6] Rajaraman, V. (2023). Frontier—world's first ExaFLOPS supercomputer. Resonance, 28(4), 567-576. [7] Silvano, C., Ielmini, D., Ferrandi, F., Fiorin, L., Curzel, S., Benini, L., ... & Perri, S. (2025). A survey on deep learning hardware accelerators for heterogeneous HPC platforms. ACM Computing Surveys, 57(11), 1-39. [8] Chang, J., Lu, K., Guo, Y., Wang, Y., Zhao, Z., Huang, L., ... & Zhang, B. (2024). A survey of compute nodes with 100 TFLOPS and beyond for supercomputers. CCF Transactions on High Performance Computing, 6(3), 243-262. [9] Godspower, O., & Anireh, V. I. E. (2022). Evolution of supercomputer architecture: A survey. International Journal of Computer Science and Mobile Applications, 10(7), 16- 26. [10] Adiga, N. R., Blumrich, M. A., Chen, D., Coteus, P., Gara, A., Giampapa, M. E., ... & Vranas, P. (2005). Blue Gene/L torus interconnection network. IBM Journal of Research and Development, 49(2.3), 265-276. [11] Matsuoka, S. (2021, June). Fugaku and A64FX: the first exascale supercomputer and its innovative ARM CPU. In 2021 Symposium on VLSI Circuits (pp. 1-3). IEEE. [12] Dongarra, J. (2020). Report on the Fujitsu Fugaku system. Technical Report ICL-UT-20- 06, University of Tennessee. [13] El-Rewini, H., & Abd-El-Barr, M. (2005). Advanced computer architecture and parallel processing (pp. 77-80). Wiley-Interscience. [14] Markidis, S., Laure, E., Markidis, S., & Laure, E. (2015). Solving software challenges for exascale. In International Conference on Exascale Applications and Software (pp. 3- 27). Springer. [15] Schönauer, W., & Häfner, H. (1994). Explaining the gap between theoretical peak performance and real performance for supercomputer architectures. Scientific Programming, 3(2), 157-168. [16] Rohr, D., Kalcher, S., Bach, M., Alaqeeliy, A. A., Alzaidy, H. M., Eschweiler, D., ... & Suliman, R. B. (2014, August). An energy-efficient multi-GPU supercomputer. In 2014 IEEE Intl Conf on High Performance Computing and Communications (pp. 42-45). IEEE. [17] Budiardja, R. D., Berrill, M., Eisenbach, M., Jansen, G. R., Joubert, W., Nichols, S., ... & Bronson Messer, O. E. (2023, May). Ready for the frontier: Preparing applications for the world's first exascale system. In International Conference on High Performance Computing (pp. 182-201). Springer. IJCSMT [18] Markomanolis, G. S., Alpay, A., Young, J., Klemm, M., Malaya, N., Esposito, A., ... & Bussmann, M. (2022, March). Evaluating GPU programming models for the LUMI supercomputer. In Asian Conference on Supercomputing Frontiers (pp. 79-101). Springer. [19] Lie, S. (2024). Inside the Cerebras wafer-scale cluster. IEEE Micro, 44(3), 49-57. [20] Kundu, Y., Kaur, M., Wig, T., Kumar, K., Kumari, P., Puri, V., & Arora, M. (2025). A comparison of the Cerebras wafer-scale integration technology with NVIDIA GPU- based systems for artificial intelligence. arXiv preprint arXiv:2503.11698. [21] Liang, Y., Galindo, A., & Yuan, C. (2024). High performance interconnect technologies for supercomputing. Technical Report, Los Alamos National Laboratory. [22] Benhari, A., Trystram, D., Dufossé, F., Denneulin, Y., & Desprez, F. (2024). Green HPC: An analysis of the domain based on Top500. arXiv preprint arXiv:2403.17466. [23] Sun, J., Gao, Z., Grant, D., Nawaz, K., Wang, P., Yang, C. M., ... & Huff, S. (2024). Energy dataset of Frontier supercomputer for waste heat recovery. Scientific Data, 11(1), 1077. [24] Domke, J., Matsuoka, S., Radanov, I., Tsushima, Y., Yuki, T., Nomura, A., ... & Dubé, N. (2019, August). The first supercomputer with HyperX topology: A viable alternative to fat-trees? In 2019 IEEE Symposium on High-Performance Interconnects (pp. 1-4). IEEE. [25] Zwinger, T., Heikonen, J., & Manninen, P. (2023). Lumi supercomputer for European researchers. Copernicus Meetings, GC11-solidearth-25. [26] Namashivayam, N. (2025). GPU-centric communication schemes for HPC and ML applications. arXiv preprint arXiv:2503.24230. [27] Lu, P. J., Lai, M. C., & Chang, J. S. (2022). A survey of high-performance interconnection networks in high-performance computer systems. Electronics, 11(9), 1369. [28] Zahn, F., Schoffer, A., & Froning, H. (2018, February). Evaluating energy-saving strategies on torus, k-Ary n-Tree, and dragonfly. In 2018 IEEE 4th International Workshop on High-Performance Interconnection Networks in the Exascale and Big- Data Era (HiPINEB) (pp. 16-23). IEEE. [29] Domke, J., Matsuoka, S., Ivanov, I. R., Tsushima, Y., Yuki, T., Nomura, A., ... & Dubé, N. (2019, November). HyperX topology: First at-scale implementation and comparison to the fat-tree. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (pp. 1-23). [30] Simonov, A., Ismagilov, T., Kazakov, D., Makagon, D., Polyakov, D., & Semenov, A. (2025). The Alpha high performance interconnect. In Russian Supercomputing Days (pp. 422-433). Springer. [31] Kamburugamuve, S., Ramasamy, K., Swany, M., & Fox, G. (2017, December). Low latency stream processing: Apache Heron with InfiniBand and Intel Omni-Path. In Proceedings of the 10th IEEE/ACM International Conference on Utility and Cloud Computing (pp. 101-110). [32] Birrittella, M. S., Debbage, M., Huggahalli, R., Kunz, J., Lovett, T., Rimmer, T., ... & Zak, R. C. (2015). Intel Omni-Path architecture technology overview. Intel Corporation, Technical Report. [33] De Sensi, D., Pichetti, L., Vella, F., De Matteis, T., Ren, Z., Fusco, L., ... & Hoefler, T. (2024, November). Exploring GPU-to-GPU communication: Insights into supercomputer interconnects. In SC24: International Conference for High Performance Computing, Networking, Storage and Analysis (pp. 1-15). IEEE. [34] Smith, A., Chapman, E., Patel, C., Swaminathan, R., Wuu, J., Huang, T., ... & Mangaser, R. (2024, February). 11.1 AMD Instinct MI300 series modular chiplet package—HPC and AI accelerator for exa-class systems. In 2024 IEEE International Solid-State Circuits Conference (Vol. 67, pp. 490-492). IEEE. IJCSMT [35] Shahi, P., Mathew, A., Saini, S., Bansode, P., Kasukurthy, R., & Agonafer, D. (2022). Assessment of reliability enhancement in high-power CPUs and GPUs using dynamic direct-to-chip liquid cooling. Journal of Enhanced Heat Transfer, 29(8). [36] Torii, S., & Ishikawa, H. (2017). ZettaScaler: Liquid immersion cooling manycore based supercomputer. CANDAR2017 Keynote, 3. [37] Ovaska, S. J., Dragseth, R. E., & Hanssen, S. A. (2016). Direct-to-chip liquid cooling for reducing power consumption in a subarctic supercomputer centre. International Journal of High Performance Computing and Networking, 9(3), 242-249. [38] Dlinnova, E., Biryukov, S., & Stegailov, V. (2020). Energy consumption of MD calculations on hybrid and CPU-only supercomputers with air and immersion cooling. In Parallel Computing: Technology Trends (pp. 574-582). IOS Press. [39] Angelelli, L. (2024, December). Power capping in supercomputing. In Job Scheduling Strategies for Parallel Processing: 27th International Workshop, JSSPP 2024 (p. 181). Springer. [40] Auweter, A., Bode, A., Brehm, M., Brochard, L., Hammer, N., Huber, H., ... & Wilde, T. (2014, June). A case study of energy aware scheduling on SuperMUC. In International Supercomputing Conference (pp. 394-409). Springer. [41] D'Amico, M., & Gonzalez, J. C. (2021). Energy hardware and workload aware job scheduling towards interconnected HPC environments. IEEE Transactions on Parallel and Distributed Systems. [42] Kramer, W. T. (2012, September). Top500 versus sustained performance: the top problems with the top500 list—and what to do about them. In Proceedings of the 21st International Conference on Parallel Architectures and Compilation Techniques (pp. 223-230). [43] Nikitenko, D., & Zheltkov, A. (2017, April). The Top50 list vivification in the evolution of HPC rankings. In International Conference on Parallel Computational Technologies (pp. 14-26). Springer. [44] Hockney, R. (1991). Performance parameters and benchmarking of supercomputers. Parallel Computing, 17(10-11), 1111-1130. [45] Marjanović, V., Gracia, J., & Glass, C. W. (2016, November). HPC benchmarking: Problem size matters. In 2016 7th International Workshop on Performance Modeling, Benchmarking and Simulation of High Performance Computer Systems (pp. 1- 10). IEEE. [46] Espino, D., Ares, G., Pedemonte, M., & Ezzatti, P. (2016, October). Overview of HPC benchmarks in hybrid hardware platforms (CPUs+GPUs). In 2016 XLII Latin American Computing Conference (pp. 1-10). IEEE. [47] Feng, W. C., & Cameron, K. (2007). The Green500 list: Encouraging sustainable supercomputing. Computer, 40(12), 50-55.

More Articles from INTERNATIONAL JOURNAL OF COMPUTER SCIENCE AND MATHEMATICAL THEORY

Advances in Algorithmic Contract Scoring for Pre-Negotiation Yield Optimization and Risk Retention

Author: Ngozi Samuel Uzougbo, Michael Ominyi, Cyril Chimelie Anichukwueze, Blessing, Chika Jones

DevTest flow: Designing a Scalable Continuous Testing Pipeline for High-Velocity Software Delivery

Author: Lawal Ahmed Oladimeji, Achori Busayo, Akeju BusayoZainab, Saka Samson, Damilare, Mbah Demian Chidi, Runsewe Similoluwa Mayowa, Oladiti Luqman, Abiodun