References
[1] Frisch, M. J. (1972). Remarks on algorithm 352 [S22], algorithm 385 [S13], algorithm 392. Communications of the ACM, 15(12), 1074. [2] Atchley, S., Zimmer, C., Lange, J., Bernholdt, D., Melesse Vergara, V., Beck, T., ... & Yeung, P. K. (2023, November). Frontier: exploring exascale. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (pp. 1-16). [3] Hill, M. D., Jouppi, N. P., & Sohi, G. (2000). Readings in computer architecture (pp. 40- 49). Morgan Kaufmann. [4] Hoffman, A. R. (1989). Supercomputers: directions in technology and applications (pp. 35- 47). National Academy Press. [5] Khan, A., Sim, H., Vazhkudai, S. S., Butt, A. R., & Kim, Y. (2021, January). An analysis of system balance and architectural trends based on top500 supercomputers. In The International Conference on High Performance Computing in Asia-Pacific Region (pp. 11-22). [6] Rajaraman, V. (2023). Frontier—world's first ExaFLOPS supercomputer. Resonance, 28(4), 567-576. [7] Silvano, C., Ielmini, D., Ferrandi, F., Fiorin, L., Curzel, S., Benini, L., ... & Perri, S. (2025). A survey on deep learning hardware accelerators for heterogeneous HPC platforms. ACM Computing Surveys, 57(11), 1-39. [8] Chang, J., Lu, K., Guo, Y., Wang, Y., Zhao, Z., Huang, L., ... & Zhang, B. (2024). A survey of compute nodes with 100 TFLOPS and beyond for supercomputers. CCF Transactions on High Performance Computing, 6(3), 243-262. [9] Godspower, O., & Anireh, V. I. E. (2022). Evolution of supercomputer architecture: A survey. International Journal of Computer Science and Mobile Applications, 10(7), 16- 26. [10] Adiga, N. R., Blumrich, M. A., Chen, D., Coteus, P., Gara, A., Giampapa, M. E., ... & Vranas, P. (2005). Blue Gene/L torus interconnection network. IBM Journal of Research and Development, 49(2.3), 265-276. [11] Matsuoka, S. (2021, June). Fugaku and A64FX: the first exascale supercomputer and its innovative ARM CPU. In 2021 Symposium on VLSI Circuits (pp. 1-3). IEEE. [12] Dongarra, J. (2020). Report on the Fujitsu Fugaku system. Technical Report ICL-UT-20- 06, University of Tennessee. [13] El-Rewini, H., & Abd-El-Barr, M. (2005). Advanced computer architecture and parallel processing (pp. 77-80). Wiley-Interscience. [14] Markidis, S., Laure, E., Markidis, S., & Laure, E. (2015). Solving software challenges for exascale. In International Conference on Exascale Applications and Software (pp. 3- 27). Springer. [15] Schönauer, W., & Häfner, H. (1994). Explaining the gap between theoretical peak performance and real performance for supercomputer architectures. Scientific Programming, 3(2), 157-168. [16] Rohr, D., Kalcher, S., Bach, M., Alaqeeliy, A. A., Alzaidy, H. M., Eschweiler, D., ... & Suliman, R. B. (2014, August). An energy-efficient multi-GPU supercomputer. In 2014 IEEE Intl Conf on High Performance Computing and Communications (pp. 42-45). IEEE. [17] Budiardja, R. D., Berrill, M., Eisenbach, M., Jansen, G. R., Joubert, W., Nichols, S., ... & Bronson Messer, O. E. (2023, May). Ready for the frontier: Preparing applications for the world's first exascale system. In International Conference on High Performance Computing (pp. 182-201). Springer. IJCSMT [18] Markomanolis, G. S., Alpay, A., Young, J., Klemm, M., Malaya, N., Esposito, A., ... & Bussmann, M. (2022, March). Evaluating GPU programming models for the LUMI supercomputer. In Asian Conference on Supercomputing Frontiers (pp. 79-101). Springer. [19] Lie, S. (2024). Inside the Cerebras wafer-scale cluster. IEEE Micro, 44(3), 49-57. [20] Kundu, Y., Kaur, M., Wig, T., Kumar, K., Kumari, P., Puri, V., & Arora, M. (2025). A comparison of the Cerebras wafer-scale integration technology with NVIDIA GPU- based systems for artificial intelligence. arXiv preprint arXiv:2503.11698. [21] Liang, Y., Galindo, A., & Yuan, C. (2024). High performance interconnect technologies for supercomputing. Technical Report, Los Alamos National Laboratory. [22] Benhari, A., Trystram, D., Dufossé, F., Denneulin, Y., & Desprez, F. (2024). Green HPC: An analysis of the domain based on Top500. arXiv preprint arXiv:2403.17466. [23] Sun, J., Gao, Z., Grant, D., Nawaz, K., Wang, P., Yang, C. M., ... & Huff, S. (2024). Energy dataset of Frontier supercomputer for waste heat recovery. Scientific Data, 11(1), 1077. [24] Domke, J., Matsuoka, S., Radanov, I., Tsushima, Y., Yuki, T., Nomura, A., ... & Dubé, N. (2019, August). The first supercomputer with HyperX topology: A viable alternative to fat-trees? In 2019 IEEE Symposium on High-Performance Interconnects (pp. 1-4). IEEE. [25] Zwinger, T., Heikonen, J., & Manninen, P. (2023). Lumi supercomputer for European researchers. Copernicus Meetings, GC11-solidearth-25. [26] Namashivayam, N. (2025). GPU-centric communication schemes for HPC and ML applications. arXiv preprint arXiv:2503.24230. [27] Lu, P. J., Lai, M. C., & Chang, J. S. (2022). A survey of high-performance interconnection networks in high-performance computer systems. Electronics, 11(9), 1369. [28] Zahn, F., Schoffer, A., & Froning, H. (2018, February). Evaluating energy-saving strategies on torus, k-Ary n-Tree, and dragonfly. In 2018 IEEE 4th International Workshop on High-Performance Interconnection Networks in the Exascale and Big- Data Era (HiPINEB) (pp. 16-23). IEEE. [29] Domke, J., Matsuoka, S., Ivanov, I. R., Tsushima, Y., Yuki, T., Nomura, A., ... & Dubé, N. (2019, November). HyperX topology: First at-scale implementation and comparison to the fat-tree. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (pp. 1-23). [30] Simonov, A., Ismagilov, T., Kazakov, D., Makagon, D., Polyakov, D., & Semenov, A. (2025). The Alpha high performance interconnect. In Russian Supercomputing Days (pp. 422-433). Springer. [31] Kamburugamuve, S., Ramasamy, K., Swany, M., & Fox, G. (2017, December). Low latency stream processing: Apache Heron with InfiniBand and Intel Omni-Path. In Proceedings of the 10th IEEE/ACM International Conference on Utility and Cloud Computing (pp. 101-110). [32] Birrittella, M. S., Debbage, M., Huggahalli, R., Kunz, J., Lovett, T., Rimmer, T., ... & Zak, R. C. (2015). Intel Omni-Path architecture technology overview. Intel Corporation, Technical Report. [33] De Sensi, D., Pichetti, L., Vella, F., De Matteis, T., Ren, Z., Fusco, L., ... & Hoefler, T. (2024, November). Exploring GPU-to-GPU communication: Insights into supercomputer interconnects. In SC24: International Conference for High Performance Computing, Networking, Storage and Analysis (pp. 1-15). IEEE. [34] Smith, A., Chapman, E., Patel, C., Swaminathan, R., Wuu, J., Huang, T., ... & Mangaser, R. (2024, February). 11.1 AMD Instinct MI300 series modular chiplet package—HPC and AI accelerator for exa-class systems. In 2024 IEEE International Solid-State Circuits Conference (Vol. 67, pp. 490-492). IEEE. IJCSMT [35] Shahi, P., Mathew, A., Saini, S., Bansode, P., Kasukurthy, R., & Agonafer, D. (2022). Assessment of reliability enhancement in high-power CPUs and GPUs using dynamic direct-to-chip liquid cooling. Journal of Enhanced Heat Transfer, 29(8). [36] Torii, S., & Ishikawa, H. (2017). ZettaScaler: Liquid immersion cooling manycore based supercomputer. CANDAR2017 Keynote, 3. [37] Ovaska, S. J., Dragseth, R. E., & Hanssen, S. A. (2016). Direct-to-chip liquid cooling for reducing power consumption in a subarctic supercomputer centre. International Journal of High Performance Computing and Networking, 9(3), 242-249. [38] Dlinnova, E., Biryukov, S., & Stegailov, V. (2020). Energy consumption of MD calculations on hybrid and CPU-only supercomputers with air and immersion cooling. In Parallel Computing: Technology Trends (pp. 574-582). IOS Press. [39] Angelelli, L. (2024, December). Power capping in supercomputing. In Job Scheduling Strategies for Parallel Processing: 27th International Workshop, JSSPP 2024 (p. 181). Springer. [40] Auweter, A., Bode, A., Brehm, M., Brochard, L., Hammer, N., Huber, H., ... & Wilde, T. (2014, June). A case study of energy aware scheduling on SuperMUC. In International Supercomputing Conference (pp. 394-409). Springer. [41] D'Amico, M., & Gonzalez, J. C. (2021). Energy hardware and workload aware job scheduling towards interconnected HPC environments. IEEE Transactions on Parallel and Distributed Systems. [42] Kramer, W. T. (2012, September). Top500 versus sustained performance: the top problems with the top500 list—and what to do about them. In Proceedings of the 21st International Conference on Parallel Architectures and Compilation Techniques (pp. 223-230). [43] Nikitenko, D., & Zheltkov, A. (2017, April). The Top50 list vivification in the evolution of HPC rankings. In International Conference on Parallel Computational Technologies (pp. 14-26). Springer. [44] Hockney, R. (1991). Performance parameters and benchmarking of supercomputers. Parallel Computing, 17(10-11), 1111-1130. [45] Marjanović, V., Gracia, J., & Glass, C. W. (2016, November). HPC benchmarking: Problem size matters. In 2016 7th International Workshop on Performance Modeling, Benchmarking and Simulation of High Performance Computer Systems (pp. 1- 10). IEEE. [46] Espino, D., Ares, G., Pedemonte, M., & Ezzatti, P. (2016, October). Overview of HPC benchmarks in hybrid hardware platforms (CPUs+GPUs). In 2016 XLII Latin American Computing Conference (pp. 1-10). IEEE. [47] Feng, W. C., & Cameron, K. (2007). The Green500 list: Encouraging sustainable supercomputing. Computer, 40(12), 50-55.