UNILOGIC-Reference-Cited by-同舟云学术

UNILOGIC

Published:2020-10 Issue:4 Volume:13 Page:1-32
ISSN:1936-7406
Container-title:ACM Transactions on Reconfigurable Technology and Systems
language:en
Short-container-title:ACM Trans. Reconfigurable Technol. Syst.

Author:

Ioannou Aggelos D.¹,Georgopoulos Konstantinos²,Malakonakis Pavlos³,Pnevmatikatos Dionisios N.⁴,Papaefstathiou Vassilis D.⁵,Papaefstathiou Ioannis⁶,Mavroidis Iakovos⁷

Affiliation:

1. School of ECE, Technical University of Crete, Greece, Telecommunication Systems Institute, Greece and Foundation for Research 8 Technology-Hellas (FORTH), Heraklion, Crete, Greece

2. Telecommunication Systems Institute, Chania, Greece

3. Telecommunication Systems Institute, Greece and Synelixis Solutions, Chalkida, Greece

4. Telecommunication Systems Institute, Greece and School of ECE, National Technical University of Athens, Athens, Greece

5. Chalmers University of Technology, Gothenburg, Sweden

6. School of ECE, Aristotle University of Thessaloniki, Thessaloniki, Greece

7. Telecommunication Systems Institute, Greece and Foundation for Research 8 Technology-Hellas (FORTH), Heraklion, Crete, Greece

Abstract

One of the main characteristics of High-performance Computing (HPC) applications is that they become increasingly performance and power demanding, pushing HPC systems to their limits. Existing HPC systems have not yet reached exascale performance mainly due to power limitations. Extrapolating from today’s top HPC systems, about 100–200 MWatts would be required to sustain an exaflop-level of performance. A promising solution for tackling power limitations is the deployment of energy-efficient reconfigurable resources (in the form of Field-programmable Gate Arrays (FPGAs)) tightly integrated with conventional CPUs. However, current FPGA tools and programming environments are optimized for accelerating a single application or even task on a single FPGA device. In this work, we present UNILOGIC (Unified Logic), a novel HPC-tailored parallel architecture that efficiently incorporates FPGAs. UNILOGIC adopts the Partitioned Global Address Space (PGAS) model and extends it to include hardware accelerators, i.e., tasks implemented on the reconfigurable resources. The main advantages of UNILOGIC are that (i) the hardware accelerators can be accessed directly by any processor in the system, and (ii) the hardware accelerators can access any memory location in the system. In this way, the proposed architecture offers a unified environment where all the reconfigurable resources can be seamlessly used by any processor/operating system. The UNILOGIC architecture also provides hardware virtualization of the reconfigurable logic so that the hardware accelerators can be shared among multiple applications or tasks. The FPGA layer of the architecture is implemented by splitting its reconfigurable resources into (i) a static partition, which provides the PGAS-related communication infrastructure, and (ii) fixed-size and dynamically reconfigurable slots that can be programmed and accessed independently or combined together to support both fine and coarse grain reconfiguration. 1 Finally, the UNILOGIC architecture has been evaluated on a custom prototype that consists of two 1U chassis, each of which includes eight interconnected daughter boards, called Quad-FPGA Daughter Boards (QFDBs); each QFDB supports four tightly coupled Xilinx Zynq Ultrascale+ MPSoCs as well as 64 Gigabytes of DDR4 memory, and thus, the prototype features a total of 64 Zynq MPSoCs and 1 Terabyte of memory. We tuned and evaluated the UNILOGIC prototype using both low-level (baremetal) performance tests, as well as two popular real-world HPC applications, one compute-intensive and one data-intensive. Our evaluation shows that UNILOGIC offers impressive performance that ranges from being 2.5 to 400 times faster and 46 to 300 times more energy efficient compared to conventional parallel systems utilizing only high-end CPUs, while it also outperforms GPUs by a factor ranging from 3 to 6 times in terms of time to solution, and from 10 to 20 times in terms of energy to solution.

Funder

European Commission under the H2020 Programme and the ECOSCALE project

Publisher

Association for Computing Machinery (ACM)

Subject

General Computer Science

Link

https://dl.acm.org/doi/pdf/10.1145/3409115

Reference77 articles.

1. AXI 2017. AXI Reference Guide. Retrieved from www.xilinx.com/support/documentation/ip_documentation/axi_ref_guide/latest/ug1037-vivado-axi-reference-guide.pdf. AXI 2017. AXI Reference Guide. Retrieved from www.xilinx.com/support/documentation/ip_documentation/axi_ref_guide/latest/ug1037-vivado-axi-reference-guide.pdf.

2. BittWare. 2019. BittWare FPGA Acceleration. Retrieved from https://www.bittware.com/. BittWare. 2019. BittWare FPGA Acceleration. Retrieved from https://www.bittware.com/.

3. Reconfigurable future for HPC

4. B. Brech J. Rubio and M. Hollinger. 2015. Data Engine for NoSQL-IBM Power Systems Edition. White Paper. B. Brech J. Rubio and M. Hollinger. 2015. Data Engine for NoSQL-IBM Power Systems Edition. White Paper.

5. HtComp: bringing reconfigurable hardware to future high-performance applications

Cited by 4 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. A Survey on FPGA-Based Heterogeneous Clusters Architectures;IEEE Access;2023

2. The Prism Bridge: Maximizing Inter-Chip AXI Throughput in the High-Speed Serial Era;IEEE Access;2023

3. Preconditioned Conjugate Gradient Acceleration on FPGA-Based Platforms;Electronics;2022-09-24

4. OmpSs@FPGA framework for high performance FPGA computing;IEEE Transactions on Computers;2021