Skip to content
GlinttsdPublic

About

(DATE'26) CG-TP²U: Accelerating Equivariant Neural Networks with Clebsch–Gordan Tensor Product Processing Unit on FPGA

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

17 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 

Repository files navigation

CG-TP²U: Accelerating Equivariant Neural Networks with Clebsch–Gordan Tensor Product Processing Unit on FPGA

License Platform: FPGA

CG-TP²U is a software-hardware co-design framework developed to accelerate the Clebsch-Gordan tensor product (CGTP)]. CGTP is the primary computational bottleneck in Equivariant Neural Networks (ENNs), which are widely used for modeling 3D geometric data in physical and biological systems.


📸 Algorithm & Architecture Overview

challenges and solutions

Challenges and solutions for accelerating ENNs on FPGA .

Key Innovations

  • Sparse-Bypass Strategy (SBS): Exploits the inherent structural sparsity of CG coefficients (>80%). It uses a novel CG data format to pack overlapping non-zeros, bypassing redundant data accesses and computations.
  • Merged-Shift Quantization (MSQ): Enables full Int8 representation for irreps, weights, and CG coefficients. It replaces complex operations with hardware-friendly, shift-only dequantization.
  • Equicore Unit: A cascaded processing unit that tightly couples FPGA logic with RAM and DSP resources. It simplifies logic data paths to achieve a high operating frequency of 500 MHz.

🏗️ Hardware Architecture

The $TP^{2}U$ system consists of a host CPU, high-bandwidth memory (HBM), and the FPGA hardware accelerator.

Illustration of the Clebsch-Gordan tensor product (CGTP) computation flow

Illustration of the interaction between the CPU, HBM, instruction cache, and the parallel Equicore tiles.

Resource Utilization (AMD Virtex VCU128)

As reported after synthesis and implementation in Vivado 2024.1:

Resource Used Available Utilization
LUT 918,138 1,303,680 70.43%
FF 1,043,328 2,607,360 40.01%
BRAM 1,896.0 2,016 94.05%
URAM 864.0 960 90.0%
DSP 5,888 9,024 65.25%

🚀 Software Compilation & ISA

The host CPU partitions MIMO tasks into independent SISO tasks and generates 32-bit customized instructions to orchestrate the hardware.

Illustration of the Clebsch-Gordan tensor product (CGTP) computation flow

Illustration of the workflow of our customized software compilation scheme.

📊 Experimental Results

Performance & Efficiency

Speedup and energy efficiency comparison of Equicore with GPU-based works

Figure 15: Speedup and energy efficiency comparison of Equicore with GPU-based works.

📄 License

This project is licensed under the Apache License 2.0. See the LICENSE file for details.

✍️ Citation

If you use this work in your research, please cite our paper:

@article{tang2026tp2u,
  title={TP2U: Accelerating Equivariant Neural Networks with Tensor Product Processing Unit on FPGA},
  author={Tang, Shidi and Zhang, Chuanzhao and Chen, Ruiqi and Lv, Yuxuan and Silva, Bruno da and Ling, Ming},
  journal={IEEE},
  year={2026}
}

About

(DATE'26) CG-TP²U: Accelerating Equivariant Neural Networks with Clebsch–Gordan Tensor Product Processing Unit on FPGA

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages