CG-TP²U: Accelerating Equivariant Neural Networks with Clebsch–Gordan Tensor Product Processing Unit on FPGA
CG-TP²U is a software-hardware co-design framework developed to accelerate the Clebsch-Gordan tensor product (CGTP)]. CGTP is the primary computational bottleneck in Equivariant Neural Networks (ENNs), which are widely used for modeling 3D geometric data in physical and biological systems.
Challenges and solutions for accelerating ENNs on FPGA .
- Sparse-Bypass Strategy (SBS): Exploits the inherent structural sparsity of CG coefficients (>80%). It uses a novel CG data format to pack overlapping non-zeros, bypassing redundant data accesses and computations.
- Merged-Shift Quantization (MSQ): Enables full Int8 representation for irreps, weights, and CG coefficients. It replaces complex operations with hardware-friendly, shift-only dequantization.
- Equicore Unit: A cascaded processing unit that tightly couples FPGA logic with RAM and DSP resources. It simplifies logic data paths to achieve a high operating frequency of 500 MHz.
The
Illustration of the interaction between the CPU, HBM, instruction cache, and the parallel Equicore tiles.
As reported after synthesis and implementation in Vivado 2024.1:
| Resource | Used | Available | Utilization |
|---|---|---|---|
| LUT | 918,138 | 1,303,680 | 70.43% |
| FF | 1,043,328 | 2,607,360 | 40.01% |
| BRAM | 1,896.0 | 2,016 | 94.05% |
| URAM | 864.0 | 960 | 90.0% |
| DSP | 5,888 | 9,024 | 65.25% |
The host CPU partitions MIMO tasks into independent SISO tasks and generates 32-bit customized instructions to orchestrate the hardware.
Illustration of the workflow of our customized software compilation scheme.
Figure 15: Speedup and energy efficiency comparison of Equicore with GPU-based works.
This project is licensed under the Apache License 2.0. See the LICENSE file for details.
If you use this work in your research, please cite our paper:
@article{tang2026tp2u,
title={TP2U: Accelerating Equivariant Neural Networks with Tensor Product Processing Unit on FPGA},
author={Tang, Shidi and Zhang, Chuanzhao and Chen, Ruiqi and Lv, Yuxuan and Silva, Bruno da and Ling, Ming},
journal={IEEE},
year={2026}
}


