Skip to content

GPU coordnum - #940

Open
HanatoK wants to merge 26 commits into
Colvars:masterfrom
HanatoK:coordnum-template-gpu-2
Open

GPU coordnum#940
HanatoK wants to merge 26 commits into
Colvars:masterfrom
HanatoK:coordnum-template-gpu-2

Conversation

@HanatoK

@HanatoK HanatoK commented Jul 15, 2026

Copy link
Copy Markdown
Member

This PR should wait for #919 and #938. #919 is necessary for wrapping the distances using the internal PBC function on GPU. This PR is independent from #926.

Regarding the pairlist implementation, the GPU kernels in this PR try to match the CPU pairlist, which is slow for the time being. A better option is to implement at least a NAMD-style GPU pairlist, but since Colvars (i) does not allow atoms in an atom group to be reordered, and (ii) uses only a single cutoff for both the switching function and the pairlist, it is impossible or very diffcult to do so.

@HanatoK HanatoK self-assigned this Jul 15, 2026
HanatoK and others added 14 commits July 28, 2026 09:29
According to the CUDA API documentation, this could improve the
performance.
Keep in mind that not every CVC use CUDA graphs for the GPU
implementation. Consequently, whether to use CUDA graphs or not is
considered to be an implementation detail. This commit strips the CUDA
graphs in colvar::cvc to a separate class to colvar_gpu_calc.*. The RMSD
GPU code is modified to use it.
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Copilot Autofix powered by AI
<175728472+Copilot@users.noreply.github.com>
Co-authored-by: Copilot Autofix powered by AI
<175728472+Copilot@users.noreply.github.com>
Change inconsistent outputFreq parameter
…ault_block_size

"constexpr static" is used for muting the compiler warnings.
@HanatoK
HanatoK force-pushed the coordnum-template-gpu-2 branch from a923b17 to 8eac9a6 Compare July 28, 2026 18:11
@HanatoK
HanatoK force-pushed the coordnum-template-gpu-2 branch from 8eac9a6 to 6c72b36 Compare July 28, 2026 19:09
@HanatoK
HanatoK marked this pull request as ready for review July 28, 2026 19:46
@HanatoK
HanatoK requested review from giacomofiorin and jhenin July 28, 2026 19:48
@HanatoK

HanatoK commented Jul 28, 2026

Copy link
Copy Markdown
Member Author

Initial performance benchmark of computing selfcoordnum of 718 atoms in NAMD:

  • CPU: 63.2372 ns/day
  • CPU without cvm::error_static in cvm::system_boundary_conditions::position_distance: 77.2318 ns/day
  • GPU: 130.812 ns/day

Nsys shows that the GPU kernel is still very computationally expensive as its execution time is multiple times of the NB kernel:

colvars_coordnum_gpu

I would prefer optimizing the overall calculations in a new PR or #926.

Test Colvars configuration:

colvarsTrajFrequency 10
colvarsRestartFrequency 1000
smp gpu

colvar {
  name cv
  outputAppliedForce on
  selfCoordNum {
    group1 {
      atomNumbers {23 26 29 32 35 38 41 44 47 50 53 56 59 62 65 68 71 74 77 80 83 86 89 92 95 98 101 104 107 110 113 116 119 122 125 128 131 134 137 140 143 146 149 152 155 158 161 164 167 170 173 176 179 182 185 188 191 194 197 200 203 206 209 212 215 218 221 224 227 230 233 236 239 242 245 248 251 254 257 260 263 266 269 272 275 278 281 284 287 290 293 296 299 302 305 308 311 314 317 320 323 326 329 332 335 338 341 344 347 350 353 356 359 362 365 368 371 374 377 380 383 386 389 392 395 398 401 404 407 410 413 416 419 422 425 428 431 434 437 440 443 446 449 452 455 458 461 464 467 470 473 476 479 482 485 488 491 494 497 500 503 506 509 512 515 518 521 524 527 530 533 536 539 542 545 548 551 554 557 560 563 566 569 572 575 578 581 584 587 590 593 596 599 602 605 608 611 614 617 620 623 626 629 632 635 638 641 644 647 650 653 656 659 662 665 668 671 674 677 680 683 686 689 692 695 698 701 704 707 710 713 716 719 722 725 728 731 734 737 740 743 746 749 752 755 758 761 764 767 770 773 776 779 782 785 788 791 794 797 800 803 806 809 812 815 818 821 824 827 830 833 836 839 842 845 848 851 854 857 860 863 866 869 872 875 878 881 884 887 890 893 896 899 902 905 908 911 914 917 920 923 926 929 932 935 938 941 944 947 950 953 956 959 962 965 968 971 974 977 980 983 986 989 992 995 998 1001 1004 1007 1010 1013 1016 1019 1022 1025 1028 1031 1034 1037 1040 1043 1046 1049 1052 1055 1058 1061 1064 1067 1070 1073 1076 1079 1082 1085 1088 1091 1094 1097 1100 1103 1106 1109 1112 1115 1118 1121 1124 1127 1130 1133 1136 1139 1142 1145 1148 1151 1154 1157 1160 1163 1166 1169 1172 1175 1178 1181 1184 1187 1190 1193 1196 1199 1202 1205 1208 1211 1214 1217 1220 1223 1226 1229 1232 1235 1238 1241 1244 1247 1250 1253 1256 1259 1262 1265 1268 1271 1274 1277 1280 1283 1286 1289 1292 1295 1298 1301 1304 1307 1310 1313 1316 1319 1322 1325 1328 1331 1334 1337 1340 1343 1346 1349 1352 1355 1358 1361 1364 1367 1370 1373 1376 1379 1382 1385 1388 1391 1394 1397 1400 1403 1406 1409 1412 1415 1418 1421 1424 1427 1430 1433 1436 1439 1442 1445 1448 1451 1454 1457 1460 1463 1466 1469 1472 1475 1478 1481 1484 1487 1490 1493 1496 1499 1502 1505 1508 1511 1514 1517 1520 1523 1526 1529 1532 1535 1538 1541 1544 1547 1550 1553 1556 1559 1562 1565 1568 1571 1574 1577 1580 1583 1586 1589 1592 1595 1598 1601 1604 1607 1610 1613 1616 1619 1622 1625 1628 1631 1634 1637 1640 1643 1646 1649 1652 1655 1658 1661 1664 1667 1670 1673 1676 1679 1682 1685 1688 1691 1694 1697 1700 1703 1706 1709 1712 1715 1718 1721 1724 1727 1730 1733 1736 1739 1742 1745 1748 1751 1754 1757 1760 1763 1766 1769 1772 1775 1778 1781 1784 1787 1790 1793 1796 1799 1802 1805 1808 1811 1814 1817 1820 1823 1826 1829 1832 1835 1838 1841 1844 1847 1850 1853 1856 1859 1862 1865 1868 1871 1874 1877 1880 1883 1886 1889 1892 1895 1898 1901 1904 1907 1910 1913 1916 1919 1922 1925 1928 1931 1934 1937 1940 1943 1946 1949 1952 1955 1958 1961 1964 1967 1970 1973 1976 1979 1982 1985 1988 1991 1994 1997 2000 2003 2006 2009 2012 2015 2018 2021 2024 2027 2030 2033 2036 2039 2042 2045 2048 2051 2054 2057 2060 2063 2066 2069 2072 2075 2078 2081 2084 2087 2090 2093 2096 2099 2102 2105 2108 2111 2114 2117 2120 2123 2126 2129 2132 2135 2138 2141 2144 2147 2150 2153 2156 2159 2162 2165 2168 2171 2174}
    }
  }
}

harmonic {
  colvars cv
  centers 4.67348234758518e+03
  forceConstant 0.001
}

@HanatoK

HanatoK commented Jul 30, 2026

Copy link
Copy Markdown
Member Author

Performance test of coordnum with Apoa1

The configuration file is

indexFile index.ndx
# smp gpu

colvar {
  name cv
  coordNum {
    group1 {
      indexGroup backbone
    }
    group2 {
      indexGroup waters
    }
    cutoff 10.0
  }
}

harmonic {
  colvars cv
  centers 16774.0
  forceConstant 0.00001
}

smp gpu was turned on when testing with this PR and CUDAGM.

  • Equilibrium simulation without Colvars: 42.4303 ns/day
  • Colvars with CPU coordNum: 0.552597 ns/day
  • Colvars with GPU coordNum (this PR): 5.82301 ns/day

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants