You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Auto-generated by python3 docs/gen_devices.py. Every row carries a
confidence tag: vendor (datasheet), paper (peer-reviewed),
press (announcement), derived (computed from published figures),
estimate (inferred — do not quote as a vendor spec).
All FLOPS are dense unless a note says otherwise; vendor figures
quoted with 2:4 structured sparsity have been halved.
GPUs
device
vendor
nm
yr
SMs
bf16
int8
memory
mem BW
L2
TDP
conf
a100_80gb_sxm
NVIDIA
7
2020
108
312.00 TFLOPS
624.00 TFLOPS
80.000 GB
2.04 TB/s
40.000 MB
400 W
vendor
amd_mi300x
AMD
5
2023
304
1.31 PFLOPS
2.62 PFLOPS
192.000 GB
5.30 TB/s
256.000 MB
750 W
vendor
ascend_910b
Huawei
7
2023
32
376.00 TFLOPS
752.00 TFLOPS
64.000 GB
1.60 TB/s
192.000 MB
400 W
press
ascend_910c
Huawei
7
2025
64
752.00 TFLOPS
1.50 PFLOPS
128.000 GB
3.20 TB/s
384.000 MB
600 W
estimate
b200_sxm
NVIDIA
4
2024
148
2.25 PFLOPS
4.50 PFLOPS
192.000 GB
8.00 TB/s
126.000 MB
1000 W
vendor
gb200_gpu
NVIDIA
4
2024
148
2.50 PFLOPS
5.00 PFLOPS
192.000 GB
8.00 TB/s
126.000 MB
1200 W
vendor
h100_sxm
NVIDIA
5
2022
132
989.00 TFLOPS
1.98 PFLOPS
80.000 GB
3.35 TB/s
50.000 MB
700 W
vendor
h200_sxm
NVIDIA
5
2023
132
989.00 TFLOPS
1.98 PFLOPS
141.000 GB
4.80 TB/s
50.000 MB
700 W
vendor
iluvatar_bi_v100
Iluvatar CoreX 天数智芯
7
2021
64
147.00 TFLOPS
295.00 TFLOPS
32.000 GB
1.20 TB/s
32.000 MB
250 W
derived
iluvatar_mr_v100
Iluvatar CoreX 天数智芯
7
2022
64
96.00 TFLOPS
384.00 TFLOPS
32.000 GB
800.00 GB/s
32.000 MB
250 W
derived
jetson_agx_orin_64
NVIDIA
8
2022
16
85.00 TFLOPS
170.00 TFLOPS
64.000 GB
204.80 GB/s
4.000 MB
60 W
vendor
jetson_agx_thor
NVIDIA
4
2025
20
250.00 TFLOPS
250.00 TFLOPS
128.000 GB
273.00 GB/s
12.000 MB
130 W
press
jetson_orin_nx_16
NVIDIA
8
2023
8
25.00 TFLOPS
50.00 TFLOPS
16.000 GB
102.40 GB/s
2.000 MB
25 W
vendor
l40s
NVIDIA
5
2023
142
362.00 TFLOPS
733.00 TFLOPS
48.000 GB
864.00 GB/s
96.000 MB
350 W
vendor
metax_c500
MetaX 沐曦
7
2023
104
280.00 TFLOPS
560.00 TFLOPS
64.000 GB
1.80 TB/s
48.000 MB
450 W
derived
mthreads_s4000
Moore Threads 摩尔线程
7
2023
128
100.00 TFLOPS
200.00 TFLOPS
48.000 GB
768.00 GB/s
32.000 MB
450 W
derived
mthreads_s5000
Moore Threads 摩尔线程
6
2026
160
250.00 TFLOPS
500.00 TFLOPS
80.000 GB
1.60 TB/s
64.000 MB
500 W
estimate
GPU sources
a100_80gb_sxm — vendor A100 datasheet
amd_mi300x — vendor MI300X datasheet
'SM' = CU count; 256 MB L2 is the Infinity Cache
ascend_910b — press + third-party measurement
'SM' count = AI Core count; L2 is the unified on-chip buffer
ascend_910c — press
dual-die 910B; public specs are approximate
b200_sxm — vendor Blackwell datasheet
dense rates; NVIDIA markets 2x these with 2:4 sparsity
gb200_gpu — vendor GB200 NVL72 datasheet
h100_sxm — vendor H100 datasheet
h200_sxm — vendor H200 datasheet
iluvatar_bi_v100 — vendor spec page + third-party comparison
天垓100 / BI-V100 training card. 24 Bn transistors, 2.5D CoWoS. This is the '较早型号' used in the reference deployment. SM count and cache sizes are estimates.
mthreads_s5000 — press
announced specs; 1000 TFLOPS FP8 vendor figure is sparse
Brain-inspired many-core dataflow chips
device
vendor
arch
nm
yr
cores
SRAM/core
SRAM total
SRAM BW
bf16
NoC bisect
chip-to-chip
board-to-board
DRAM
TDP
conf
brainscales2
Heidelberg
snn
65
2020
4
32.000 KB
128.000 KB
2.00 GB/s
5.00 GFLOPS
2.00 GB/s
1.00 GB/s
500.00 MB/s
none
1 W
paper
cerebras_wse3
Cerebras
wafer
5
2024
900000
48.000 KB
41.199 GB
20700.00 TB/s
62.50 PFLOPS
26750.00 TB/s
1.20 TB/s
600.00 GB/s
1.172 TB @ 150.00 GB/s
23000 W
vendor
darwin3
Zhejiang Univ 浙大
snn
22
2023
240
64.000 KB
15.000 MB
312.00 GB/s
600.00 GFLOPS
20.00 GB/s
2.00 GB/s
1.00 GB/s
none
2 W
estimate
groq_lpu_v1
Groq
dataflow
14
2020
320
736.256 KB
230.080 MB
80.00 TB/s
188.00 TFLOPS
80.00 TB/s
200.00 GB/s
100.00 GB/s
none
275 W
paper
groq_lpu_v2
Groq
dataflow
4
2025
320
1.600 MB
512.000 MB
192.00 TB/s
375.00 TFLOPS
192.00 TB/s
400.00 GB/s
200.00 GB/s
none
350 W
estimate
ipu_bow
Graphcore
dataflow
7
2022
1472
624.000 KB
897.000 MB
65.95 TB/s
350.00 TFLOPS
11.00 TB/s
320.00 GB/s
128.00 GB/s
128.000 GB @ 20.00 GB/s
350 W
vendor
ipu_gc200
Graphcore
dataflow
7
2020
1472
624.000 KB
897.000 MB
47.55 TB/s
250.00 TFLOPS
8.00 TB/s
320.00 GB/s
128.00 GB/s
64.000 GB @ 20.00 GB/s
300 W
vendor
loihi2
Intel
snn
7
2021
128
192.000 KB
24.000 MB
1.02 TB/s
500.00 GFLOPS
90.00 GB/s
6.00 GB/s
3.00 GB/s
none
1 W
estimate
lynxi_hp300_proj
Lynxi 灵汐科技
brain
12
2025
256
3.000 MB
768.000 MB
51.20 TB/s
128.00 TFLOPS
2.00 TB/s
200.00 GB/s
100.00 GB/s
64.000 GB @ 200.00 GB/s
150 W
estimate
lynxi_ka200
Lynxi 灵汐科技
brain
28
2021
30
1.600 MB
48.000 MB
2.40 TB/s
16.00 TFLOPS
350.00 GB/s
50.00 GB/s
25.00 GB/s
16.000 GB @ 68.00 GB/s
15 W
estimate
sambanova_sn40l
SambaNova
dataflow
5
2023
1040
512.000 KB
520.000 MB
14.56 TB/s
638.00 TFLOPS
4.00 TB/s
400.00 GB/s
200.00 GB/s
1.500 TB @ 200.00 GB/s
700 W
paper
spinnaker2
TU Dresden / SpiNNcloud
snn
22
2024
152
128.000 KB
19.000 MB
364.80 GB/s
300.00 GFLOPS
30.00 GB/s
6.00 GB/s
3.00 GB/s
2.000 GB @ 10.00 GB/s
3 W
paper
tesla_dojo_d1
Tesla
dataflow
7
2021
354
1.250 MB
442.500 MB
40.00 TB/s
362.00 TFLOPS
10.00 TB/s
4.00 TB/s
900.00 GB/s
none
400 W
press
tianjic
Tsinghua 清华
brain
28
2019
156
40.000 KB
6.094 MB
187.20 GB/s
320.00 GFLOPS
15.00 GB/s
4.00 GB/s
2.00 GB/s
none
1 W
paper
tianjic2_proj
Tsinghua 清华
brain
12
2024
320
256.000 KB
80.000 MB
20.48 TB/s
64.00 TFLOPS
1.00 TB/s
100.00 GB/s
50.00 GB/s
16.000 GB @ 120.00 GB/s
35 W
estimate
tianjicx
Tsinghua 清华
brain
28
2022
16
80.000 KB
1.250 MB
64.00 GB/s
500.00 GFLOPS
16.00 GB/s
2.00 GB/s
1.00 GB/s
none
1 W
paper
truenorth
IBM
snn
28
2014
4096
1.300 KB
5.200 MB
5.32 GB/s
58.00 GFLOPS
400.00 MB/s
50.00 MB/s
20.00 MB/s
none
0 W
paper
SRAM die-area plausibility
SRAM array area implied by the quoted capacity and process node
(6T bitcell + 1.35x array/periphery overhead), against an ~800 mm2
reticle. A part needing more than ~0.65 reticles of SRAM alone is a
multi-die package whether or not the datasheet says so.
device
SRAM
node
SRAM area
reticles
single die?
brainscales2
128.000 KB
65 nm
0.7 mm2
0.00
yes
cerebras_wse3
41.199 GB
5 nm
10032.9 mm2
12.54
no
darwin3
15.000 MB
22 nm
15.6 mm2
0.02
yes
groq_lpu_v1
230.080 MB
14 nm
166.8 mm2
0.21
yes
groq_lpu_v2
512.000 MB
4 nm
121.8 mm2
0.15
yes
ipu_bow
897.000 MB
7 nm
274.3 mm2
0.34
yes
ipu_gc200
897.000 MB
7 nm
274.3 mm2
0.34
yes
loihi2
24.000 MB
7 nm
7.3 mm2
0.01
yes
lynxi_hp300_proj
768.000 MB
12 nm
478.4 mm2
0.60
yes
lynxi_ka200
48.000 MB
28 nm
69.0 mm2
0.09
yes
sambanova_sn40l
520.000 MB
5 nm
123.7 mm2
0.15
yes
spinnaker2
19.000 MB
22 nm
19.8 mm2
0.02
yes
tesla_dojo_d1
442.500 MB
7 nm
135.3 mm2
0.17
yes
tianjic
6.094 MB
28 nm
8.8 mm2
0.01
yes
tianjic2_proj
80.000 MB
12 nm
49.8 mm2
0.06
yes
tianjicx
1.250 MB
28 nm
1.8 mm2
0.00
yes
truenorth
5.200 MB
28 nm
7.5 mm2
0.01
yes
Dataflow sources
brainscales2 — Pehle et al., Front. Neurosci. (2022)
Analog/mixed-signal accelerated neuromorphic (1000x biological real time). 512 neurons, 130k synapses. Fundamentally not an LLM inference part; included for architectural comparison.
cerebras_wse3 — Cerebras WSE-3 datasheet
Whole-wafer engine. Vendor headline 125 PFLOPS is FP16 with sparsity; dense halved here. 'dram' models MemoryX weight streaming. TDP is the full CS-3 system including cooling.
darwin3 — Ma et al. 2023 + press
Darwin3: 2.35M neurons, 100M+ synapses, custom SNN ISA. SRAM/NoC figures partly estimated.
groq_lpu_v1 — Groq ISCA'20 TSP paper + vendor
TSP: functionally-sliced, software-scheduled, no caches, fully deterministic. 320x320 fused dot-product unit, 5120 vector ALUs. NO DRAM AT ALL -- a 70B model needs hundreds of chips. Core count here models the 320 SIMD slices.
lynxi_hp300_proj — projection
PROJECTED next-generation Lynxi part sized to make the GPU+brain-chip rack in the 2026 中国算力大会 announcement physically consistent (KA200 at 16 TFLOPS / 48 MB cannot host a DeepSeek-class MoE FFN). ALL NUMBERS ARE THE SIMULATOR AUTHOR'S CONSTRUCTION, NOT VENDOR SPECS. 768 MB at 12 nm is ~478 mm2 of SRAM array alone (60% of a reticle) -- buildable but aggressive; 150 W reflects that. Sweep sram_per_core and tdp with flowgpu sweep to test how sensitive the conclusions are to these two numbers.
lynxi_ka200 — vendor press + Science Robotics Tianjic lineage
领启 KA200: 30 brain-computing cores, heterogeneous ANN+SNN, compute-in-memory, any-to-any core comms with multicast. 250k neurons / 25M synapses dense. 12-15 W. SRAM capacity/bandwidth and NoC figures are ESTIMATES -- Lynxi publishes neuron/synapse counts, not SRAM bytes. LPDDR4 is on the HP300/HM100 module, not on package.
sambanova_sn40l — SambaNova SN40L paper + vendor
RDU with three-tier memory (SRAM/HBM/DDR). 'cores' models the PCU+PMU array. The DDR tier is what lets one node hold a trillion-parameter model -- at 200 GB/s.
spinnaker2 — Mayr/Hoeppner SpiNNaker2 papers
152 ARM Cortex-M4F PEs + ML accelerators, 22nm FDSOI. Software-defined neurons -- flexible but low density.
tesla_dojo_d1 — Tesla AI Day / Hot Chips 34
D1 die: 354 training nodes, 1.25 MB SRAM each, no DRAM. 25 D1 dies per training tile. CFP8 native.
tianjic — Pei et al., Nature 572 (2019)
Nature 2019 cover chip. 156 FCores, 40k neurons, 10M synapses. 1.28 TOPS/W ANN, 649 GSOPS/W SNN. Research part -- far too small for LLM inference; included for completeness.
tianjic2_proj — projection
PROJECTED Tianjic-2 class part. Not a vendor spec.
tianjicx — Ma et al., Science Robotics 7 (2022)
TianjicX: spatiotemporal-elastic neuromorphic chip for robots (Science Robotics 2022). Optimised for latency-critical multi-task robotics, not datacentre LLM serving.
truenorth — Merolla et al., Science 345 (2014)
1M neurons / 256M synapses at 70 mW. Binary spikes only, no multiply. CANNOT run a transformer -- the simulator will report an unsupported-dtype fallback if you try.
Calibrated energy coefficients
Produced by flowgpu.power.model.calibrate. Apportionment comes
from measured Synopsys DC + PrimeTime PX runs (lsi10k, activity derived-rho0.5); absolute scale is anchored to
each device's TDP. Off-chip DRAM energy is not rescaled --
3.5 pJ/bit for HBM3 is interface physics, not a free parameter.
device
source
bf16 pJ/FLOP
int8 pJ/FLOP
SRAM pJ/B
L2 pJ/B
DRAM pJ/B
NoC pJ/B/hop
static W
h100_sxm
eda:lsi10k + TDP
0.637
0.159
1.82
4.24
28.00
0.140
154
b200_sxm
eda:lsi10k + TDP
0.414
0.104
1.18
3.59
25.60
0.091
220
a100_80gb_sxm
eda:lsi10k + TDP
0.953
0.238
2.44
5.88
29.60
0.209
88
iluvatar_bi_v100
eda:lsi10k + TDP
1.390
0.348
3.27
7.90
31.20
0.304
55
ascend_910b
eda:lsi10k + TDP
0.946
0.237
5.38
9.08
29.60
0.207
88
groq_lpu_v1
eda:lsi10k + TDP
0.642
0.161
3.23
0.00
0.00
0.141
50
ipu_gc200
eda:lsi10k + TDP
0.901
0.225
4.24
0.00
144.00
0.197
54
lynxi_ka200
eda:lsi10k + TDP
0.606
0.152
4.07
0.00
64.00
0.133
3
lynxi_hp300_proj
eda:lsi10k + TDP
0.379
0.095
3.07
0.00
40.00
0.083
27
cerebras_wse3
eda:lsi10k + TDP
0.342
0.085
0.51
0.00
40.00
0.075
4140
Models
model
family
L
d
total params
active
experts
top-k
attn
KV/token
source
deepseek_r1
llm
61
7168
671.0 B
36.6 B
256
8
MLA
68.6 KiB
deepseek-ai/DeepSeek-R1 config.json (same arch as V3)