Skip to content

Latest commit

 

History

History
196 lines (178 loc) · 16.5 KB

File metadata and controls

196 lines (178 loc) · 16.5 KB

Device and model reference

Auto-generated by python3 docs/gen_devices.py. Every row carries a confidence tag: vendor (datasheet), paper (peer-reviewed), press (announcement), derived (computed from published figures), estimate (inferred — do not quote as a vendor spec).

All FLOPS are dense unless a note says otherwise; vendor figures quoted with 2:4 structured sparsity have been halved.

GPUs

device vendor nm yr SMs bf16 int8 memory mem BW L2 TDP conf
a100_80gb_sxm NVIDIA 7 2020 108 312.00 TFLOPS 624.00 TFLOPS 80.000 GB 2.04 TB/s 40.000 MB 400 W vendor
amd_mi300x AMD 5 2023 304 1.31 PFLOPS 2.62 PFLOPS 192.000 GB 5.30 TB/s 256.000 MB 750 W vendor
ascend_910b Huawei 7 2023 32 376.00 TFLOPS 752.00 TFLOPS 64.000 GB 1.60 TB/s 192.000 MB 400 W press
ascend_910c Huawei 7 2025 64 752.00 TFLOPS 1.50 PFLOPS 128.000 GB 3.20 TB/s 384.000 MB 600 W estimate
b200_sxm NVIDIA 4 2024 148 2.25 PFLOPS 4.50 PFLOPS 192.000 GB 8.00 TB/s 126.000 MB 1000 W vendor
gb200_gpu NVIDIA 4 2024 148 2.50 PFLOPS 5.00 PFLOPS 192.000 GB 8.00 TB/s 126.000 MB 1200 W vendor
h100_sxm NVIDIA 5 2022 132 989.00 TFLOPS 1.98 PFLOPS 80.000 GB 3.35 TB/s 50.000 MB 700 W vendor
h200_sxm NVIDIA 5 2023 132 989.00 TFLOPS 1.98 PFLOPS 141.000 GB 4.80 TB/s 50.000 MB 700 W vendor
iluvatar_bi_v100 Iluvatar CoreX 天数智芯 7 2021 64 147.00 TFLOPS 295.00 TFLOPS 32.000 GB 1.20 TB/s 32.000 MB 250 W derived
iluvatar_mr_v100 Iluvatar CoreX 天数智芯 7 2022 64 96.00 TFLOPS 384.00 TFLOPS 32.000 GB 800.00 GB/s 32.000 MB 250 W derived
jetson_agx_orin_64 NVIDIA 8 2022 16 85.00 TFLOPS 170.00 TFLOPS 64.000 GB 204.80 GB/s 4.000 MB 60 W vendor
jetson_agx_thor NVIDIA 4 2025 20 250.00 TFLOPS 250.00 TFLOPS 128.000 GB 273.00 GB/s 12.000 MB 130 W press
jetson_orin_nx_16 NVIDIA 8 2023 8 25.00 TFLOPS 50.00 TFLOPS 16.000 GB 102.40 GB/s 2.000 MB 25 W vendor
l40s NVIDIA 5 2023 142 362.00 TFLOPS 733.00 TFLOPS 48.000 GB 864.00 GB/s 96.000 MB 350 W vendor
metax_c500 MetaX 沐曦 7 2023 104 280.00 TFLOPS 560.00 TFLOPS 64.000 GB 1.80 TB/s 48.000 MB 450 W derived
mthreads_s4000 Moore Threads 摩尔线程 7 2023 128 100.00 TFLOPS 200.00 TFLOPS 48.000 GB 768.00 GB/s 32.000 MB 450 W derived
mthreads_s5000 Moore Threads 摩尔线程 6 2026 160 250.00 TFLOPS 500.00 TFLOPS 80.000 GB 1.60 TB/s 64.000 MB 500 W estimate

GPU sources

  • a100_80gb_sxm — vendor A100 datasheet
  • amd_mi300x — vendor MI300X datasheet
    'SM' = CU count; 256 MB L2 is the Infinity Cache
  • ascend_910b — press + third-party measurement
    'SM' count = AI Core count; L2 is the unified on-chip buffer
  • ascend_910c — press
    dual-die 910B; public specs are approximate
  • b200_sxm — vendor Blackwell datasheet
    dense rates; NVIDIA markets 2x these with 2:4 sparsity
  • gb200_gpu — vendor GB200 NVL72 datasheet
  • h100_sxm — vendor H100 datasheet
  • h200_sxm — vendor H200 datasheet
  • iluvatar_bi_v100 — vendor spec page + third-party comparison
    天垓100 / BI-V100 training card. 24 Bn transistors, 2.5D CoWoS. This is the '较早型号' used in the reference deployment. SM count and cache sizes are estimates.
  • iluvatar_mr_v100 — vendor spec page
    智铠100 / MR-V100 inference card
  • jetson_agx_orin_64 — vendor Jetson AGX Orin datasheet
    LPDDR5 unified memory; 275 INT8 TOPS figure is 2:4-sparse
  • jetson_agx_thor — vendor Jetson AGX Thor announcement
    2070 FP4 TFLOPS vendor figure is sparse; dense halved here
  • jetson_orin_nx_16 — vendor Jetson Orin NX datasheet
  • l40s — vendor L40S datasheet
    GDDR6 not HBM; listed under 'hbm' level for uniformity
  • metax_c500 — vendor page + third-party survey
    曦云 C500. No FP8 support in current stack. Caches estimated.
  • mthreads_s4000 — vendor product page
    GDDR6 (not HBM). MUSA arch. Cache sizes estimated.
  • mthreads_s5000 — press
    announced specs; 1000 TFLOPS FP8 vendor figure is sparse

Brain-inspired many-core dataflow chips

device vendor arch nm yr cores SRAM/core SRAM total SRAM BW bf16 NoC bisect chip-to-chip board-to-board DRAM TDP conf
brainscales2 Heidelberg snn 65 2020 4 32.000 KB 128.000 KB 2.00 GB/s 5.00 GFLOPS 2.00 GB/s 1.00 GB/s 500.00 MB/s none 1 W paper
cerebras_wse3 Cerebras wafer 5 2024 900000 48.000 KB 41.199 GB 20700.00 TB/s 62.50 PFLOPS 26750.00 TB/s 1.20 TB/s 600.00 GB/s 1.172 TB @ 150.00 GB/s 23000 W vendor
darwin3 Zhejiang Univ 浙大 snn 22 2023 240 64.000 KB 15.000 MB 312.00 GB/s 600.00 GFLOPS 20.00 GB/s 2.00 GB/s 1.00 GB/s none 2 W estimate
groq_lpu_v1 Groq dataflow 14 2020 320 736.256 KB 230.080 MB 80.00 TB/s 188.00 TFLOPS 80.00 TB/s 200.00 GB/s 100.00 GB/s none 275 W paper
groq_lpu_v2 Groq dataflow 4 2025 320 1.600 MB 512.000 MB 192.00 TB/s 375.00 TFLOPS 192.00 TB/s 400.00 GB/s 200.00 GB/s none 350 W estimate
ipu_bow Graphcore dataflow 7 2022 1472 624.000 KB 897.000 MB 65.95 TB/s 350.00 TFLOPS 11.00 TB/s 320.00 GB/s 128.00 GB/s 128.000 GB @ 20.00 GB/s 350 W vendor
ipu_gc200 Graphcore dataflow 7 2020 1472 624.000 KB 897.000 MB 47.55 TB/s 250.00 TFLOPS 8.00 TB/s 320.00 GB/s 128.00 GB/s 64.000 GB @ 20.00 GB/s 300 W vendor
loihi2 Intel snn 7 2021 128 192.000 KB 24.000 MB 1.02 TB/s 500.00 GFLOPS 90.00 GB/s 6.00 GB/s 3.00 GB/s none 1 W estimate
lynxi_hp300_proj Lynxi 灵汐科技 brain 12 2025 256 3.000 MB 768.000 MB 51.20 TB/s 128.00 TFLOPS 2.00 TB/s 200.00 GB/s 100.00 GB/s 64.000 GB @ 200.00 GB/s 150 W estimate
lynxi_ka200 Lynxi 灵汐科技 brain 28 2021 30 1.600 MB 48.000 MB 2.40 TB/s 16.00 TFLOPS 350.00 GB/s 50.00 GB/s 25.00 GB/s 16.000 GB @ 68.00 GB/s 15 W estimate
sambanova_sn40l SambaNova dataflow 5 2023 1040 512.000 KB 520.000 MB 14.56 TB/s 638.00 TFLOPS 4.00 TB/s 400.00 GB/s 200.00 GB/s 1.500 TB @ 200.00 GB/s 700 W paper
spinnaker2 TU Dresden / SpiNNcloud snn 22 2024 152 128.000 KB 19.000 MB 364.80 GB/s 300.00 GFLOPS 30.00 GB/s 6.00 GB/s 3.00 GB/s 2.000 GB @ 10.00 GB/s 3 W paper
tesla_dojo_d1 Tesla dataflow 7 2021 354 1.250 MB 442.500 MB 40.00 TB/s 362.00 TFLOPS 10.00 TB/s 4.00 TB/s 900.00 GB/s none 400 W press
tianjic Tsinghua 清华 brain 28 2019 156 40.000 KB 6.094 MB 187.20 GB/s 320.00 GFLOPS 15.00 GB/s 4.00 GB/s 2.00 GB/s none 1 W paper
tianjic2_proj Tsinghua 清华 brain 12 2024 320 256.000 KB 80.000 MB 20.48 TB/s 64.00 TFLOPS 1.00 TB/s 100.00 GB/s 50.00 GB/s 16.000 GB @ 120.00 GB/s 35 W estimate
tianjicx Tsinghua 清华 brain 28 2022 16 80.000 KB 1.250 MB 64.00 GB/s 500.00 GFLOPS 16.00 GB/s 2.00 GB/s 1.00 GB/s none 1 W paper
truenorth IBM snn 28 2014 4096 1.300 KB 5.200 MB 5.32 GB/s 58.00 GFLOPS 400.00 MB/s 50.00 MB/s 20.00 MB/s none 0 W paper

SRAM die-area plausibility

SRAM array area implied by the quoted capacity and process node (6T bitcell + 1.35x array/periphery overhead), against an ~800 mm2 reticle. A part needing more than ~0.65 reticles of SRAM alone is a multi-die package whether or not the datasheet says so.

device SRAM node SRAM area reticles single die?
brainscales2 128.000 KB 65 nm 0.7 mm2 0.00 yes
cerebras_wse3 41.199 GB 5 nm 10032.9 mm2 12.54 no
darwin3 15.000 MB 22 nm 15.6 mm2 0.02 yes
groq_lpu_v1 230.080 MB 14 nm 166.8 mm2 0.21 yes
groq_lpu_v2 512.000 MB 4 nm 121.8 mm2 0.15 yes
ipu_bow 897.000 MB 7 nm 274.3 mm2 0.34 yes
ipu_gc200 897.000 MB 7 nm 274.3 mm2 0.34 yes
loihi2 24.000 MB 7 nm 7.3 mm2 0.01 yes
lynxi_hp300_proj 768.000 MB 12 nm 478.4 mm2 0.60 yes
lynxi_ka200 48.000 MB 28 nm 69.0 mm2 0.09 yes
sambanova_sn40l 520.000 MB 5 nm 123.7 mm2 0.15 yes
spinnaker2 19.000 MB 22 nm 19.8 mm2 0.02 yes
tesla_dojo_d1 442.500 MB 7 nm 135.3 mm2 0.17 yes
tianjic 6.094 MB 28 nm 8.8 mm2 0.01 yes
tianjic2_proj 80.000 MB 12 nm 49.8 mm2 0.06 yes
tianjicx 1.250 MB 28 nm 1.8 mm2 0.00 yes
truenorth 5.200 MB 28 nm 7.5 mm2 0.01 yes

Dataflow sources

  • brainscales2 — Pehle et al., Front. Neurosci. (2022)
    Analog/mixed-signal accelerated neuromorphic (1000x biological real time). 512 neurons, 130k synapses. Fundamentally not an LLM inference part; included for architectural comparison.
  • cerebras_wse3 — Cerebras WSE-3 datasheet
    Whole-wafer engine. Vendor headline 125 PFLOPS is FP16 with sparsity; dense halved here. 'dram' models MemoryX weight streaming. TDP is the full CS-3 system including cooling.
  • darwin3 — Ma et al. 2023 + press
    Darwin3: 2.35M neurons, 100M+ synapses, custom SNN ISA. SRAM/NoC figures partly estimated.
  • groq_lpu_v1 — Groq ISCA'20 TSP paper + vendor
    TSP: functionally-sliced, software-scheduled, no caches, fully deterministic. 320x320 fused dot-product unit, 5120 vector ALUs. NO DRAM AT ALL -- a 70B model needs hundreds of chips. Core count here models the 320 SIMD slices.
  • groq_lpu_v2 — press
    Samsung 4nm LPU v2. Capacity/BW figures are projections.
  • ipu_bow — vendor Bow datasheet
    Bow: wafer-on-wafer power delivery, +40% clock over GC200
  • ipu_gc200 — vendor GC200 datasheet + Graphcore papers
    MK2 IPU: 1472 tiles x 6 threads, BSP execution model. 'dram' is IPU-M2000 Streaming Memory (DDR4) -- note the very low 20 GB/s: spilling weights off-chip is catastrophic here.
  • loihi2 — Intel Loihi 2 tech brief
    Intel 4 process, 2.1 Bn transistors, 1M neurons, 120M synapses, programmable neuron models. Throughput figures estimated from published spike rates.
  • lynxi_hp300_proj — projection
    PROJECTED next-generation Lynxi part sized to make the GPU+brain-chip rack in the 2026 中国算力大会 announcement physically consistent (KA200 at 16 TFLOPS / 48 MB cannot host a DeepSeek-class MoE FFN). ALL NUMBERS ARE THE SIMULATOR AUTHOR'S CONSTRUCTION, NOT VENDOR SPECS. 768 MB at 12 nm is ~478 mm2 of SRAM array alone (60% of a reticle) -- buildable but aggressive; 150 W reflects that. Sweep sram_per_core and tdp with flowgpu sweep to test how sensitive the conclusions are to these two numbers.
  • lynxi_ka200 — vendor press + Science Robotics Tianjic lineage
    领启 KA200: 30 brain-computing cores, heterogeneous ANN+SNN, compute-in-memory, any-to-any core comms with multicast. 250k neurons / 25M synapses dense. 12-15 W. SRAM capacity/bandwidth and NoC figures are ESTIMATES -- Lynxi publishes neuron/synapse counts, not SRAM bytes. LPDDR4 is on the HP300/HM100 module, not on package.
  • sambanova_sn40l — SambaNova SN40L paper + vendor
    RDU with three-tier memory (SRAM/HBM/DDR). 'cores' models the PCU+PMU array. The DDR tier is what lets one node hold a trillion-parameter model -- at 200 GB/s.
  • spinnaker2 — Mayr/Hoeppner SpiNNaker2 papers
    152 ARM Cortex-M4F PEs + ML accelerators, 22nm FDSOI. Software-defined neurons -- flexible but low density.
  • tesla_dojo_d1 — Tesla AI Day / Hot Chips 34
    D1 die: 354 training nodes, 1.25 MB SRAM each, no DRAM. 25 D1 dies per training tile. CFP8 native.
  • tianjic — Pei et al., Nature 572 (2019)
    Nature 2019 cover chip. 156 FCores, 40k neurons, 10M synapses. 1.28 TOPS/W ANN, 649 GSOPS/W SNN. Research part -- far too small for LLM inference; included for completeness.
  • tianjic2_proj — projection
    PROJECTED Tianjic-2 class part. Not a vendor spec.
  • tianjicx — Ma et al., Science Robotics 7 (2022)
    TianjicX: spatiotemporal-elastic neuromorphic chip for robots (Science Robotics 2022). Optimised for latency-critical multi-task robotics, not datacentre LLM serving.
  • truenorth — Merolla et al., Science 345 (2014)
    1M neurons / 256M synapses at 70 mW. Binary spikes only, no multiply. CANNOT run a transformer -- the simulator will report an unsupported-dtype fallback if you try.

Calibrated energy coefficients

Produced by flowgpu.power.model.calibrate. Apportionment comes from measured Synopsys DC + PrimeTime PX runs (lsi10k, activity derived-rho0.5); absolute scale is anchored to each device's TDP. Off-chip DRAM energy is not rescaled -- 3.5 pJ/bit for HBM3 is interface physics, not a free parameter.

device source bf16 pJ/FLOP int8 pJ/FLOP SRAM pJ/B L2 pJ/B DRAM pJ/B NoC pJ/B/hop static W
h100_sxm eda:lsi10k + TDP 0.637 0.159 1.82 4.24 28.00 0.140 154
b200_sxm eda:lsi10k + TDP 0.414 0.104 1.18 3.59 25.60 0.091 220
a100_80gb_sxm eda:lsi10k + TDP 0.953 0.238 2.44 5.88 29.60 0.209 88
iluvatar_bi_v100 eda:lsi10k + TDP 1.390 0.348 3.27 7.90 31.20 0.304 55
ascend_910b eda:lsi10k + TDP 0.946 0.237 5.38 9.08 29.60 0.207 88
groq_lpu_v1 eda:lsi10k + TDP 0.642 0.161 3.23 0.00 0.00 0.141 50
ipu_gc200 eda:lsi10k + TDP 0.901 0.225 4.24 0.00 144.00 0.197 54
lynxi_ka200 eda:lsi10k + TDP 0.606 0.152 4.07 0.00 64.00 0.133 3
lynxi_hp300_proj eda:lsi10k + TDP 0.379 0.095 3.07 0.00 40.00 0.083 27
cerebras_wse3 eda:lsi10k + TDP 0.342 0.085 0.51 0.00 40.00 0.075 4140

Models

model family L d total params active experts top-k attn KV/token source
deepseek_r1 llm 61 7168 671.0 B 36.6 B 256 8 MLA 68.6 KiB deepseek-ai/DeepSeek-R1 config.json (same arch as V3)
deepseek_v2_lite llm 27 2048 15.7 B 2.5 B 64 6 MLA 30.4 KiB deepseek-ai/DeepSeek-V2-Lite config.json
deepseek_v3 llm 61 7168 671.0 B 36.6 B 256 8 MLA 68.6 KiB deepseek-ai/DeepSeek-V3 config.json
deepseek_v4_flash_proj llm 48 4096 153.0 B 8.9 B 256 8 MLA 48.0 KiB PROJECTION -- not a published config
gpt_oss_120b llm 36 2880 116.8 B 5.1 B 128 4 GQA 72.0 KiB openai/gpt-oss-120b config.json
llama3_1b llm 16 2048 1.5 B 1.2 B - - GQA 32.0 KiB Meta Llama-3.2-1B config.json
llama3_405b llm 126 16384 405.9 B 403.8 B - - GQA 504.0 KiB Meta Llama-3.1-405B config.json
llama3_70b llm 80 8192 70.6 B 69.5 B - - GQA 320.0 KiB Meta Llama-3.1-70B config.json
llama3_8b llm 32 4096 8.0 B 7.5 B - - GQA 128.0 KiB Meta Llama-3.1-8B config.json
mixtral_8x7b llm 32 4096 46.7 B 12.7 B 8 2 GQA 128.0 KiB Mistral Mixtral-8x7B config.json
qwen3_235b_a22b llm 94 4096 235.1 B 21.6 B 128 8 GQA 188.0 KiB Qwen3-235B-A22B config.json
qwen3_32b llm 64 5120 32.8 B 32.0 B - - GQA 256.0 KiB Qwen3-32B config.json
qwen3_8b llm 36 4096 8.2 B 7.6 B - - GQA 144.0 KiB Qwen3-8B config.json
internvl3_38b vlm 64 5120 32.8 B 32.0 B - - GQA 256.0 KiB InternVL3-38B config.json
qwen2_5_vl_7b vlm 28 3584 7.6 B 7.1 B - - GQA 56.0 KiB Qwen2.5-VL-7B config.json
flux_dev diffusion - - - - - - dit -
sd3_medium diffusion - - - - - - dit -
sdxl diffusion - - - - - - unet -
resnet18 cnn - - - - - - resnet18 -
resnet50 cnn - - - - - - resnet50 -
vit_l_16 cnn - - - - - - vit_l_16 -
yolov8l cnn - - - - - - yolov8l -