Skip to content
#

a8w4

Here is 1 public repository matching this topic...

Apple Neural Engine (ANE/NPU) vs GPU for local LLM inference on Apple Silicon / macOS. Core ML (CoreML), Core AI (CoreAI), MLX and Metal benchmarks: prefill, throughput, latency, memory, thermals, INT4/INT8, W4A16/A8W4 quantization, grouped scales and FP16 arithmetic. Reproducible component tests, compatibility findings, English/Chinese articles.

  • Updated Sep 12, 2026
  • Python

Add this topic to your repo

To associate your repository with the a8w4 topic, visit your repo's landing page and select "manage topics."

Learn more