Apple Neural Engine (ANE/NPU) vs GPU for local LLM inference on Apple Silicon / macOS. Core ML (CoreML), Core AI (CoreAI), MLX and Metal benchmarks: prefill, throughput, latency, memory, thermals, INT4/INT8, W4A16/A8W4 quantization, grouped scales and FP16 arithmetic. Reproducible component tests, compatibility findings, English/Chinese articles.
macos benchmark metal gpu quantization ane mlx int8 coreml prefill npu apple-silicon int4 llm apple-neural-engine local-llm llm-inference coreai w4a16 a8w4
-
Updated
Sep 12, 2026 - Python