This repository contains fully automated, deterministic benchmarking scripts comparing Zenthrix runtime performance against industry baselines:
- Apple CoreML Tools (v0.8+)
- llama.cpp (Metal & Hexagon builds)
- PyTorch ExecuTorch (v0.3+)
Results are intended for reproducibility and review. Only validated result files should be published; each result must identify its implementation, hardware, model, and metric values.
- Primary Metrics Evaluated
- Verified Benchmark Results
- Reproducing the Results
- Integrity and Hardware Telemetry
- Changelog
- Contributing
- License
- Time To First Token (TTFT) — Milliseconds elapsed before emitting the initial completion token.
- Tokens Per Second (TPS) — Sustained autoregressive generation throughput under fixed sequence lengths (context = 512 tokens, generated = 128 tokens).
- Peak Resident Set Size (RSS) — Physical RAM overhead measured throughout the execution lifecycle.
- Thermal Throttling Threshold — Latency variation across continuous 15-minute inference stress tests.
An example validated result file is available at
examples/results.json. Validate it with:
uv run zbench validate examples/results.json
uv run zbench report examples/results.jsonThe result validator accepts ttft_ms, tokens_per_second, peak_memory_mb,
and load_time_ms. Values must be finite and positive. Versioned result
envelopes also record an ISO 8601 timestamp, source commit, environment, and
measurement configuration; all results in one file must share the same
hardware and model context.
git clone https://github.com/withbrian-technologies/zenthrix-benchmarks.git
cd zenthrix-benchmarks
uv syncExecute the automated harness across all installed framework runtimes:
python scripts/run_zenthrix_eval.py --config configs/llama_3.2_1b.yaml --device local
python scripts/run_coreml_eval.py --config configs/llama_3.2_1b.yaml
python scripts/run_llamacpp_eval.sh --model llama-3.2-1b.Q4_K_M.ggufProcess the raw telemetry JSONs into visual performance plots:
python scripts/generate_charts.py --input results/ --output-dir ./chartsThe zbench utility validates result JSON files and renders stable Markdown
summaries:
uv run zbench validate results.json
uv run zbench report results.json
Report columns are emitted in a stable order, implementations are sorted by
name, and the best value in each metric is bolded. Higher throughput is better;
lower latency, memory, and load time are better. Missing measurements are shown
as — rather than being treated as zero.
All metrics are captured via OS-level tracing utilities (powermetrics on macOS; perf and sysfs on Linux/Android) to avoid runtime measurement distortion. System-level CPU/GPU/NPU core frequencies and thermal zones are logged continuously to verify the absence of thermal throttling during comparative iterations.
The benchmarking harness, scripts, and aggregated result datasets are released under the Apache License 2.0.