Hardware-Adaptive Edge Neural Graph Compiler
Zenthrix is an edge-native model compiler frontend for compiling open-weight neural networks (LLMs, SLMs, and vision models) into zero-copy, memory-optimized binaries tailored for consumer edge silicon — Apple Silicon, Qualcomm Snapdragon NPU, and Arm Cortex/Ethos.
- Direct Ingestion — Native loaders for SafeTensors, ONNX, PyTorch Export (AOTInductor), and GGUF architectures, with no intermediate format conversion required.
- Unified Memory Tiling — Schedules compute passes against unified memory architectures, reducing peak active RAM allocation by up to 40%.
- Zero-Copy Runtime — Emits standalone, relocatable
.zxbinaries that execute locally without a heavy Python runtime dependency. - Privacy-First Compilation — Models compile entirely on-device; weights and computational graphs never leave the local environment.
Install the precompiled command-line client and runtime via pip:
pip install zenthrixYou can also install and run it with uv:
uv tool install zenthrix
zenthrix --versionFor a one-off invocation without installing the command globally:
uvx zenthrix --version| Platform | Minimum Version |
|---|---|
| macOS | 14.0+ (Apple Silicon M1/M2/M3/M4) |
| Linux | Ubuntu 22.04+ (aarch64 / x86_64) |
| Android | NDK r25+ (for targeting Snapdragon platforms) |
Compilation requires the separately distributed native engine. Without it, the
CLI reports an actionable error rather than producing an invalid .zx file.
Compile an ONNX or GGUF model targeting local hardware execution:
zenthrix compile \
--model meta-llama/Llama-3.2-1B-Instruct \
--format onnx \
--target auto \
--quantization int4 \
--output ./llama-3.2-1b.zxInspect uncompiled source models (.safetensors, .gguf, .onnx) directly to view parameter counts, tensor shapes, and estimated memory footprint without requiring the native engine:
zenthrix inspect ./model.safetensors
zenthrix inspect ./model.gguf --jsonFor compiled .zx binaries, inspect projected memory profiles and operator fusions:
zenthrix inspect ./llama-3.2-1b.zx --memory-profileVerify compiled throughput directly in your terminal:
zenthrix run \
--model ./llama-3.2-1b.zx \
--prompt "Explain quantum decoherence in two sentences." \
--max-tokens 128import zenthrix
# Load and initialize the compiled runtime
engine = zenthrix.Engine(model_path="./llama-3.2-1b.zx")
# Execute a deterministic inference pass
output = engine.generate(
prompt="Synthesize the primary risks of high inference latency.",
temperature=0.2,
max_tokens=256,
)
print(output.text)
print(f"Time to First Token (TTFT): {output.ttft_ms} ms")
print(f"Throughput: {output.tokens_per_second} tokens/sec")The public package does not include the proprietary native compiler/runtime.
compile, inspect, and run therefore return a non-zero status until a
compatible runtime is provisioned; they never create placeholder artifacts or
claim successful inference.
For automation, add --json to run, inspect, or compile. Expected
failures are emitted as a JSON object with error and message fields:
{"error": "EngineUnavailableError", "message": "..."}Successful inference uses text, ttft_ms, and tokens_per_second fields.
| Silicon Target | Optimization Backend | Compute Units |
|---|---|---|
| Apple Silicon (M-Series / A-Series) | Metal MSL & AMX Matrix Intrinsics | GPU / Neural Engine |
| Qualcomm Snapdragon (8 Gen 2/3/4) | Hexagon HTP Architecture (C++) | NPU / HVX |
| Arm Neoverse / Cortex | Arm NEON / SVE2 Assembly | CPU Vector Extensions |
We welcome community contributions to adapters, loaders, and frontend parsers. All contributions require signing our Contributor License Agreement (CLA) during the pull request process.
See CONTRIBUTING.md for local environment setup instructions.
The Zenthrix CLI and client adapters are distributed under the Apache License 2.0. The underlying compilation engine dynamic binary is subject to the WithBrian Technologies Commercial EULA embedded in binary distributions.