Skip to content
#

byte-level

Here are 19 public repositories matching this topic...

A tiny byte-level multi-head content classifier (~1.5M params, ~200KB ONNX, <6ms). Classifies code, text, markup, config, images, binary, secrets, 62 code languages, 30 text languages, 90 MIME types from raw bytes — no tokenizer needed.

  • Updated Sep 26, 2026
  • Python
purebyte

Tiny byte-level AI on CPU: a zero-dependency C++17 runtime and neural decision models (0.8 ms per 64-byte decision, 9.4 MB RAM). Reads raw bytes, no tokenizer. Specialists for secret scanning (F1 0.797 vs 0.337 for gitleaks on CredData) and PII redaction. Code under Apache-2.0; model weights under their own license.

  • Updated Sep 28, 2026
  • C++
purebyte-train

The training stack for PureByte: from a task definition to a verified GGUF specialist in 25 minutes to an hour on one consumer GPU. Synthetic look-alikes grounded in real bytes, leak checks against the exams, criteria frozen before training, multi-seed runs with controls, and proof that the C++ runtime gives the same answers.

  • Updated Sep 28, 2026
  • Python

Add this topic to your repo

To associate your repository with the byte-level topic, visit your repo's landing page and select "manage topics."

Learn more