Cozy Creator's model ETL: hub ingest (HF + Civitai), dtype cast / quantization, repackage, and Tensorhub publish.
- Ingest: HuggingFace (
HfApi.list_repo_files+ classifier +snapshot_download(allow_patterns=…)) and Civitai (bounded provider API). - Convert: streaming dtype cast + fp8-E4M3 storage cast (
#fp8flavor), bitsandbytes nf4/fp4, GGUF (llama.cpp toolchain), singlefile↔diffusers repackage. - Publish: one commit call against Tensorhub's HF-shaped
/commitswrite API (mode: merge|replace). - Tenant SDK:
Source,Component,Dataset,ProducedFlavor, streaming cast/fp8 writers, calibration policy — for@endpoint(kind="conversion")endpoints.
Model size must never dictate converter hardware. What each operation actually needs:
| Operation | Peak anonymous RAM | Why |
|---|---|---|
dtype cast (streaming_dtype_cast / streaming_cast_snapshot) |
≈ largest single tensor | two-pass streaming: header-only shard plan, then one tensor at a time |
fp8-E4M3 storage flavor (streaming_fp8_storage_cast / streaming_fp8_snapshot) |
≈ largest single tensor | same engine; clamp ±448 + layerwise-cast skip patterns |
byte-offset reshard / shard merge (shard_safetensors_by_offset) |
O(1) (8 MB copy chunks) | raw byte-range copy, no tensor decode |
| bnb nf4/fp4 | full component | from_pretrained load, inherent to bnb |
| singlefile↔diffusers repackage | full model | from_single_file / whole-keyspace remap needs the full tensor set — the one legitimate big-RAM operation |
| GGUF | full model (llama.cpp converter) | external toolchain |
Casts and fp8 flavor production run on the standard 32 GB CPU class regardless of model size; only repackage/GGUF of huge models still needs RAM sized to the model.
from gen_worker.convert import clone
result = clone.from_huggingface(ctx, payload) # download → convert → one commit per flavor