Skip to content

Pin, verify and de-pickle every model artifact this repo loads - #4

Merged
lgoyal6 merged 2 commits into
mainfrom
supply-chain/model-artifact-guard
Sep 6, 2026
Merged

lgoyal6 merged 2 commits into
mainfrom
supply-chain/model-artifact-guard

Conversation

@lgoyal6

@lgoyal6 lgoyal6 commented Sep 6, 2026

Copy link
Copy Markdown
Collaborator

The trust boundary this closes

Every model in this repo enters through from_pretrained(...) reading a persistent, mutable Modal Volume used as the Hugging Face cache (llmlingua2-hf-cache, attentionrag-hf-cache, turboquant-hf-cache). The volume is populate-once-then-read and is committed back from inside the GPU container, so whatever sits in it at load time is what gets loaded. No load site pinned a revision, forced safetensors, or checked a digest.

It is not a hypothetical exposure. BAAI/bge-small-en-v1.5, the SmallEmbedder in two_stage_compressor, publishes both model.safetensors and pytorch_model.bin. A .bin checkpoint is a zipped Python pickle, and unpickling is arbitrary code execution that runs during load, before any shape or dtype is inspected. Whether the installed transformers happens to default weights_only=True is not something this repo controls, because the Modal images pip-install transformers unpinned. So the file is taken out of the decision rather than the loader trusted.

What model_guard enforces

provenance a pinned commit SHA per model id, read from the hub API on 2026-09-05, so a change to the hub's main cannot silently change the weights
no code trust_remote_code=False, and verify_snapshot refuses a snapshot containing any *.py
no pickle use_safetensors=True; a pickle that must be read at all goes through scan_pickle, an opcode-level allowlist that runs before any unpickling, and only then through torch.load(weights_only=True)
validation before activation digests and byte sizes against a manifest, then tensor key order, shapes, dtypes and finiteness
rollback KnownGood keeps the last manifest that passed, so a rejected candidate leaves the previous revision in place

Only the stdlib is on the validation path (pickletools, hashlib, json, zipfile). torch and safetensors are imported lazily and only when tensors are actually read, so the module imports and its tests run on a CPU-only box with neither installed.

Nothing bypasses it

The second commit routes all 17 load sites across 10 files through the guard: 14 from_pretrained calls become guarded_from_pretrained(...), and 3 snapshot_download calls become pinned_snapshot_download, wrapped in assert_no_pickled_weights wherever the snapshot is the thing later loaded.

llmlingua 0.2.2 defaults trust_remote_code to True inside PromptCompressor where the caller cannot reach it through kwargs, so the two llmlingua entrypoints pass llmlingua_model_config(...) to set it back to False explicitly.

test_model_guard.py adds three repo-wide AST sweeps, so a future load site that calls from_pretrained bare, or reaches for torch.load / pickle.load without the guard, fails the suite rather than passing unnoticed.

Tests

.agent-work/ is excluded from discovery: worktrees under it hold duplicate copies of these same test files. Explicit list of the four test entrypoints outside it.

file result
test_token_merge.py 14/14
test_model_guard.py 21/21
attentionrag/test_core.py 10/10
experiments/test_data.py rc=0, 0 tests (fixture module)
total 45 passing, 0 failing

Unchanged from the baseline. The 21 model_guard tests were already running from the working tree before this branch; this commit is what puts them under version control.

🤖 Generated with Claude Code

Every model in this repo enters through from_pretrained() reading a
persistent, mutable Modal Volume used as the Hugging Face cache. The
volume is populate-once-then-read and is committed back from inside the
GPU container, so whatever sits in it at load time is what gets loaded.
No load site pinned a revision, forced safetensors, or checked a digest,
which made the cache an unverified trust boundary.

That is not hypothetical here. BAAI/bge-small-en-v1.5, the SmallEmbedder
in two_stage_compressor, publishes both model.safetensors and
pytorch_model.bin. A .bin checkpoint is a zipped Python pickle, and
unpickling is arbitrary code execution that runs during load, before any
shape or dtype is inspected. Whether the installed transformers happens
to default weights_only=True is not something this repo controls, because
the Modal images pip-install transformers unpinned. So the file is taken
out of the decision rather than the loader trusted.

model_guard enforces five things:

  provenance  a pinned commit SHA per model id, read from the hub API on
              2026-09-05, so a change to the hub's main branch cannot
              silently change the weights
  no code     trust_remote_code=False, and verify_snapshot refuses a
              snapshot containing any *.py
  no pickle   use_safetensors=True; a pickle that must be read at all
              goes through scan_pickle, an opcode-level allowlist that
              runs BEFORE any unpickling, and only then through
              torch.load(weights_only=True)
  validation  digests and byte sizes against a manifest, then tensor key
              order, shapes, dtypes and finiteness, all before activation
  rollback    KnownGood keeps the last manifest that passed, so a
              rejected candidate leaves the previous revision in place

Only the stdlib is on the validation path (pickletools, hashlib, json,
zipfile). torch and safetensors are imported lazily and only when tensors
are actually read, so the module imports and its tests run on a CPU-only
box with neither installed.

test_model_guard.py covers the guard directly and adds three repo-wide
AST sweeps, so a future load site that calls from_pretrained bare, or
reaches for torch.load / pickle.load without the guard, fails the suite
rather than passing unnoticed.
The guard is only worth having if nothing bypasses it, so all seventeen
load sites across ten files now go through it instead of calling the hub
directly.

  * fourteen from_pretrained call sites become guarded_from_pretrained(...),
    which injects the pinned revision, use_safetensors=True and
    trust_remote_code=False: attentionrag/hf_backend.py,
    two_stage_compressor.py (both SmallEmbedder and CrossEncoderReranker),
    experiments/bench/bench_llm_modal.py, experiments/bench/run_compress.py,
    experiments/bench/compress_devpost.py, turboquant_modal.py and
    turboquant-poc/modal_app.py.
  * three snapshot_download call sites become pinned_snapshot_download, and
    where the snapshot is the thing that later gets loaded it is wrapped in
    assert_no_pickled_weights so a pickled checkpoint in the cache volume
    fails the container start rather than the inference:
    attentionrag/modal_app.py, llmlingua2_modal.py, experiments/eval_modal.py.
  * llmlingua 0.2.2 defaults trust_remote_code to True inside
    PromptCompressor, which the caller cannot reach through kwargs, so the
    two llmlingua entrypoints pass llmlingua_model_config(...) to set it
    back to False explicitly.
  * every Modal image that loads a model now ships model_guard alongside
    its own sources via add_local_python_source, otherwise the guarded
    import would not resolve inside the container.

After this, a repo-wide grep for a bare .from_pretrained( outside
model_guard finds nothing, which is what the AST sweep in
test_model_guard.py asserts.

No behaviour changes beyond the guard: same models, same dtypes, same
device maps. Prose em dashes in touched comments are replaced with
hyphens, per the repository's style.
@lgoyal6
lgoyal6 merged commit d579b9b into main Sep 6, 2026
1 check passed
@lgoyal6
lgoyal6 deleted the supply-chain/model-artifact-guard branch September 6, 2026 23:10
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant