Skip to content

Package the CPU-side core as winnow-core, and lock what the GPU half needs - #5

Merged
lgoyal6 merged 2 commits into
mainfrom
packaging/winnow-core
Sep 6, 2026
Merged

lgoyal6 merged 2 commits into
mainfrom
packaging/winnow-core

Conversation

@lgoyal6

@lgoyal6 lgoyal6 commented Sep 6, 2026

Copy link
Copy Markdown
Collaborator

What this changes

The compression core is the part of this repo an outside consumer can actually use: the AttentionRAG selection logic, the two-compressor token merge, and the model-artifact guard. None of it needs a GPU, a Modal account or an API key, but until now the only way to get it was to clone the repo and hope the imports resolved.

pyproject.toml builds it as winnow-core, exporting the attentionrag package plus the token_merge and model_guard top-level modules.

dependencies is deliberately empty. attentionrag.core is the paper's selection logic behind a Backend protocol, token_merge is pure string and span arithmetic, and model_guard's validation path uses only pickletools, hashlib, json and zipfile. That is what makes this half installable and testable without a GPU.

Verified, not assumed

Run on this machine before opening the PR:

step result
python -m build winnow_core-0.1.0-py3-none-any.whl
pip install <wheel> in a clean venv pip freeze lists winnow-core and nothing else
cd /tmp && python -m attentionrag.test_core 10/10 core tests passed
import token_merge, model_guard, attentionrag from /tmp all resolve from site-packages

The /tmp working directory matters: it proves the install works off the wheel rather than off the checkout sitting in the current directory.

The [hf] extra is pinned, not floated

The guard's whole point is that "the same model name" must not mean different bytes tomorrow, and the same argument applies to the loader. numpy is held at 2.4.6 rather than the newest release because safetensors needs numpy to convert torch tensors and numpy 2.5.x publishes no cp311 wheel, which is the interpreter the Modal images use.

requirements-hf.lock is the pip freeze of that extra in a clean 3.11 venv and records the command to reproduce it, so the resolved set is a fact rather than a wish. requirements.txt is the separate Modal/FastAPI client freeze for the local proxy in server.py.

One deliberate correction

The Homepage URL read https://github.com/winnow, which is not this repository and does not resolve. It now points at https://github.com/TC960/winnow. Flagging it because it is the one line here that is not simply the working tree as it stood.

Commit split

The em dash to hyphen sweep in README.md and .gitignore is a separate first commit, so the packaging diff is readable rather than a page of punctuation.

Tests

.agent-work/ is excluded from discovery: worktrees under it hold duplicate copies of these same test files.

file result
test_token_merge.py 14/14
test_model_guard.py 21/21
attentionrag/test_core.py 10/10
experiments/test_data.py rc=0, 0 tests (fixture module)
total 45 passing, 0 failing

Unchanged from the baseline. Packaging adds no runnable test of its own; the wheel build and clean-venv install above are its evidence.

🤖 Generated with Claude Code

Repository style is hyphens, and these two files were the last places in
the tree still using em dashes in prose. Text only: 19 lines in README.md
and the one comment line in .gitignore. No rules, no instructions and no
code change, so the two files behave exactly as before.

Split out from the packaging change on the next commit so that one is a
readable diff rather than a page of punctuation.
…needs

The compression core is the part of this repo an outside consumer can
actually use: the AttentionRAG selection logic, the two-compressor token
merge, and the model-artifact guard. None of it needs a GPU, a Modal
account or an API key, but until now the only way to get it was to clone
the repo and hope the imports resolved.

pyproject.toml builds it as `winnow-core` with setuptools, exporting the
attentionrag package plus the token_merge and model_guard top-level
modules. dependencies is deliberately empty: attentionrag.core is the
paper's selection logic behind a Backend protocol, token_merge is pure
string and span arithmetic, and model_guard's validation path uses only
pickletools, hashlib, json and zipfile. That is what makes this half
installable and testable without a GPU.

Verified rather than assumed, on this machine:

  python -m build                                    -> winnow_core-0.1.0-py3-none-any.whl
  pip install <that wheel> in a clean venv           -> pip freeze lists winnow-core and nothing else
  cd /tmp && python -m attentionrag.test_core        -> 10/10 core tests passed
  import token_merge, model_guard, attentionrag      -> all resolve from site-packages

The model-touching half goes in an [hf] extra, pinned rather than
floated: the guard's whole point is that "the same model name" must not
mean different bytes tomorrow, and the same argument applies to the
loader. numpy is held at 2.4.6 rather than the newest release because
safetensors needs numpy to convert torch tensors and numpy 2.5.x
publishes no cp311 wheel, which is the interpreter the Modal images use.

requirements-hf.lock is the pip freeze of that extra in a clean 3.11 venv
and records the reproduce command, so the resolved set is a fact rather
than a wish. requirements.txt is the separate Modal/FastAPI client freeze
for the local proxy in server.py.

.gitignore picks up /dist/, /build/ and *.egg-info/ so a local build does
not show up as untracked noise.

One deliberate correction while landing this: the Homepage URL read
https://github.com/winnow, which is not this repository and does not
resolve. It now points at https://github.com/TC960/winnow.
@lgoyal6
lgoyal6 merged commit f28f698 into main Sep 6, 2026
1 check passed
@lgoyal6
lgoyal6 deleted the packaging/winnow-core branch September 6, 2026 23:14
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant