Skip to content

Fix canonicalize_license_expression mangling case-distinct LicenseRef identifiers - #1412

Open
Str0k wants to merge 1 commit into
pypa:mainfrom
Str0k:githubpower/t_f074219d
Open

Str0k wants to merge 1 commit into
pypa:mainfrom
Str0k:githubpower/t_f074219d

Conversation

@Str0k

@Str0k Str0k commented Sep 14, 2026

Copy link
Copy Markdown

When a license expression contains two LicenseRef- identifiers that differ only in case, e.g. 'licenseref-MIT AND LicenseRef-mit', canonicalize_license_expression() rewrote every occurrence to the spelling of the last one seen, returning 'LicenseRef-mit AND LicenseRef-mit'. This happens because the module lowercases the expression for SPDX lookup and then rebuilds each LicenseRef token from a dict keyed by the lowercased token, so one casing silently overwrites the other. After the fix, each occurrence keeps its own spelling and 'licenseref-MIT AND LicenseRef-mit' canonicalizes to 'LicenseRef-MIT AND LicenseRef-mit'.

SPDX LicenseRef- identifiers are case-sensitive: LicenseRef-mit and LicenseRef-MIT name two different licenses. Tests in tests/test_metadata.py already document that a single ref's own spelling is preserved, and the lowercased dict broke that contract for expressions mixing casings of the same base identifier. The fix records the original spelling of each token occurrence (keyed by token index, with a strict zip tie to the lowercased tokens) and looks it up by index while emitting normalized tokens, leaving all other behavior unchanged.

Validation:

  • Regression on unchanged base: 4 assertion failure(s); with patch: 5 tests, exit 0.
  • Full suite: base 62440 tests (exit 1); patch 62440 tests (exit 0).
  • Repository checks: coverage_gaps: exit 0, license_header: exit 1, mypy: exit 0, ruff: exit 0, ruff_format: exit 0, typos: exit 0.

AI assistance: implementation and independent review used Hermes with GLM 5.3. Test evidence was reproduced in clean checkouts. This does not represent a human review.

@sylvesterkaczmarek sylvesterkaczmarek left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The core assumption here conflicts with SPDX's case-sensitivity rules. For user-defined identifiers, the variable part after LicenseRef- is case-insensitive, so LicenseRef-Name and LicenseRef-name refer to the same license rather than two distinct licenses. Preserving each occurrence's casing as a distinct identifier can therefore change the meaning of an expression. Could case variants canonicalize consistently to one identifier instead, with a regression covering mixed-case occurrences?

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants