Skip to content

perf(keys): use binary GCD in the batch-GCD shared-prime pass - #20

Merged
nmatt0 merged 1 commit into
masterfrom
perf/keys-binary-gcd
Sep 15, 2026
Merged

nmatt0 merged 1 commit into
masterfrom
perf/keys-binary-gcd

Conversation

@nmatt0

@nmatt0 nmatt0 commented Sep 15, 2026

Copy link
Copy Markdown
Owner

Summary

Fixes the --keys pass hanging for minutes on a firmware rootfs that ships a
full CA certificate store. Closes #19.

The cross-file batch-GCD shared-prime check runs a pairwise BigUint::gcd over
every RSA modulus in the tree. The O(n^2) loop is capped (256 keys) and
expected; the cost was per-gcd. BigUint::gcd was Euclidean (a % b), and
operator% routes through the bit-by-bit long-division divmod, which iterates
over every bit of the dividend, allocates fresh vectors per bit, and computes a
quotient it then discards. A single 2048-bit gcd cost ~37 ms, so a standard
/etc/ssl/certs bundle (hundreds of RSA CA certs) produced enough moduli to
turn the pass into a multi-minute hang.

Change

Replace the Euclidean BigUint::gcd with Stein's binary GCD (shifts and
subtraction only, no division), keeping divmod off the hot path.

This is a constant-factor fix, not an asymptotic one — both forms are
O(bits^2) worst case. The win is removing the per-bit allocation churn that made
each gcd pathologically slow. The O(n^2) pairwise loop and its key-count / bit
caps are unchanged, so the DoS bound on a hostile tree is unaffected.

Effect

The --keys pass on a rootfs with a full CA store drops from 6+ minutes to
~5s; the remaining time is the now-cheap O(n^2) loop. Quadratic shape before the
fix, for reference (whole --keys pass over N certs):

N=20 ->  7.0s
N=40 -> 29.1s
N=60 -> 50.0s

Testing

  • Full unit suite green (1339 checks), including the bigint gcd KATs and the
    Fermat / batch-GCD recovery KATs.
  • tests/run.sh integration suite green, including the shared-prime recovery
    fixture — the detector still recovers real shared factors.
  • On a rootfs carrying a full CA store the pass now returns in ~5s and reports
    only the expected weak-modulus findings, with no spurious shared-prime hits.

Follow-up (not in this PR)

If the O(n^2) loop ever needs to scale to very large key counts, a product-tree
batch-GCD gives near-linear big-integer work. That is a larger change and not
needed to fix this hang.

The cross-file batch-GCD check runs a pairwise BigUint::gcd over every RSA
modulus in the tree. The O(n^2) loop is capped and expected; the cost was
per-gcd. BigUint::gcd was Euclidean (a % b), and operator% goes through the
bit-by-bit long-division divmod, which iterates over every bit of the
dividend, allocates fresh vectors per bit, and computes a quotient it then
discards. A single 2048-bit gcd cost ~37ms, so a rootfs shipping a full CA
certificate store (hundreds of RSA CA certs) turned the pass into a
multi-minute hang.

Replace it with Stein's binary GCD (shifts and subtraction, no division),
keeping divmod off the hot path. Same O(bits^2) worst case, but without the
per-bit allocation churn: the --keys pass on such a tree drops from 6+ minutes
to ~5s, and the remaining time is the now-cheap O(n^2) loop.

Correctness is unchanged: the bigint gcd KATs and the shared-prime recovery
fixtures still pass, and the detector still recovers real shared factors.

Closes #19
@nmatt0
nmatt0 merged commit 9d8e320 into master Sep 15, 2026
4 checks passed
@nmatt0
nmatt0 deleted the perf/keys-binary-gcd branch September 15, 2026 17:17
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

--keys batch-GCD hangs for minutes on a rootfs with a full CA certificate store

1 participant