Skip to content

feat(leads): step through paginated directories - #145

Merged
ralyodio merged 1 commit into
masterfrom
feat/directory-pagination
Jul 28, 2026
Merged

ralyodio merged 1 commit into
masterfrom
feat/directory-pagination

Conversation

@ralyodio

Copy link
Copy Markdown
Contributor

A listing page is a fraction of the directory behind it, and a fraction reported as a total is indistinguishable from a small directory. Seeds now walk the pager before opening anything.

URL-following first, so fetch-first survives

Following the next page by URL is tried before anything else, so an ordinary paginated directory costs one cheap request per page rather than one browser render per page. Clicking a JS-only control is the expensive fallback.

Ordered by how much the page is asserting

Signal Why it ranks there
rel="next" the site stating it outright
link text / aria-label a convention
increment ?page=N a guess — and one that can loop

Link text is matched whole — a link reading "next" is a pager, one reading "next steps" is prose. aria-label is matched on the word, because labels are written to be read aloud ("Next page").

The page-parameter guess only fires when the URL already carries one. Inventing ?page=2 asks every site for a page that may not exist. Disabled controls are refused — a greyed-out arrow means last page, and following it loops.

The guard a live run made necessary

The guess needed a second stop condition that no unit test would have surfaced. A site that ignores its page parameter answers every guess with the same page — so the walk fetched 21 identical pages before hitting the cap.

A page contributing no new entries is now the end of the list, whatever its pager claims:

before: walked 21 listing pages → 12 people
after:  walked 2 listing pages  → 12 people

Checks

  • tsc --noEmit clean
  • 907/907 tests pass, 17 new
  • production build compiles
  • verified live against a real directory

🤖 Generated with Claude Code

A listing page is a fraction of the directory behind it, and a fraction
reported as a total is indistinguishable from a small directory. Seeds
now walk the pager before opening anything.

Following the next page by URL is tried before anything else, so the
fetch-first path still applies and an ordinary paginated directory costs
one cheap request per page rather than one browser render per page.

The three ways of finding that URL are ordered by how much the page is
actually asserting. rel="next" is the site stating it outright. Link text
is a convention, and is matched whole — a link reading "next" is a pager,
one reading "next steps" is prose — while an aria-label is matched on the
word, because labels are written to be read aloud and say "Next page".
Incrementing a page parameter is a guess, so it comes last and only when
the URL already carries one; inventing ?page=2 asks every site on the
internet for a page that may not exist.

Disabled controls are refused, since a greyed-out arrow means this is the
last page and following it loops.

The guess needed a second guard that only a live run revealed. A site that
ignores its page parameter answers every guess with the same page, and the
walk dutifully fetched twenty-one identical pages before stopping at the
cap. A page that contributes no new entries is now the end of the list
whatever its pager claims, which took the same run from twenty-one page
loads to two.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown

vu1nz Security Review

0 finding(s) in PR #?

No security issues found.

@ralyodio
ralyodio merged commit 2f23f8e into master Jul 28, 2026
8 checks passed
@ralyodio
ralyodio deleted the feat/directory-pagination branch July 28, 2026 08:59
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant