Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

anchor-check

Crawl a site and find the #fragment links that lead nowhere. It checks the id (or a[name]) the link points to in the HTML the server actually sent, in-page and across pages, and it also looks at skip-links, duplicate ids and aria-* / label[for] references. Standard library only; Playwright is optional.

python -m anchor_check https://example.com/                 # text report
anchor-check https://example.com/ --format markdown -o anchors.md
anchor-check https://example.com/ --render --fail-on warning   # CI: exit 1 on warnings too

Why: a link like /pricing#faq keeps working as a URL after someone renames the heading, it just silently stops scrolling anywhere. Nothing in a normal link checker notices, because the page itself still answers 200.

What it reports

Kind Severity Meaning
broken-anchor error #id link (in-page or to another page of the site) whose target has no element with that id or a[name]
skip-link-missing error a skip-link (text/class such as "Skip to content", "Saltar al contenido", "Aller au contenu") whose target does not exist
skip-link-not-focusable warning the skip-link target exists but is not interactive and has no tabindex. Modern browsers move the focus starting point to the target anyway, so this often works; older browsers and some assistive technology do not. tabindex="-1" makes it reliable. Hence a warning, not an error
duplicate-id warning the same id more than once on a page (a fragment goes to the first one)
aria-ref-missing warning aria-controls, aria-labelledby or aria-describedby naming an id that is not in the served HTML
label-for-missing warning label[for] naming a missing id
target-unavailable warning the linked page answered 4xx/5xx, so the fragment cannot be checked
runtime-only-id info with --render: the id is missing from the served HTML but exists once scripts have run

Rules from the HTML spec: fragments are percent-decoded before matching, an empty # and #top (any case) are always valid, and text fragments (#:~:text=) and client-side routes (#/app, #!route) are accepted without a lookup. <base href> is honoured, because browsers resolve #x against it. Content inside <template> is ignored.

How it crawls

  • Same host only, at most 30 pages (--max-pages), and at least 1 s between requests (--delay, lower it only for servers you run). A Crawl-delay in robots.txt is honoured up to 10 s.
  • User-Agent: anchor-check/1.0 (+https://github.com/dreadmoreeee/anchor-check). Only GET requests, no forms, no cookies, no clicks.
  • robots.txt is read with an RFC 9309 matcher (anchor_check/rfc9309.py: the group of our product token, longest matching rule wins, Allow wins ties, * and $, 4xx means allowed, 5xx or unreachable means disallowed). urllib.robotparser is not used because it applies the first matching rule.
  • /cdn-cgi/ links are skipped (Cloudflare's email-protection#... links point at nothing you own). Non-HTML files (pdf, images, css, js...) are not fetched.
  • Queue order when the cap is tight: hreflang alternates (/en/, /fr/, /es/) first, then pages that some #fragment link points to, then everything else. A target that did not fit under the cap is listed under "Not checked", never reported as broken.
  • Redirects are followed by hand and only within the host; each hop is throttled and checked against robots.txt.

--render re-opens only the pages where an id was missing, in headless Chromium (pip install playwright && python -m playwright install chromium), and moves ids that exist after scripts ran from error/warning to info (runtime-only-id). Chromium loads the page's own css, js and images like any browser; nothing is clicked or typed.

Options and exit codes

Option Default
--max-pages N 30 pages to fetch
--delay SECONDS 1.0 minimum gap between requests
--timeout SECONDS 15 per request
--render off re-check missing ids in Chromium
-f, --format text|markdown|json text
-o, --output FILE stdout
--fail-on error|warning|info|none error lowest severity that gives exit code 1
-v list every crawled page (text)

Exit codes: 0 nothing at or above --fail-on, 1 findings at or above it, 2 usage or network error (start page unreachable, blocked by robots.txt, not HTML, --render without Chromium, bad option). A CI example is in examples/ci-github-actions.yml.

Measured result

Run on 2026-09-29 against my own three sites, default settings (30 pages, 1 s between requests, GET only). Each command was run twice (once for the text report, once for the JSON one), so each site saw the crawl two times.

$ python -m anchor_check https://marvin.demarkstudio.ca/ -v
anchor-check 1.0.0  https://marvin.demarkstudio.ca/
6 page(s) crawled (cap 30), 24 fragment link(s): 24 checked, 0 not checked; 0 special (#, #top, routes) accepted
note: robots.txt found

  200  /
  200  /es/
  200  /work/demark-studio/
  200  /work/triple-a-woodworks/
  200  /work/stopping-form-spam/
  200  /work/auditing-my-own-sites/

WARNING skip-link-not-focusable  /:29
        skip link "Skip to content" targets <main> #main, which is not focusable (no tabindex, not interactive). Modern browsers move [...]
WARNING skip-link-not-focusable  /es/:29
        skip link "Ir al contenido" targets <main> #main, which is not focusable (no tabindex, not interactive). Modern browsers move [...]

Summary: 0 error(s), 2 warning(s), 0 info

$ python -m anchor_check https://demarkstudio.ca/en/ -v
anchor-check 1.0.0  https://demarkstudio.ca/en/
29 page(s) crawled (cap 30), 403 fragment link(s): 403 checked, 0 not checked; 1 special (#, #top, routes) accepted
note: robots.txt found
note: page cap reached; raise --max-pages to crawl more
[... 29 page lines: /en/, /en/websites, /en/product, ..., /docs/api, /docs/api?lang=fr ...]

ERROR   broken-anchor            /en/for/contractors:155
        cross-page link /en/product#trabajos: no element with id="trabajos" (or a[name]) in https://demarkstudio.ca/en/product
ERROR   broken-anchor            /en/for/contractors:164
        cross-page link /en/product#trabajo: no element with id="trabajo" (or a[name]) in https://demarkstudio.ca/en/product
ERROR   broken-anchor            /en/for/contractors:173
        cross-page link /en/product#presupuesto-firma: no element with id="presupuesto-firma" (or a[name]) in ...
ERROR   broken-anchor            /en/for/contractors:182
        cross-page link /en/product#orden-cambio: no element with id="orden-cambio" (or a[name]) in ...
ERROR   broken-anchor            /en/for/contractors:191
        cross-page link /en/product#debito-bancario: no element with id="debito-bancario" (or a[name]) in ...
ERROR   broken-anchor            /en/for/contractors:200
        cross-page link /en/product#exportar: no element with id="exportar" (or a[name]) in ...
WARNING skip-link-not-focusable  /docs/api:44
        skip link "Skip to content" targets <main> #contenido, which is not focusable (no tabindex, not interactive). Modern browsers move [...]
[... 26 more skip-link-not-focusable warnings, one per page, same <main id="contenido"> ...]
WARNING duplicate-id             /en/for/salons:220
        id="titulo-calculadora" is used 2 times (<h2> line 210, <h2> line 220); a fragment link goes to the first one

Summary: 6 error(s), 28 warning(s), 0 info

$ python -m anchor_check https://demo.demarkstudio.ca/riverside-pub/site -v
anchor-check 1.0.0  https://demo.demarkstudio.ca/riverside-pub/site
29 page(s) crawled (cap 30), 61 fragment link(s): 60 checked, 1 not checked; 0 special (#, #top, routes) accepted
note: robots.txt found
note: page cap reached; raise --max-pages to crawl more
[... 29 page lines: /riverside-pub/site, /riverside-pub/site/servicios, ..., /en/your-data ...]

WARNING skip-link-not-focusable  /en/:43
        skip link "Skip to content" targets <main> #contenido, which is not focusable (no tabindex, not interactive). Modern browsers move [...]
[... 26 more skip-link-not-focusable warnings, on <main id="contenido"> and <main id="principal"> ...]
WARNING duplicate-id             /en/for/salons:220
        id="titulo-calculadora" is used 2 times (<h2> line 210, <h2> line 220); a fragment link goes to the first one

Not checked (1):
  /en/ -> /en/websites#plantillas: page cap reached

Summary: 0 error(s), 29 warning(s), 0 info

(Long messages are shortened with [...] here; the tool prints the full text. The [...] lines are my trimming, everything else is verbatim output.)

What it found, honestly:

  • 6 real broken links on /en/for/contractors: six links to /en/product#... (#trabajos, #trabajo, #presupuesto-firma, #orden-cambio, #debito-bancario, #exportar). I checked by hand: /en/product has no such ids in the served HTML nor after Chromium ran its scripts (it has #presupuestos, #facturas, #cobros, #solicitudes, ...), so the links land at the top of the page instead of the section. These are mine to fix.
  • One duplicate id (titulo-calculadora, two <h2> on /en/for/salons) that a shared template also carries into the demo site's /en/for/salons.
  • Skip-link warnings everywhere: every page's skip-link targets a <main> without tabindex="-1". This is a warning on purpose: current browsers usually continue Tab from the target anyway; adding tabindex="-1" makes it dependable.
  • No missing aria-controls / aria-labelledby / aria-describedby / label[for] targets on any of the three sites, and no broken in-page anchors.
  • marvin.demarkstudio.ca has 0 errors; its only findings are the two skip-link warnings.
  • The two larger sites hit the 30-page cap, so pages beyond it were not looked at; one link target was left unchecked for that reason and is listed as such.

Fixed on my portfolio the same morning. The two warnings were right: the skip link pointed at <main>, which could not take focus. <main id="main" tabindex="-1"> makes the jump reliable for keyboard and screen-reader users in every browser. Redeployed and checked again:

$ python -m anchor_check https://marvin.demarkstudio.ca/
Summary: 0 error(s), 0 warning(s), 0 info

Install

Python 3.10+.

pip install .                 # `anchor-check` command, no dependencies
pip install ".[render]"       # adds Playwright for --render
python -m pytest -q -p no:cacheprovider --import-mode=importlib

Tests use a local http.server fixture site on a random free port and need no internet. The --render test is skipped when Chromium is not installed.

68 passed in 12.02s

Limitations

  • It reads the HTML as served (plus what --render adds). A page that needs a login, or ids that appear only after a click, cannot be seen.
  • Only <a>/<area> links are followed; the skip-link detection is a text/class heuristic (skip, "jump to", "saltar", "aller au contenu", ...), so a skip-link labelled some other way is checked as an ordinary anchor.
  • Redirects to another host, mailto:, external sites and non-HTML files are out of scope by design.
  • aria-* references that a script adds later are warnings unless you run --render.

Author

Marvin Palencia, founder of DeMark Studio, Miramichi, New Brunswick, Canada. Portfolio: marvin.demarkstudio.ca

MIT License.

About

Find broken #fragment links across a site, skip links that go nowhere, duplicate ids and aria references to missing ids, optionally after JavaScript runs.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages