Skip to content

Exclude scripts, styles and chrome from word_count - #103

Open
amedipiran wants to merge 1 commit into
PhialsBasement:mainfrom
amedipiran:fix/word-count
Open

amedipiran wants to merge 1 commit into
PhialsBasement:mainfrom
amedipiran:fix/word-count

Conversation

@amedipiran

Copy link
Copy Markdown

seo_extractor.py counts words with soup.get_text() on the whole document. That includes the contents of <script> and <style>, plus the navigation, header and footer that repeat on every page.

On a WooCommerce site with a large mega menu I measured 1333 words on a product page, of which 774 were chrome and inline JS. The homepage went from 1250 to 461.

This quietly breaks two things that depend on word_count:

  • issue_detector flags low content below 300 words. On a site with a decent menu that threshold can never be reached, so the check silently never fires.
  • word_count carries 0.10 weight in near-duplicate detection, computed on a number that is mostly boilerplate.

The fix strips script, style, noscript, template, nav, header, footer and aside, then prefers <main> over <body>. It works on a copy.copy of the soup, so the caller keeps using the same soup afterwards for links, headings and images. I checked that: after counting, the link count on the original soup is unchanged.

Worth noting this will lower word_count across the board, so any thresholds people have tuned against the old numbers will behave differently. That seemed better than leaving the number meaningless, but your call.

Generated with Claude Code

soup.get_text() on the whole document counted the contents of <script> and
<style> as words, along with the navigation, header and footer that repeat
on every page. On a site with a large menu that is over a thousand words of
chrome on every page.

Measured on one WooCommerce product page: 1333 words before, 559 after.

This quietly broke everything built on word_count. The low-content check
fires below 300 words and never fired at all, and word_count carries 0.10
weight in near-duplicate detection.

Counts on a copy so the caller's soup keeps its links, headings and images.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant