Repository navigation
Exclude scripts, styles and chrome from word_count - #103
Open
amedipiran wants to merge 1 commit into
Open
amedipiran wants to merge 1 commit into
amedipiran wants to merge 1 commit into
Conversation
soup.get_text() on the whole document counted the contents of <script> and <style> as words, along with the navigation, header and footer that repeat on every page. On a site with a large menu that is over a thousand words of chrome on every page. Measured on one WooCommerce product page: 1333 words before, 559 after. This quietly broke everything built on word_count. The low-content check fires below 300 words and never fired at all, and word_count carries 0.10 weight in near-duplicate detection. Counts on a copy so the caller's soup keeps its links, headings and images.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
seo_extractor.pycounts words withsoup.get_text()on the whole document. That includes the contents of<script>and<style>, plus the navigation, header and footer that repeat on every page.On a WooCommerce site with a large mega menu I measured 1333 words on a product page, of which 774 were chrome and inline JS. The homepage went from 1250 to 461.
This quietly breaks two things that depend on
word_count:issue_detectorflags low content below 300 words. On a site with a decent menu that threshold can never be reached, so the check silently never fires.word_countcarries 0.10 weight in near-duplicate detection, computed on a number that is mostly boilerplate.The fix strips
script,style,noscript,template,nav,header,footerandaside, then prefers<main>over<body>. It works on acopy.copyof the soup, so the caller keeps using the same soup afterwards for links, headings and images. I checked that: after counting, the link count on the original soup is unchanged.Worth noting this will lower
word_countacross the board, so any thresholds people have tuned against the old numbers will behave differently. That seemed better than leaving the number meaningless, but your call.Generated with Claude Code