Skip to content

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Dynamic Content Scraping Examples (Python)

HasData, the web scraping API one example calls

Ten Python examples for pages that build their content with JavaScript, ordered the way the decision actually goes, find the data before you reach for a browser. They follow our dynamic content scraping guide, and the study behind its advice is included.

Table of Contents

Requirements

Python 3.10 or newer.

pip install requests beautifulsoup4 lxml playwright selenium
playwright install chromium

ex09 reads the API key from the HASDATA_API_KEY environment variable, the rest need no account anywhere.

The Examples

ex01_probe.py checks what the raw HTML already carries. ex02_inline_data.py reads embedded JSON straight from the page source. ex03_endpoint_years.py and ex04_endpoint_pages.py call the JSON endpoint the page itself uses, by year and by page. ex05 through ex08 bring in the browser where nothing lighter works, Playwright waits, Selenium waits, in-page JavaScript, and infinite scroll. ex09_hasdata_api.py renders through an API instead of a local browser, and ex10_session_token.py replays an endpoint that wants the page's own session token first.

The Survey

study/survey.py visited 51 JS-heavy pages and recorded what a browser sees, what plain HTTP sees, what sits inline in the source, and whether the page's largest JSON endpoint replays without cookies. study/results/survey_final.json holds the verdicts. 25 of 51 pages were server-rendered after all (11 of them with inline JSON on top), 8 more carried the data in inline JSON, 4 exposed an open or header-gated JSON endpoint, 10 blocked both routes, and 4 were skipped for robots.txt. Only 16 of the 51 were JS-dependent at all.

Stacked bar chart of 51 surveyed pages by category, showing server-rendered HTML, inline JSON, open endpoints, one endpoint that needs headers, blocked pages, and pages skipped for robots.txt

Every bar's rows are in study/results/survey_final.json with the per-page evidence.

The Timing Study

study/timing.py collects the same 100 quotes through four routes on a sandbox site, and study/results/timing.json keeps the cumulative per-page medians of 3 runs. The endpoint route finishes in 3.2 seconds sequential and 1.0 second with 10 concurrent requests, server-rendered HTML parses in 2.9, and the browser routes trail far behind on the same data.

Line chart of cumulative seconds per page for four ways of fetching the same 100 quotes, the concurrent endpoint path finishing in about one second and the page-by-page browser path in about nine

The curves come straight from study/results/timing.json.

Disclaimer

The examples and studies fetch publicly available pages, and the survey respects robots.txt, which is why four of its pages went unmeasured. Whether and how such collection is appropriate depends on jurisdiction, the site, and the use, and nothing in this repository is legal advice. Is Web Scraping Legal? covers how we think about the question.

More Resources

About

Ten Python examples for pages that build their content with JavaScript, from reading embedded JSON and the page's own endpoints to Playwright, Selenium and infinite scroll. Includes the survey and timing study behind the guide.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages