Ten Python examples for pages that build their content with JavaScript, ordered the way the decision actually goes, find the data before you reach for a browser. They follow our dynamic content scraping guide, and the study behind its advice is included.
Python 3.10 or newer.
pip install requests beautifulsoup4 lxml playwright selenium
playwright install chromiumex09 reads the API key from the HASDATA_API_KEY environment variable, the rest need no account anywhere.
ex01_probe.py checks what the raw HTML already carries. ex02_inline_data.py reads embedded JSON straight from the page source. ex03_endpoint_years.py and ex04_endpoint_pages.py call the JSON endpoint the page itself uses, by year and by page. ex05 through ex08 bring in the browser where nothing lighter works, Playwright waits, Selenium waits, in-page JavaScript, and infinite scroll. ex09_hasdata_api.py renders through an API instead of a local browser, and ex10_session_token.py replays an endpoint that wants the page's own session token first.
study/survey.py visited 51 JS-heavy pages and recorded what a browser sees, what plain HTTP sees, what sits inline in the source, and whether the page's largest JSON endpoint replays without cookies. study/results/survey_final.json holds the verdicts. 25 of 51 pages were server-rendered after all (11 of them with inline JSON on top), 8 more carried the data in inline JSON, 4 exposed an open or header-gated JSON endpoint, 10 blocked both routes, and 4 were skipped for robots.txt. Only 16 of the 51 were JS-dependent at all.
Every bar's rows are in study/results/survey_final.json with the per-page evidence.
study/timing.py collects the same 100 quotes through four routes on a sandbox site, and study/results/timing.json keeps the cumulative per-page medians of 3 runs. The endpoint route finishes in 3.2 seconds sequential and 1.0 second with 10 concurrent requests, server-rendered HTML parses in 2.9, and the browser routes trail far behind on the same data.
The curves come straight from study/results/timing.json.
The examples and studies fetch publicly available pages, and the survey respects robots.txt, which is why four of its pages went unmeasured. Whether and how such collection is appropriate depends on jurisdiction, the site, and the use, and nothing in this repository is legal advice. Is Web Scraping Legal? covers how we think about the question.
- How to Scrape Dynamic Content in Python, the guide these examples follow
- Web Scraping Using Selenium Python, when the browser route is the right one


