Turn any public catalog, directory, or listing site into a clean, structured dataset — in minutes, not days.
A battle-tested reference implementation for a common freelance request: "extract everything from this catalog/directory into an Excel or CSV file." This isn't a quick-and-dirty script — it's built the way a production data pipeline would be built, just scoped down to a fast, fixed-price deliverable.
I use this exact architecture for client projects. It's designed to be cloned, pointed at a new target site, and shipped within a 1-2 hour engagement.
The demo target is books.toscrape.com, a public sandbox site built specifically for practicing web scraping.
- 🔁 Automatic retry with exponential backoff — transient network errors (timeouts, connection failures) are retried automatically via
tenacity; permanent errors (404, 403) fail fast instead of hanging. - 🛡️ Never crashes on malformed HTML — every field extraction uses safe fallbacks, so a missing
<span>never throwsAttributeError: 'NoneType' object has no attribute 'text'(the #1 failure mode of amateur scrapers). - ✅ Schema-validated output — every record is validated through a
Pydanticmodel before export, guaranteeing correct types (no"$19.99 "strings pretending to be prices, no broken URLs). - 🔗 Correct relative URL resolution — links are resolved against the page URL with
urljoin, so paginated pages under subdirectories (e.g./catalogue/page-2.html) produce working absolute URLs, not silent 404s. - 🔍 Detail-page enrichment — fields missing from listing pages (like
category, which only exists in detail-page breadcrumbs) are fetched per record with--enrich(default). A failed detail fetch logs a warning and keeps the listing data — one bad page never kills the run. - 📊 Clean exports out of the box — CSV (UTF-8 with BOM, Excel-compatible) and formatted
.xlsx(bold headers, autosized columns) with zero manual cleanup needed. - 🕵️ Polite by design — rotating realistic User-Agents, configurable delay between requests, and respect for basic scraping etiquette.
- 🧪 Actually tested — unit tests for the HTTP client (mocked retries/failures via
respx), parser (HTML fixtures), exporter, and the orchestration layer. No live network needed in CI. - ⚙️ One-command CLI — no need to touch the code to run a new extraction.
Real output from data/sample_output.csv, generated by running this exact scraper against a public book catalog:
| Title | Category | Price (£) | Availability | Rating | URL |
|---|---|---|---|---|---|
| A Light in the Attic | Poetry | 51.77 | In stock | 3 | link |
| Tipping the Velvet | Historical Fiction | 53.74 | In stock | 1 | link |
| Soumission | Fiction | 50.10 | In stock | 1 | link |
| Sharp Objects | Mystery | 47.82 | In stock | 4 | link |
| Sapiens: A Brief History of Humankind | History | 54.23 | In stock | 5 | link |
📁 Full dataset (40 records, 2 pages):
data/sample_output.csv
# 1. Clone and enter the project
git clone https://github.com/kanderson-ai-dev/python-web-scraper-template.git
cd python-web-scraper-template
# 2. Install dependencies
python -m venv .venv && source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install -r requirements.txt
# 3. Run the scraper
python main.py scrape --pages 2 --output data/products --format bothThat's it — you'll get data/products.csv and data/products.xlsx, both clean and ready to open.
Run python main.py scrape --help for all options:
| Option | Default | Description |
|---|---|---|
--pages, -p |
1 |
Number of listing pages to scrape |
--output, -o |
data/products |
Output file path (extension optional) |
--format, -f |
csv |
csv, xlsx, or both |
--delay, -d |
1.0 |
Delay between requests in seconds |
--enrich / --no-enrich |
--enrich |
Fetch detail pages to fill fields missing from listings (e.g. category) |
⏱️ With
--enrichon, each listing page costs 1 + N requests (one per record). Use--no-enrichfor a faster, lighter scrape when you don't need detail-only fields.
python-web-scraper-template/
├── .github/workflows/
│ └── ci.yml # Lint + tests on every push/PR
├── scraper/
│ ├── client.py # Resilient HTTP client (retries, backoff, UA rotation)
│ ├── parser.py # BeautifulSoup + lxml parsing with safe fallbacks
│ ├── models.py # Pydantic schema — guarantees clean, typed output
│ ├── exporter.py # Pandas-powered CSV/Excel export
│ ├── scraper.py # Pagination + validation orchestration
│ ├── config.py # Centralized configuration
│ └── logging_setup.py # Console + rotating-file logging
├── tests/ # Unit tests — fixtures + mocked HTTP, no live network
├── data/
│ └── sample_output.csv # Real sample output, committed for transparency
├── main.py # CLI entry point (Typer + Rich)
├── requirements.txt # Runtime dependencies
└── requirements-dev.txt # Dev/test tooling (ruff, black, pytest, respx)
Most quick scraping scripts break the moment a site has one missing field, one slow response, or one blocked request. This one is built to not do that:
- Network hiccup? → retried automatically with exponential backoff.
- Missing HTML element on one row? → logged and skipped, not a crashed script.
- Weird price formatting (
"£19.99 ","$1,299")? → normalized before it ever reaches your spreadsheet. - Field only exists on detail pages? → fetched per-record, with per-record failure isolation.
- One record fails validation? → logged and skipped; the other 999 rows still export.
That reliability is exactly what turns a one-off script into a repeat client relationship.
pip install -r requirements-dev.txt
ruff check . # lint
black --check . # format check
pytest # test suite (no network required)CI runs the same three checks on every push and pull request across Python 3.10–3.12.
I build custom scrapers for directories, catalogs, and e-commerce sites — delivering clean datasets and production-ready code, typically within 24 hours.
If you have a specific site you need extracted (with pagination, login, anti-bot measures, or scheduled monitoring), I can adapt this exact architecture to your use case.
MIT — free to use as a reference or starting point for your own projects.