Skip to content

Repository files navigation

🕷️ Production-Grade Python Web Scraper & Directory Extractor

Turn any public catalog, directory, or listing site into a clean, structured dataset — in minutes, not days.

Python 3.10+ Pandas BeautifulSoup Pydantic CI Code style: black Ruff License: MIT


📌 What this is

A battle-tested reference implementation for a common freelance request: "extract everything from this catalog/directory into an Excel or CSV file." This isn't a quick-and-dirty script — it's built the way a production data pipeline would be built, just scoped down to a fast, fixed-price deliverable.

I use this exact architecture for client projects. It's designed to be cloned, pointed at a new target site, and shipped within a 1-2 hour engagement.

The demo target is books.toscrape.com, a public sandbox site built specifically for practicing web scraping.


⚡ Key Features

  • 🔁 Automatic retry with exponential backoff — transient network errors (timeouts, connection failures) are retried automatically via tenacity; permanent errors (404, 403) fail fast instead of hanging.
  • 🛡️ Never crashes on malformed HTML — every field extraction uses safe fallbacks, so a missing <span> never throws AttributeError: 'NoneType' object has no attribute 'text' (the #1 failure mode of amateur scrapers).
  • ✅ Schema-validated output — every record is validated through a Pydantic model before export, guaranteeing correct types (no "$19.99 " strings pretending to be prices, no broken URLs).
  • 🔗 Correct relative URL resolution — links are resolved against the page URL with urljoin, so paginated pages under subdirectories (e.g. /catalogue/page-2.html) produce working absolute URLs, not silent 404s.
  • 🔍 Detail-page enrichment — fields missing from listing pages (like category, which only exists in detail-page breadcrumbs) are fetched per record with --enrich (default). A failed detail fetch logs a warning and keeps the listing data — one bad page never kills the run.
  • 📊 Clean exports out of the box — CSV (UTF-8 with BOM, Excel-compatible) and formatted .xlsx (bold headers, autosized columns) with zero manual cleanup needed.
  • 🕵️ Polite by design — rotating realistic User-Agents, configurable delay between requests, and respect for basic scraping etiquette.
  • 🧪 Actually tested — unit tests for the HTTP client (mocked retries/failures via respx), parser (HTML fixtures), exporter, and the orchestration layer. No live network needed in CI.
  • ⚙️ One-command CLI — no need to touch the code to run a new extraction.

📋 Data Sample Preview

Real output from data/sample_output.csv, generated by running this exact scraper against a public book catalog:

Title Category Price (£) Availability Rating URL
A Light in the Attic Poetry 51.77 In stock 3 link
Tipping the Velvet Historical Fiction 53.74 In stock 1 link
Soumission Fiction 50.10 In stock 1 link
Sharp Objects Mystery 47.82 In stock 4 link
Sapiens: A Brief History of Humankind History 54.23 In stock 5 link

📁 Full dataset (40 records, 2 pages): data/sample_output.csv


🚀 Quickstart (under 2 minutes)

# 1. Clone and enter the project
git clone https://github.com/kanderson-ai-dev/python-web-scraper-template.git
cd python-web-scraper-template

# 2. Install dependencies
python -m venv .venv && source .venv/bin/activate  # Windows: .venv\Scripts\activate
pip install -r requirements.txt

# 3. Run the scraper
python main.py scrape --pages 2 --output data/products --format both

That's it — you'll get data/products.csv and data/products.xlsx, both clean and ready to open.

Run python main.py scrape --help for all options:

Option Default Description
--pages, -p 1 Number of listing pages to scrape
--output, -o data/products Output file path (extension optional)
--format, -f csv csv, xlsx, or both
--delay, -d 1.0 Delay between requests in seconds
--enrich / --no-enrich --enrich Fetch detail pages to fill fields missing from listings (e.g. category)

⏱️ With --enrich on, each listing page costs 1 + N requests (one per record). Use --no-enrich for a faster, lighter scrape when you don't need detail-only fields.


🏗️ Project Structure

python-web-scraper-template/
├── .github/workflows/
│   └── ci.yml             # Lint + tests on every push/PR
├── scraper/
│   ├── client.py          # Resilient HTTP client (retries, backoff, UA rotation)
│   ├── parser.py          # BeautifulSoup + lxml parsing with safe fallbacks
│   ├── models.py          # Pydantic schema — guarantees clean, typed output
│   ├── exporter.py        # Pandas-powered CSV/Excel export
│   ├── scraper.py         # Pagination + validation orchestration
│   ├── config.py          # Centralized configuration
│   └── logging_setup.py   # Console + rotating-file logging
├── tests/                 # Unit tests — fixtures + mocked HTTP, no live network
├── data/
│   └── sample_output.csv  # Real sample output, committed for transparency
├── main.py                # CLI entry point (Typer + Rich)
├── requirements.txt       # Runtime dependencies
└── requirements-dev.txt   # Dev/test tooling (ruff, black, pytest, respx)

🧠 Why this matters for your project

Most quick scraping scripts break the moment a site has one missing field, one slow response, or one blocked request. This one is built to not do that:

  • Network hiccup? → retried automatically with exponential backoff.
  • Missing HTML element on one row? → logged and skipped, not a crashed script.
  • Weird price formatting ("£19.99 ", "$1,299")? → normalized before it ever reaches your spreadsheet.
  • Field only exists on detail pages? → fetched per-record, with per-record failure isolation.
  • One record fails validation? → logged and skipped; the other 999 rows still export.

That reliability is exactly what turns a one-off script into a repeat client relationship.


🛠️ Development

pip install -r requirements-dev.txt

ruff check .        # lint
black --check .     # format check
pytest              # test suite (no network required)

CI runs the same three checks on every push and pull request across Python 3.10–3.12.


💼 Need this for your own project?

I build custom scrapers for directories, catalogs, and e-commerce sites — delivering clean datasets and production-ready code, typically within 24 hours.

If you have a specific site you need extracted (with pagination, login, anti-bot measures, or scheduled monitoring), I can adapt this exact architecture to your use case.


📄 License

MIT — free to use as a reference or starting point for your own projects.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages