Skip to content

Repository files navigation

AO Scraper

Quick Start

1. UV

Open PowerShell as Administrator and run:

powershell -ExecutionPolicy ByPass -c "irm https://astral.sh/uv/install.ps1 | iex"

Close and reopen your terminal, then verify: uv --version

2. Set Up

  1. Open terminal in the ao_scraper folder
  2. Install dependencies (UV will automatically get Python 3.12):
uv sync

3. VS Code

  1. Install VS Code from https://code.visualstudio.com
  2. Install the Python extension
  3. Open the ao_scraper folder in VS Code
  4. Press Ctrl+Shift+P, type "Python: Select Interpreter"
  5. Choose the interpreter ending with ao_scraper\.venv\Scripts\python.exe
  6. Open terminal in VS Code: Terminal > New Terminal
  7. Run the scraper: uv run python scraper.py

4. Input

Create urls.csv in the project folder:

URL
https://ao.com/product/wth485001gb-bosch-serie-6-heat-pump-tumble-dryer-white-391297
https://ao.com/product/dv90cgc0a0abeu-samsung-series-5-optimaldrya-heat-pump-tumble-dryer-black-481062

5. Run the Scraper

uv run python scraper.py

The scraper will:

  • Visit each URL in your CSV file
  • Save product data as JSON files in scraped/ folder
  • Skip URLs already scraped (tracked in scraped_urls.json)

6. Convert to Excel-Friendly CSV

After scraping, convert all JSON files to a single CSV:

uv run python converter.py

This creates output/products.csv with columns:

  • sku - Product model number (extracted from specifications)
  • scraped_at - Timestamp when scraped
  • url - Original product URL
  • title - Product name
  • non_member_price - Regular price (numeric, e.g., 599.00)
  • ao_member_price - Member price (numeric, e.g., 549.00)
  • product_summary - Product description
  • All specification fields - Dynamic columns like "Dimensions - Height", "Features - Energy Rating", etc.

Configuration Files

1. config.yaml (Scraper Settings)

crawling:
  max_concurrency: 3          # Browser tabs open at once (1-5)
  request_interval_min: 2.0   # Min seconds between requests
  request_interval_max: 5.0   # Max seconds between requests
  page_timeout: 30            # Seconds to wait for page load
  wait_after_load: 3          # Extra wait after page loads
  max_retries: 3              # Retry failed pages
  retry_delay: 5              # Seconds between retries

browser:
  headless: true              # false = see browser window
  rotate_user_agent: true     # Randomize browser identity
  stealth_mode: true          # Hide automation markers

output:
  directory: "scraped"        # Where JSON files are saved
  pretty_print: true          # Format JSON nicely

state:
  state_file: "scraped_urls.json"  # Tracks completed URLs
  skip_existing: true              # Skip already scraped URLs

2. converter_config.yaml (CSV Converter Settings)

input:
  directory: "scraped"        # Read JSON files from here

output:
  file: "output/products.csv" # Save CSV here

Troubleshooting

"Getting blocked by Cloudflare"

  1. Set headless: false to see what's happening
  2. Reduce max_concurrency to 1
  3. Increase request_interval_min to 5-10 seconds
  4. Take a break - try again in a few hours

"No data extracted from page"

  1. Page structure may have changed
  2. Check if you're on the right page (set headless: false)
  3. Product might be unavailable

"UnicodeDecodeError"

Already handled - the scraper automatically tries different encodings

"Can't find uv command"

Restart your terminal after installing UV, or use full path:

%USERPROFILE%\.local\bin\uv.exe run python scraper.py

Project Structure

ao_scraper/
│
├── scraper.py              # Main scraper
├── converter.py            # JSON to CSV converter
├── scraper_config.yaml             # Scraper settings
├── converter_config.yaml   # Converter settings
├── urls.csv                # Your input URLs
├── scraped_urls.json       # Tracks completed URLs (auto-created)
│
├── scraped/                # JSON output files
│   ├── WTH485001GB.json
│   ├── DV90CGC0A0ABEU.json
│   └── ...
│
├── output/                 # CSV output
│   └── products.csv
│
├── tests/                  # Test files
└── pyproject.toml          # Dependencies

Common Tasks

Resume After Interruption

Just run uv run python scraper.py again, it automatically skips completed URLs.

Start Fresh

Delete scraped_urls.json to re-scrape everything.

Add More URLs

Edit urls.csv and run the scraper, it only processes new URLs.

Extract Specific Fields

The converter automatically finds ALL specification fields across all products and creates columns for each.

Debug a Specific URL

  1. Create a test CSV with just that URL
  2. Set headless: false in config.yaml
  3. Run and watch the browser

Advanced Usage

Running Tests

uv run pytest

Custom Field Extraction

Edit converter.py to add custom fields

Technical Details

Data Flow

  1. urls.csv → scraper.py → scraped/*.json
  2. scraped/*.json → converter.py → output/products.csv

State Management

scraped_urls.json contains:

[
  "https://ao.com/product/abc123",
  "https://ao.com/product/def456"
]

About

ao_scraper

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages