Open PowerShell as Administrator and run:
powershell -ExecutionPolicy ByPass -c "irm https://astral.sh/uv/install.ps1 | iex"Close and reopen your terminal, then verify: uv --version
- Open terminal in the
ao_scraperfolder - Install dependencies (UV will automatically get Python 3.12):
uv sync- Install VS Code from https://code.visualstudio.com
- Install the Python extension
- Open the
ao_scraperfolder in VS Code - Press
Ctrl+Shift+P, type "Python: Select Interpreter" - Choose the interpreter ending with
ao_scraper\.venv\Scripts\python.exe - Open terminal in VS Code:
Terminal > New Terminal - Run the scraper:
uv run python scraper.py
Create urls.csv in the project folder:
URL
https://ao.com/product/wth485001gb-bosch-serie-6-heat-pump-tumble-dryer-white-391297
https://ao.com/product/dv90cgc0a0abeu-samsung-series-5-optimaldrya-heat-pump-tumble-dryer-black-481062uv run python scraper.pyThe scraper will:
- Visit each URL in your CSV file
- Save product data as JSON files in
scraped/folder - Skip URLs already scraped (tracked in
scraped_urls.json)
After scraping, convert all JSON files to a single CSV:
uv run python converter.pyThis creates output/products.csv with columns:
- sku - Product model number (extracted from specifications)
- scraped_at - Timestamp when scraped
- url - Original product URL
- title - Product name
- non_member_price - Regular price (numeric, e.g., 599.00)
- ao_member_price - Member price (numeric, e.g., 549.00)
- product_summary - Product description
- All specification fields - Dynamic columns like "Dimensions - Height", "Features - Energy Rating", etc.
crawling:
max_concurrency: 3 # Browser tabs open at once (1-5)
request_interval_min: 2.0 # Min seconds between requests
request_interval_max: 5.0 # Max seconds between requests
page_timeout: 30 # Seconds to wait for page load
wait_after_load: 3 # Extra wait after page loads
max_retries: 3 # Retry failed pages
retry_delay: 5 # Seconds between retries
browser:
headless: true # false = see browser window
rotate_user_agent: true # Randomize browser identity
stealth_mode: true # Hide automation markers
output:
directory: "scraped" # Where JSON files are saved
pretty_print: true # Format JSON nicely
state:
state_file: "scraped_urls.json" # Tracks completed URLs
skip_existing: true # Skip already scraped URLsinput:
directory: "scraped" # Read JSON files from here
output:
file: "output/products.csv" # Save CSV here- Set
headless: falseto see what's happening - Reduce
max_concurrencyto 1 - Increase
request_interval_minto 5-10 seconds - Take a break - try again in a few hours
- Page structure may have changed
- Check if you're on the right page (set
headless: false) - Product might be unavailable
Already handled - the scraper automatically tries different encodings
Restart your terminal after installing UV, or use full path:
%USERPROFILE%\.local\bin\uv.exe run python scraper.pyao_scraper/
│
├── scraper.py # Main scraper
├── converter.py # JSON to CSV converter
├── scraper_config.yaml # Scraper settings
├── converter_config.yaml # Converter settings
├── urls.csv # Your input URLs
├── scraped_urls.json # Tracks completed URLs (auto-created)
│
├── scraped/ # JSON output files
│ ├── WTH485001GB.json
│ ├── DV90CGC0A0ABEU.json
│ └── ...
│
├── output/ # CSV output
│ └── products.csv
│
├── tests/ # Test files
└── pyproject.toml # Dependencies
Just run uv run python scraper.py again, it automatically skips completed URLs.
Delete scraped_urls.json to re-scrape everything.
Edit urls.csv and run the scraper, it only processes new URLs.
The converter automatically finds ALL specification fields across all products and creates columns for each.
- Create a test CSV with just that URL
- Set
headless: falsein config.yaml - Run and watch the browser
uv run pytestEdit converter.py to add custom fields
urls.csv→scraper.py→scraped/*.jsonscraped/*.json→converter.py→output/products.csv
scraped_urls.json contains:
[
"https://ao.com/product/abc123",
"https://ao.com/product/def456"
]