A robust, modular web crawler designed for collecting datasets for GUI-graph + GNN-based dark pattern detection research. This crawler automatically captures full-page screenshots, viewport screenshots, and rich metadata for multi-step user flows across various website categories.
- Modular Flow System: State machine-based flows for different user interaction patterns
- Heuristic Selectors: No hardcoded selectors - uses intelligent text/role/aria-label matching
- Rich Metadata: JSON metadata for every screenshot to support later GUI graph construction
- Config-Driven: YAML-based site definitions for easy extensibility
- Retry Logic: Automatic retries with graceful failure handling
- Full-Page + Viewport Screenshots: Captures both full-page and viewport (1366x768) screenshots
- Extensible Architecture: Easy to add new flows and site categories
- Node.js 18+
- npm or yarn
- Playwright browsers (installed automatically)
# Install dependencies
npm install
# Install Playwright browsers
npx playwright install chromiumbtp-crawler/
├── src/
│ ├── core/
│ │ ├── BrowserManager.ts # Browser lifecycle management
│ │ ├── Flow.ts # Abstract base flow class
│ │ ├── FlowRunner.ts # Central orchestrator
│ │ └── ConfigLoader.ts # YAML config loader
│ ├── flows/
│ │ ├── FlowFactory.ts # Flow factory
│ │ ├── LandingCookieConsentFlow.ts
│ │ ├── ForcedSignupPaywallFlow.ts
│ │ ├── ProductCartCheckoutFlow.ts
│ │ ├── TrialFunnelFlow.ts
│ │ └── CancellationRoachMotelFlow.ts
│ ├── utils/
│ │ ├── logger.ts # Winston logger
│ │ ├── selectors.ts # Heuristic selector utilities
│ │ └── screenshot.ts # Screenshot capture utilities
│ ├── types/
│ │ └── index.ts # TypeScript type definitions
│ └── index.ts # Main entry point
├── configs/
│ ├── example_ecommerce.yaml
│ └── example_saas.yaml
├── dataset/ # Output directory (created automatically)
└── README.md
# Run all sites in configs/ directory (headless)
npm run crawl
# Run in non-headless mode (for debugging)
npm run dev -- --no-headless
# Run specific site config
npm run dev -- --config=configs/example_ecommerce.yaml
# Run all configs in a directory
npm run dev -- --config-dir=configs/# Run with TypeScript directly (faster iteration)
npm run dev- Create a YAML file in
configs/directory:
site_id: "site_003"
category: "ecommerce" # or saas_subscription, travel_booking, etc.
base_url: "https://example.com"
notes: "Description of the site"
flows:
- flow_id: "F1_landing"
flow_type: "landing_cookie_consent"
enabled: true
start_url: "https://example.com"
- flow_id: "F2_checkout"
flow_type: "product_cart_checkout"
enabled: true
start_url: "https://example.com/products/item"
selectors:
# Optional: site-specific selectors
add_to_cart: "button.custom-add-cart"- Run the crawler:
npm run dev -- --config=configs/your_site.yaml- Create a new flow class extending
Flow:
// src/flows/YourNewFlow.ts
import { Flow } from "../core/Flow.js";
export class YourNewFlow extends Flow {
protected async runSteps(): Promise<{
stepsCompleted: number;
screenshots: string[];
}> {
const screenshots: string[] = [];
let stepsCompleted = 0;
// Step 1: Initial state
screenshots.push(...(await this.captureStep("initial")));
stepsCompleted++;
// Step 2: Perform action
const element = await findSomeElement(this.page);
if (element) {
await this.safeClick(element);
screenshots.push(...(await this.captureStep("after_action", "click:element")));
stepsCompleted++;
}
return { stepsCompleted, screenshots };
}
}- Register in
FlowFactory.ts:
case "your_flow_type":
return new YourNewFlow(context);- Add flow type to
types/index.ts:
export type FlowType =
| "landing_cookie_consent"
| ...
| "your_flow_type"; // Add here- Use in site config:
flows:
- flow_id: "F1_your_flow"
flow_type: "your_flow_type"
enabled: trueThe crawler generates the following directory structure:
dataset/
├── ecommerce/
│ ├── site_001/
│ │ ├── flow_E1_checkout/
│ │ │ ├── step_01_full.png
│ │ │ ├── step_01_viewport.png
│ │ │ ├── step_01.json
│ │ │ ├── step_02_full.png
│ │ │ └── ...
│ │ └── flow_E2_landing/
│ │ └── ...
├── saas_subscription/
│ └── ...
Each step generates a JSON file with:
{
"site_id": "site_001",
"category": "ecommerce",
"flow_id": "E1_checkout",
"step_index": 3,
"url": "https://example.com/checkout",
"timestamp": "2024-01-15T10:30:00.000Z",
"viewport": [1366, 768],
"interaction": "click:add_to_cart",
"flow_step_name": "after_add_to_cart",
"notes": "Optional notes about this step"
}landing_cookie_consent: Captures landing page, attempts to reject cookies, then acceptforced_signup_paywall: Detects modals/overlays and attempts to close them
product_cart_checkout: Product page → Add to cart → Cart → Checkout
trial_funnel: Pricing page → Start trial → Payment infocancellation_roach_motel: Account settings → Subscription → Cancel attempt
addon_sneaking: Detect pre-selected checkboxes/add-onstravel_booking_flow: Search → Booking → Extras → Paymentfintech_login_kyc: Feature → Login/KYC gate
site_id: string # Unique identifier
category: SiteCategory # One of: ecommerce, saas_subscription, etc.
base_url: string # Base URL of the site
flows: FlowConfig[] # Array of flow configurations
selectors?: object # Optional site-specific selectors
notes?: string # Optional notesflow_id: string # Unique flow identifier
flow_type: FlowType # Type of flow to execute
enabled: boolean # Whether to run this flow
start_url?: string # Optional: override base_url for this flow
options?: object # Flow-specific optionsRun with --no-headless to see the browser:
npm run dev -- --no-headlessLogs are written to:
logs/combined.log- All logslogs/error.log- Errors only
Console output is enabled in development mode.
-
Element not found: The heuristic selectors may not match. Check logs and consider adding site-specific selectors in config.
-
Timeout errors: Increase timeout in
BrowserManageror flow-specific waits. -
Screenshots not captured: Check that the output directory is writable and has sufficient space.
- This crawler respects
robots.txtwhere reasonable - No private user data is collected
- Screenshots are for academic research purposes
- Use responsibly and in accordance with website terms of service
Each flow is implemented as a state machine:
- Initialize page and navigate
- Execute steps sequentially
- Capture screenshots at each step
- Handle errors gracefully
- Return result with success status
The crawler uses heuristic selectors to avoid hardcoding:
- Text matching (case-insensitive, partial)
- ARIA labels and roles
- Common class patterns
- Fallback strategies
Site-specific selectors can be provided in config for edge cases.
- Flows never crash the crawler
- Failed steps are logged but don't stop execution
- Retry logic at the flow level
- Graceful degradation when elements aren't found
To add new flows or improve existing ones:
- Follow the existing flow pattern
- Use heuristic selectors from
utils/selectors.ts - Capture screenshots at meaningful steps
- Include rich metadata
- Handle errors gracefully
- Update this README
MIT
Built for academic research on dark pattern detection using GUI-graph and GNN approaches.