Skip to content

Repository files navigation

BTP Crawler: Dark Pattern Detection Dataset Collector

A robust, modular web crawler designed for collecting datasets for GUI-graph + GNN-based dark pattern detection research. This crawler automatically captures full-page screenshots, viewport screenshots, and rich metadata for multi-step user flows across various website categories.

🎯 Features

  • Modular Flow System: State machine-based flows for different user interaction patterns
  • Heuristic Selectors: No hardcoded selectors - uses intelligent text/role/aria-label matching
  • Rich Metadata: JSON metadata for every screenshot to support later GUI graph construction
  • Config-Driven: YAML-based site definitions for easy extensibility
  • Retry Logic: Automatic retries with graceful failure handling
  • Full-Page + Viewport Screenshots: Captures both full-page and viewport (1366x768) screenshots
  • Extensible Architecture: Easy to add new flows and site categories

📋 Requirements

  • Node.js 18+
  • npm or yarn
  • Playwright browsers (installed automatically)

🚀 Installation

# Install dependencies
npm install

# Install Playwright browsers
npx playwright install chromium

📁 Project Structure

btp-crawler/
├── src/
│   ├── core/
│   │   ├── BrowserManager.ts    # Browser lifecycle management
│   │   ├── Flow.ts               # Abstract base flow class
│   │   ├── FlowRunner.ts         # Central orchestrator
│   │   └── ConfigLoader.ts       # YAML config loader
│   ├── flows/
│   │   ├── FlowFactory.ts        # Flow factory
│   │   ├── LandingCookieConsentFlow.ts
│   │   ├── ForcedSignupPaywallFlow.ts
│   │   ├── ProductCartCheckoutFlow.ts
│   │   ├── TrialFunnelFlow.ts
│   │   └── CancellationRoachMotelFlow.ts
│   ├── utils/
│   │   ├── logger.ts             # Winston logger
│   │   ├── selectors.ts          # Heuristic selector utilities
│   │   └── screenshot.ts         # Screenshot capture utilities
│   ├── types/
│   │   └── index.ts              # TypeScript type definitions
│   └── index.ts                  # Main entry point
├── configs/
│   ├── example_ecommerce.yaml
│   └── example_saas.yaml
├── dataset/                      # Output directory (created automatically)
└── README.md

🎮 Usage

Basic Usage

# Run all sites in configs/ directory (headless)
npm run crawl

# Run in non-headless mode (for debugging)
npm run dev -- --no-headless

# Run specific site config
npm run dev -- --config=configs/example_ecommerce.yaml

# Run all configs in a directory
npm run dev -- --config-dir=configs/

Development Mode

# Run with TypeScript directly (faster iteration)
npm run dev

📝 Adding New Sites

  1. Create a YAML file in configs/ directory:
site_id: "site_003"
category: "ecommerce"  # or saas_subscription, travel_booking, etc.
base_url: "https://example.com"
notes: "Description of the site"

flows:
  - flow_id: "F1_landing"
    flow_type: "landing_cookie_consent"
    enabled: true
    start_url: "https://example.com"

  - flow_id: "F2_checkout"
    flow_type: "product_cart_checkout"
    enabled: true
    start_url: "https://example.com/products/item"

selectors:
  # Optional: site-specific selectors
  add_to_cart: "button.custom-add-cart"
  1. Run the crawler:
npm run dev -- --config=configs/your_site.yaml

🔄 Adding New Flows

  1. Create a new flow class extending Flow:
// src/flows/YourNewFlow.ts
import { Flow } from "../core/Flow.js";

export class YourNewFlow extends Flow {
  protected async runSteps(): Promise<{
    stepsCompleted: number;
    screenshots: string[];
  }> {
    const screenshots: string[] = [];
    let stepsCompleted = 0;

    // Step 1: Initial state
    screenshots.push(...(await this.captureStep("initial")));
    stepsCompleted++;

    // Step 2: Perform action
    const element = await findSomeElement(this.page);
    if (element) {
      await this.safeClick(element);
      screenshots.push(...(await this.captureStep("after_action", "click:element")));
      stepsCompleted++;
    }

    return { stepsCompleted, screenshots };
  }
}
  1. Register in FlowFactory.ts:
case "your_flow_type":
  return new YourNewFlow(context);
  1. Add flow type to types/index.ts:
export type FlowType =
  | "landing_cookie_consent"
  | ...
  | "your_flow_type";  // Add here
  1. Use in site config:
flows:
  - flow_id: "F1_your_flow"
    flow_type: "your_flow_type"
    enabled: true

📊 Dataset Structure

The crawler generates the following directory structure:

dataset/
├── ecommerce/
│   ├── site_001/
│   │   ├── flow_E1_checkout/
│   │   │   ├── step_01_full.png
│   │   │   ├── step_01_viewport.png
│   │   │   ├── step_01.json
│   │   │   ├── step_02_full.png
│   │   │   └── ...
│   │   └── flow_E2_landing/
│   │       └── ...
├── saas_subscription/
│   └── ...

Metadata Format

Each step generates a JSON file with:

{
  "site_id": "site_001",
  "category": "ecommerce",
  "flow_id": "E1_checkout",
  "step_index": 3,
  "url": "https://example.com/checkout",
  "timestamp": "2024-01-15T10:30:00.000Z",
  "viewport": [1366, 768],
  "interaction": "click:add_to_cart",
  "flow_step_name": "after_add_to_cart",
  "notes": "Optional notes about this step"
}

🔧 Available Flows

Universal Flows

  • landing_cookie_consent: Captures landing page, attempts to reject cookies, then accept
  • forced_signup_paywall: Detects modals/overlays and attempts to close them

Ecommerce Flows

  • product_cart_checkout: Product page → Add to cart → Cart → Checkout

SaaS Flows

  • trial_funnel: Pricing page → Start trial → Payment info
  • cancellation_roach_motel: Account settings → Subscription → Cancel attempt

Planned Flows

  • addon_sneaking: Detect pre-selected checkboxes/add-ons
  • travel_booking_flow: Search → Booking → Extras → Payment
  • fintech_login_kyc: Feature → Login/KYC gate

🎛️ Configuration Options

Site Config Schema

site_id: string              # Unique identifier
category: SiteCategory       # One of: ecommerce, saas_subscription, etc.
base_url: string            # Base URL of the site
flows: FlowConfig[]         # Array of flow configurations
selectors?: object          # Optional site-specific selectors
notes?: string              # Optional notes

Flow Config Schema

flow_id: string             # Unique flow identifier
flow_type: FlowType         # Type of flow to execute
enabled: boolean            # Whether to run this flow
start_url?: string          # Optional: override base_url for this flow
options?: object            # Flow-specific options

🐛 Debugging

Non-Headless Mode

Run with --no-headless to see the browser:

npm run dev -- --no-headless

Logs

Logs are written to:

  • logs/combined.log - All logs
  • logs/error.log - Errors only

Console output is enabled in development mode.

Common Issues

  1. Element not found: The heuristic selectors may not match. Check logs and consider adding site-specific selectors in config.

  2. Timeout errors: Increase timeout in BrowserManager or flow-specific waits.

  3. Screenshots not captured: Check that the output directory is writable and has sufficient space.

🔒 Ethics & Legal

  • This crawler respects robots.txt where reasonable
  • No private user data is collected
  • Screenshots are for academic research purposes
  • Use responsibly and in accordance with website terms of service

📚 Architecture Notes

Flow System

Each flow is implemented as a state machine:

  1. Initialize page and navigate
  2. Execute steps sequentially
  3. Capture screenshots at each step
  4. Handle errors gracefully
  5. Return result with success status

Selector Strategy

The crawler uses heuristic selectors to avoid hardcoding:

  • Text matching (case-insensitive, partial)
  • ARIA labels and roles
  • Common class patterns
  • Fallback strategies

Site-specific selectors can be provided in config for edge cases.

Error Handling

  • Flows never crash the crawler
  • Failed steps are logged but don't stop execution
  • Retry logic at the flow level
  • Graceful degradation when elements aren't found

🤝 Contributing

To add new flows or improve existing ones:

  1. Follow the existing flow pattern
  2. Use heuristic selectors from utils/selectors.ts
  3. Capture screenshots at meaningful steps
  4. Include rich metadata
  5. Handle errors gracefully
  6. Update this README

📄 License

MIT

🙏 Acknowledgments

Built for academic research on dark pattern detection using GUI-graph and GNN approaches.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages