Skip to content

Repository files navigation

Image Email Extractor

A general-purpose OCR pipeline that extracts email addresses from images — screenshots, scanned documents, business cards, "contact us" graphics, or any other image where an email address appears as pixels rather than text.

The core engine (extractor.py) has no knowledge of any particular website or organization. An example application (examples/) shows it being used to scrape a university's faculty directory, where emails are published as images — a common lightweight anti-scraping technique.

Why this exists

Sites sometimes render contact emails as images specifically to block simple text-scraping. Reading them back out requires an actual OCR pipeline: locating the image, cleaning it up, running it through a text recognition model, and then repairing the inevitable character-recognition noise into a valid email address.

How it works

  1. Multi-strategy preprocessing — the same image is prepared three ways (untouched, contrast-enhanced via CLAHE, and aggressively binarized/inverted), since no single preprocessing approach works best across every font weight and background.
  2. Dual OCR engines — each variant is run through EasyOCR first, then Tesseract (both with default settings and with a tighter single-line character-whitelist config), and every result is pooled together.
  3. Normalization & repair — raw OCR text is cleaned up: symbol substitutions people/sites use ((at) → @, full-width digits, stray punctuation) are fixed, and malformed fragments are discarded.
  4. Optional fuzzy domain correction — if you already know the expected domain (e.g. scraping one organization's directory), you can pass target_domain="example.com" and a small hand-written Levenshtein-distance function will auto-correct near-miss OCR reads (exampl3.con → example.com). This is entirely optional — the extractor works fine for general-purpose use without it.

Quick start

pip install -r requirements.txt
python demo.py

demo.py runs the extractor against the synthetic sample images in sample_images/ — no network access or target website required.

from extractor import extract_emails_from_image

emails = extract_emails_from_image("business_card.png")
print(emails)  # ['jane.doe@example.com']

Testing

pip install -r requirements-dev.txt
pytest

Tests cover email normalization, fuzzy/compound domain repair, OCR result deduplication, and end-to-end extraction against the bundled sample images. See tests/test_extractor.py.

Project structure

extractor.py                              # generic OCR engine — the core of the project
demo.py                                    # runs the extractor on sample_images/
sample_images/                             # synthetic demo images (no real personal data)
tests/
  test_extractor.py                        # unit + integration tests (pytest)
examples/
  scrape_university_directory.py           # example: Playwright scraper + extractor.py
requirements.txt
requirements-dev.txt                       # testing dependencies (pytest)

Example application: directory scraping

examples/scrape_university_directory.py demonstrates a real use case: crawling a public faculty directory, finding each profile's email image, and handing it to extractor.py. Usage:

python examples/scrape_university_directory.py --start-page 1 --end-page 5

Requires:

playwright install chromium

Results are written to the CSV incrementally, one row at a time — if the run is interrupted (Ctrl+C, lost connection, crash), everything extracted up to that point is already saved to disk rather than lost.

Tesseract OCR must also be installed separately and available on your system:

  • Windows: install from the Tesseract releases page, then either add it to your PATH or set:
    set TESSERACT_CMD=C:\Program Files\Tesseract-OCR\tesseract.exe
    
  • macOS: brew install tesseract
  • Linux: sudo apt install tesseract-ocr

A note on the example script's intended use

The example scraper bypasses an image-based email obfuscation scheme that a site likely uses on purpose to deter automated harvesting. It's included to demonstrate the automation + OCR technique, not as an invitation to scrape any particular site. If you adapt it:

  • Respect the target site's robots.txt and terms of service.
  • Keep request rates polite (see DELAY_SECONDS in the script).
  • Don't use extracted contact data for spam or unsolicited outreach.

Tech highlights

  • Generic, reusable OCR extraction API — not tied to any one site or domain
  • Multi-strategy preprocessing + dual-engine OCR pooling for robustness across varied fonts, resolutions, and backgrounds
  • Custom Levenshtein-distance fuzzy domain repair (no external fuzzy-match dependency), optional and off by default
  • Playwright-driven example scraper with resilient multi-selector element discovery

About

A generic OCR pipeline for extracting email addresses from images — dual-engine (EasyOCR + Tesseract), fuzzy domain repair, and confidence-based deduplication. Includes a real-world scraping example.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages