Skip to content

Rebuild as tgscraper: library, CLI, MCP server, Claude skill, dashboard - #1

Merged
specialteam merged 3 commits into
mainfrom
claude/quirky-lovelace-bty4zt
Sep 26, 2026
Merged

specialteam merged 3 commits into
mainfrom
claude/quirky-lovelace-bty4zt

Conversation

@specialteam

@specialteam specialteam commented Sep 26, 2026 •

Copy link
Copy Markdown
Owner

Summary

Turns the single-class scraper into the tgscraper package, with a one-line API (tg.scrape("durov")) and a CLI (tgscraper durov).

  • Core: new data models for messages (views, reactions, media, replies, forwards, hashtags, links) and channel info. Sync and async clients with timeouts, retries/backoff, Retry-After handling, rate limiting and proxy rotation.
  • Fixes to the old code: requests now time out; messages are deduplicated by id instead of text; line breaks and links are kept; media-only posts are no longer skipped.
  • Features: filters (dates, keywords, regex, hashtag, media type, min views), Telegram server-side search, incremental scraping, media download, monitoring with webhook / Telegram bot notifiers, analytics with English/Persian sentiment.
  • Export: JSON, JSONL, CSV, XLSX, SQLite (upsert) and Markdown.
  • AI integration:
    • MCP server tgscraper-mcp with 8 tools, 2 prompts and 1 resource; works with mcp 1.x and 2.x.
    • .mcp.json registers the server automatically in this repo.
    • Claude skill in .claude/skills/telegram-scraper.
    • llms.txt and AGENTS.md describe the project for AI tools.
  • Tooling: Streamlit dashboard, Dockerfile and docker-compose, pyproject.toml with optional extras, CI on Python 3.9–3.13.
  • Docs: new README with a Persian section.
  • Compatibility: from telegram_scraper import TelegramScraper still works as a shim.

Testing

  • 43 offline tests (respx mocks and an HTML fixture) pass on Python 3.9–3.13; flake8 reports no errors.

  • New Live test workflow scrapes the real t.me and runs on PRs, weekly, and on demand. It covers:

    • channel info, paging, search, single message, not-found handling, multi-channel scraping, and reactions;
    • CLI commands and XLSX/SQLite export;
    • an MCP stdio call.

    It passes.

  • The live run caught one real bug: all reactions were merged into a single "custom" key. Fixed in e751afd, which parses image-encoded, custom and paid reactions.

  • The Streamlit dashboard was not run, only compiled.

🤖 Generated with Claude Code

https://claude.ai/code/session_01Koo8a1vU5D6JttUkzjxE6H

- Structured Message/Channel models (views, reactions, media, replies, forwards, hashtags...)
- Sync + async clients with retries, rate limiting, proxy rotation, timeouts
- Filters, server-side search, incremental scraping, media download, monitoring
- Export to JSON/JSONL/CSV/XLSX/SQLite/Markdown; analytics with EN/FA sentiment
- tgscraper CLI, MCP server (8 tools, mcp 1.x and 2.x), Claude skill, .mcp.json
- Streamlit dashboard, Docker, pyproject packaging, offline test suite
- New README, llms.txt and AGENTS.md; legacy TelegramScraper kept as a shim

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Koo8a1vU5D6JttUkzjxE6H
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Koo8a1vU5D6JttUkzjxE6H
Live run on t.me showed every reaction merged into one "custom" key.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Koo8a1vU5D6JttUkzjxE6H
@specialteam
specialteam merged commit 6ec7f5a into main Sep 26, 2026
6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants