Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
75 changes: 75 additions & 0 deletions .claude/skills/telegram-scraper/SKILL.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,75 @@
---
name: telegram-scraper
description: Read, search, export and analyze PUBLIC Telegram channels (t.me/<channel>) without an API key or login, using the tgscraper CLI / Python library. Use when the user mentions a Telegram channel, a t.me link, wants the latest posts, news or announcements from Telegram, wants to monitor a channel, export posts to CSV/JSON/Excel/SQLite, download channel photos/videos, or get channel statistics (views, top posts, posting times, hashtags, sentiment).
---

# Telegram Scraper

`tgscraper` reads the public web preview of Telegram channels (`https://t.me/s/<channel>`).
No API key, phone number or login. Only **public channels with web preview** work — not private
groups, chats or bots.

## 0. Setup (once)

```bash
tgscraper --version || pip install "tgscraper[all] @ git+https://github.com/specialteam/TelegramScraper"
```

If the `telegram-scraper` MCP tools (`get_messages`, `search_messages`, `get_channel_info`,
`analyze_channel`, ...) are available, prefer them over the shell — same features.

Channel arguments accept `durov`, `@durov`, `t.me/durov` or `https://t.me/s/durov`.
For a post link `https://t.me/durov/123` the channel is `durov` and the message id is `123`.

## 1. Pick the right command

| User wants | Command |
|---|---|
| Latest posts | `tgscraper durov -n 20 --json` |
| Posts in a date range | `tgscraper durov -n 0 --since 2026-01-01 --until 2026-01-31 --json` |
| Posts about a topic (whole history) | `tgscraper search durov "privacy" -n 30 --json` |
| Filter recent posts by words | `tgscraper durov -n 200 -k bitcoin -k btc --json` |
| Only posts with photos/videos | `tgscraper durov --media-only --media-type photo --json` |
| Popular posts | `tgscraper durov -n 300 --min-views 100000 --json` |
| Channel info (subscribers, description) | `tgscraper info durov --json` |
| Statistics / analysis | `tgscraper stats durov -n 300 --json` |
| Save to a file | `tgscraper durov -n 500 -o durov.csv` (`.json .jsonl .xlsx .db .md`) |
| Several channels | `tgscraper chan1 chan2 chan3 -n 50 -o all.db` |
| Only new posts since last run | `tgscraper durov --incremental -o archive.db` |
| Download photos/videos | `tgscraper media durov -n 30 -d ./media` |
| Monitor for new posts | `tgscraper watch durov -i 120 --webhook URL` (long-running; run in background) |

`-n 0` means "no limit" (whole history — can be slow for big channels; combine with `--since`).

## 2. Output (`--json`)

A JSON array, newest first. Each message:

```json
{"id": 123, "channel": "durov", "url": "https://t.me/durov/123", "date": "2026-01-10T09:30:00+00:00",
"text": "...", "views": 1250000, "author": null, "edited": false, "forwarded_from": null, "reply_to": null,
"media": [{"type": "photo", "url": "https://cdn.../x.jpg"}], "reactions": {"👍": 15000},
"hashtags": ["news"], "mentions": ["telegram"], "links": ["https://..."], "html": "..."}
```

Pipe large results through `jq` instead of reading everything, e.g.
`tgscraper durov -n 200 --json | jq '[.[] | {url, views, text: .text[:120]}]'`.

## 3. Python (for custom processing)

```python
import tgscraper as tg
posts = tg.scrape("durov", limit=100, since="2026-01-01", keywords=["ton"])
tg.export(posts, "out.xlsx")
stats = tg.summarize(posts) # dict: top_posts, posts_by_hour, top_hashtags, sentiment...
results = tg.scrape_many(["a", "b"]) # concurrent, {channel: [Message] | Exception}
```

## 4. Answering well

- Always cite posts with their `url` and date.
- Report views as numbers (`1.2M`), and say how many posts you analyzed.
- `sentiment` is a rough lexicon score (-1..1), mention that it is approximate.
- Errors: `ChannelNotFound` → the channel is private, misspelled, or has web preview disabled.
Network / HTTP 429 errors → wait and retry, or pass a proxy with `-p socks5://host:port`.
- Respect privacy and Telegram's terms: only public data, reasonable request volume.
67 changes: 67 additions & 0 deletions .github/workflows/live-test.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,67 @@
name: Live test

# Scrapes the real t.me to catch Telegram HTML changes.
on:
pull_request:
workflow_dispatch:
inputs:
channel:
description: "Channel to test"
default: "durov"
schedule:
- cron: "17 6 * * 1" # weekly

jobs:
live:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: "3.12"
- run: pip install -e ".[dev]"
- name: Live pytest
env:
TGSCRAPER_LIVE: "1"
TGSCRAPER_LIVE_CHANNEL: ${{ inputs.channel || 'durov' }}
run: pytest tests/test_live.py -v -s
- name: Show raw reaction markup
run: |
python - <<'EOF'
import httpx
from bs4 import BeautifulSoup
soup = BeautifulSoup(httpx.get("https://t.me/s/durov").text, "html.parser")
for r in soup.select(".tgme_reaction")[:6]:
print(r)
EOF
- name: CLI against real channels
run: |
tgscraper info durov telegram
tgscraper durov -n 5
tgscraper stats durov -n 50
tgscraper durov -n 30 -o out/durov.xlsx
tgscraper durov -n 30 -o out/durov.db
tgscraper durov -n 3 --json | python -c "import json,sys; d=json.load(sys.stdin); print(json.dumps(d[0], ensure_ascii=False, indent=2))"
- name: MCP server over stdio
run: |
python - <<'EOF'
import asyncio
from mcp import ClientSession, StdioServerParameters
from mcp.client.stdio import stdio_client

async def main():
async with stdio_client(StdioServerParameters(command="tgscraper-mcp")) as (r, w):
async with ClientSession(r, w) as s:
await s.initialize()
res = await s.call_tool("get_messages", {"channel": "durov", "limit": 2})
text = res.content[0].text
print(text[:1500])
assert '"error"' not in text, text

asyncio.run(main())
EOF
- uses: actions/upload-artifact@v4
if: always()
with:
name: live-output
path: out/
24 changes: 9 additions & 15 deletions .github/workflows/python-package.yml
Original file line number Diff line number Diff line change
@@ -1,6 +1,3 @@
# This workflow will install Python dependencies, run tests and lint with a variety of Python versions
# For more information see: https://docs.github.com/en/actions/automating-builds-and-tests/building-and-testing-python

name: Python package

on:
Expand All @@ -11,30 +8,27 @@ on:

jobs:
build:

runs-on: ubuntu-latest
strategy:
fail-fast: false
matrix:
python-version: ["3.9", "3.10", "3.11"]
python-version: ["3.9", "3.10", "3.11", "3.12", "3.13"]

steps:
- uses: actions/checkout@v4
- name: Set up Python ${{ matrix.python-version }}
uses: actions/setup-python@v3
uses: actions/setup-python@v5
with:
python-version: ${{ matrix.python-version }}
- name: Install dependencies
- name: Install
run: |
python -m pip install --upgrade pip
python -m pip install flake8 pytest
if [ -f requirements.txt ]; then pip install -r requirements.txt; fi
pip install -e ".[dev]"
- name: Lint with flake8
run: |
# stop the build if there are Python syntax errors or undefined names
flake8 . --count --select=E9,F63,F7,F82 --show-source --statistics
# exit-zero treats all errors as warnings. The GitHub editor is 127 chars wide
flake8 . --count --exit-zero --max-complexity=10 --max-line-length=127 --statistics
- name: Test with pytest
run: |
pytest
flake8 . --count --exit-zero --max-complexity=12 --max-line-length=127 --statistics
- name: Test
run: pytest -q
- name: CLI smoke test
run: tgscraper --help
8 changes: 8 additions & 0 deletions .mcp.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,8 @@
{
"mcpServers": {
"telegram-scraper": {
"command": "uvx",
"args": ["--from", ".[mcp]", "tgscraper-mcp"]
}
}
}
34 changes: 34 additions & 0 deletions AGENTS.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,34 @@
# AGENTS.md

Guide for AI coding agents working on this repository. Users of the tool should read
[README.md](README.md) or the skill in `.claude/skills/telegram-scraper/SKILL.md`.

## Setup & checks

```bash
pip install -e ".[dev]"
pytest -q # offline; HTTP is mocked with respx
flake8 . --max-line-length=127
```

The public site `t.me` is never called in tests. Parser tests use `tests/fixtures/page1.html`;
paging tests build pages with `tests/conftest.py::make_page`.

## Architecture

- `tgscraper/parser.py` — the only place that knows Telegram's HTML classes (`tgme_widget_message*`).
If Telegram changes its markup, fix it here and update the fixture.
- `tgscraper/client.py` — `Scraper` (sync) and `AsyncScraper` share paging logic in `_Walk.feed`.
Pages come oldest→newest; the public API yields newest→oldest.
- `tgscraper/__init__.py` — the simple functional API (`scrape`, `search`, ...). Keep it simple.
- `tgscraper/cli.py` — `tgscraper <channel>` is shorthand for `tgscraper scrape <channel>`.
Every command should support `--json`.
- `tgscraper/mcp_server.py` — MCP tools; return JSON-friendly dicts, never raise to the client
(return `{"error": ...}`), strip `html`, cap limits. Works with `mcp` 1.x (FastMCP) and 2.x (MCPServer).

## Conventions

- Python 3.9+ (`from __future__ import annotations`, `typing.Optional`), core deps only `httpx` + `beautifulsoup4`;
everything else is an optional extra (`excel`, `mcp`, `dashboard`).
- New features need a test and a line in README (and the skill / MCP tool if agents should use them).
- `telegram_scraper.py` is a backward-compatibility shim; don't remove it.
14 changes: 14 additions & 0 deletions Dockerfile
Original file line number Diff line number Diff line change
@@ -0,0 +1,14 @@
FROM python:3.12-slim

WORKDIR /app
COPY pyproject.toml README.md LICENSE telegram_scraper.py ./
COPY tgscraper ./tgscraper
RUN pip install --no-cache-dir ".[all]"

# Data (exports, state file) goes here: docker run -v "$PWD/data:/data" ...
WORKDIR /data
ENV TGSCRAPER_STATE=/data/.tgscraper-state.json
EXPOSE 8501

ENTRYPOINT ["tgscraper"]
CMD ["--help"]
Loading
Loading