3.6 KiB
3.6 KiB
Telegram Scraper
Telegram scraper on top of Telethon with:
- CLI workflow for scraping, exports, media recovery, and forwarding
- lightweight Web UI served by Python
- SQLite storage per tracked chat/channel
- Docker/Compose setup for "run it and leave it on the server"
What It Does
- Scrapes messages with metadata: views, forwards, reactions, post author
- Downloads media into per-channel folders
- Stores messages in SQLite
- Exports to CSV and JSON
- Supports continuous scraping
- Supports forwarding rules
- Lets you browse tracked channels and local exports in a web panel
Project Layout
.
├── main.py # starts the web server
├── webui_server.py # HTTP server + API
├── webui/ # plain HTML/CSS/JS frontend
├── telegram_scraper_with_forwarding.py
├── data/
│ ├── state.json # shared settings and credentials
│ └── <channel_id>/<channel_id>.db # SQLite databases
└── session/ # Telethon session files
Requirements
- Python 3.11+
- Telegram account
api_idandapi_hashfrom https://my.telegram.org
Telegram API Credentials
- Go to https://my.telegram.org/auth
- Sign in with your Telegram account
- Open
API development tools - Create an application
- Save:
api_idapi_hash
Install
Local
pip install -r requirements.txt
Or with uv:
uv sync
Running
Web UI
Starts on 0.0.0.0:8080 by default:
python main.py
Or:
uv run python main.py
Open:
http://<server-ip>:8080
The current web panel includes:
- tracked channels overview
- background job queue
- scrape/export/media actions
- shared
scrape_mediatoggle - local message viewer powered by SQLite + media files
CLI
If you still want the old interactive mode:
python telegram_scraper_with_forwarding.py
Docker
Build and run:
docker compose up -d --build
Then open:
http://<server-ip>:8080
Volumes
./data:/app/datafor databases, exports, andstate.jsonsession:/app/sessionfor Telethon sessions
If you already logged in before with the same compose volume and did not remove it, the session should be reused.
Shared State
The Web UI and CLI share the same files:
data/state.jsonsession/session.session
That means:
- credentials can stay in
state.json - both entry points use the same tracked channels
- both entry points can reuse the same Telegram login session
Data Storage
SQLite
Each tracked channel gets its own database:
data/<channel_id>/<channel_id>.db
Main fields include:
message_iddatesender_idfirst_namelast_nameusernamemessagemedia_typemedia_pathreply_topost_authorviewsforwardsreactions
Media
Media files are stored in:
data/<channel_id>/media/
Exports
Generated into the same channel folder:
data/<channel_id>/<channel_id>_<username>.csvdata/<channel_id>/<channel_id>_<username>.json
Notes
- The web server is intentionally simple: plain HTML/CSS/JS, no frontend framework
- The viewer currently reads local SQLite data, not Telegram export HTML directly
- If dependencies are installed but the Telegram session is missing, the Web UI falls back to read-only status until login is available
Disclaimer
Use responsibly and make sure your usage complies with Telegram rules, local law, and any privacy obligations.