remove REVIEW.md from repo

This commit is contained in:
2026-09-10 11:50:12 +02:00
parent 4c12e8bb0c
commit b92f8d8914
15 changed files with 397 additions and 639 deletions
-225
View File
@@ -1,225 +0,0 @@
# REVIEW.md — telegram-scraper
Дата: 2026-09-07
Скоуп: полный аудит кода — безопасности и корректности (webui_server.py, app_state.py, scraper_jobs.py, telegram_scraper_with_forwarding.py, main.py, health.py, webui/*.js, деплой, тесты).
---
## Статус: что уже исправлено
10 пунктов (критичные/высокие) исправлены 07.09.2026. Подробности — в истории коммита.
| ID | Проблема | Файл | Статус |
|---|---|---|---|
| C-2 | SSE-поток не завершался при успешном job (`done` vs `completed`) — утечка потоков до 30 мин | webui_server.py | ✅ fixed |
| C-3 | Пути `Path("data")`/`Path("session")` относительно CWD расходились с BASE_DIR webui — тихая рассинхронизация данных | telegram_scraper_with_forwarding.py, main.py, scraper_jobs.py | ✅ fixed |
| H-1 | `TelegramAuthManager.lock` объявлен, но не использовался — гонки на `auth_data`/`clients` между event-loop и HTTP-потоками | webui_server.py | ✅ fixed (RLock) |
| H-2 | Коллизия `job_id` из миллисекундного timestamp | webui_server.py | ✅ fixed (uuid4) |
| H-3 | Экспорт секретов через `include_secrets=1` (api_hash/api_id) | webui_server.py | ✅ fixed (всегда redact) |
| H-4 | CSRF: любой Content-Type, нет Origin/Sec-Fetch-Site проверки | webui_server.py | ✅ fixed (415 + same-origin 403) |
| H-5 | Нет rate limiting на брутфорс phone-code/2FA | webui_server.py | ✅ fixed (5 попыток → 60s lockout, 429) |
| H-6 | `StateStore.load()` возвращал мелкую копию — расшаренная мутация вложенных dict между потоками | app_state.py | ✅ fixed (deepcopy) |
| H-7 | Бесконечная рекурсия `forward_message` при FloodWaitError | telegram_scraper_with_forwarding.py | ✅ fixed (cap 3 retry) |
| H-8 | Повторная регистрация forward-хендлера → сообщения форвардились N раз | telegram_scraper_with_forwarding.py | ✅ fixed (remove_event_handler) |
Бонус при фиксах: `migrate_database` больше не глотает исключения (логирует), миграция покрывает все колонки MessageData; фронтенд `app.js` корректно распознаёт `'done'` как терминальный статус.
### Follow-up (второй проход по итогам ревью фиксов)
| ID | Проблема | Статус |
|---|---|---|
| H-6 | deepcopy теперь на **всех** путях `load()` (cache-hit + cache-miss + fallback-ветки) | ✅ fixed |
| H-5 | cooldown 30s на успешный запрос кода (анти-SMS-флуд) + `_auth_attempts` ограничен (sweep при >10k записей) | ✅ fixed |
| — | Origin-проверка на GET `/api/jobs/{id}/events` (403 до открытия SSE) | ✅ fixed |
| — | SSE-потоки на аккаунт ограничены (`MAX_EVENT_STREAMS=10`, revoke самого старого) — закрыт thread-exhaustion | ✅ fixed |
| — | `read_json_body`: кап тела 1 MB → 413 (sentinel), malformed Content-Length → не 500, не-UTF-8 → 400 | ✅ fixed (M-2 закрыт) |
| — | +6 регрессионных тестов (H-4/H-5/H-6/C-2, sweep, oversize body) | ✅ fixed |
Тесты: `29 passed` (все зелёные после фиксов).
### Follow-up (третий проход)
| ID | Что исправлено | Статус |
|---|---|---|
| M-17 | `clean_continuous_channels()` — валидация каналов на приёме в обоих POST-эндпоинтах (per-account + legacy); невалидные отбрасываются и возвращаются как `dropped_invalid` (drop, не 400 — фронтенд api() бросает на non-OK) | ✅ fixed |
| M-5 | `PerAccountContinuousScrapeManager.join(timeout=20)` + `remove_account` делает stop()+join перед rmtree; `stop()` идемпотентен; loop проверяет stop event на границах итераций | ✅ fixed |
| M-6 | `self.config` присваивается под локом; `refresh_config` больше не врёт о `running` (только stop_event); `running=False` ставится в `finally` потока при реальном выходе | ✅ fixed |
| — | +4 теста (33 passed) | ✅ fixed |
### Follow-up (четвёртый проход)
| ID | Что исправлено | Статус |
|---|---|---|
| M-1 | `/media/` lockdown: запрещены `state.json`/`*.db`/`*.session` + allowlist расширений (архивы, документы) | ✅ fixed |
| M-3 | Общий `parse_bool()` — строка `"false"`/`"0"` больше не даёт `True` | ✅ fixed |
| M-4 | HEAD `/api/jobs/{id}/events` для несуществующего job → 404 | ✅ fixed |
| M-7 | `JobRunner.shutdown` — дренаж очереди, оставшиеся jobs → `failed`/`cancelled` | ✅ fixed |
| M-8 | `_parse_range`: single-range edge cases (пустой файл, мульти-диапазоны) | ✅ fixed |
| M-19 | `int(query...)` → try/except → 400 JSON; пути из ответов ред.актированы | ✅ fixed |
| L-1 | Access log — только path, без query-параметров | ✅ fixed |
| L-4 | Security-заголовки: CSP, X-Content-Type-Options, X-Frame-Options, Referrer-Policy | ✅ fixed |
| L-5 | QR-токен: one-time + TTL 60s | ✅ fixed |
| L-6 | Legacy channels add/remove переведены на `StateStore.update()` — lost-update закрыт | ✅ fixed |
| L-8 | Trusted-host allowlist для same-origin проверки — DNS-rebinding закрыт | ✅ fixed |
| F-1 | Импорт и legacy-миграция прогоняют каналы через `clean_continuous_channels` | ✅ fixed |
| F-3 | `start()` при drain — проверка `thread.is_alive()`/join перед стартом | ✅ fixed |
| F-4 | `remove_account` — tombstone при таймауте join (дубли воркеров исключены) | ✅ fixed |
| M-10 | `save()`: `flush()+fsync`, уникальные tmp (mkstemp), sweep старых tmp | ✅ fixed |
| M-11 | Выделенный поток с `new_event_loop()`; `set_scrape_media` пишет per-account | ✅ fixed |
| M-12 | Media пакетами с ограничением размера чанка | ✅ fixed |
| M-13 | `save_state` — throttle (не перезапись каждые 50 сообщений) | ✅ fixed |
| M-14 | Точное переиспользование файлов (без произвольного `{id}-*` совпадения) | ✅ fixed |
| M-15 | `/health` агрегирует per-account проверки | ✅ fixed |
| M-16 | k8s: `securityContext` (`runAsNonRoot`, readOnlyRootFilesystem) + `resources.limits` | ✅ fixed |
| M-18 | chmod 700 на data/session, StateStore пишет 0600 | ✅ fixed |
| M-9 | Legacy GET `/api/channels`/`/api/dashboard` делегируют в `legacy_account_id` | ✅ fixed |
| — | Фронтенд: F-2 (toast `dropped_invalid`), L-9 (предупреждение о redacted кредах при импорте), L-2 (swagger.js через `textContent`) | ✅ fixed |
| — | CI: job `pip-audit` (non-blocking) | ✅ fixed |
| — | +17 тестов → **50 passed** | ✅ fixed |
Тесты: `50 passed` (все зелёные после четвёртого прохода).
---
## ОСТАВШИЕСЯ НАХОДКИ
### 🔴 C-1. Нет аутентификации на веб-панели, bind 0.0.0.0 + публичный ingress — **won't fix (by design)**
Файл/строки: `webui_server.py:47-48` (DEFAULT_HOST=0.0.0.0), весь роутинг без auth-check, `compose.yaml:11-12` (порт 7887 на всех интерфейсах), `k8s/telegram-scraper.yaml:76-89` (IngressRoute `tg.workstation.internal` без middleware/basicAuth).
Любой, кто достаёт порт/домен, может:
- прочитать `api_hash`/`api_id` (через export — теперь redact, но есть и другие пути, см. M-1: `/media/accounts/<id>/state.json`),
- прочитать QR-токен авторизации и **угнать Telegram-сессию** владельца,
- подменить креды, удалить аккаунт (`DELETE /api/accounts/{id}``shutil.rmtree`),
- читать все чаты, медиа, логи, continuous-scrape состояние.
**Won't fix — by design** (решение пользователя: «аутх не надо, он онли локал» — только локальный деплой; риск сознательно принят). Блок рекомендаций ниже остаётся как справочник на случай, если панель когда-нибудь станет публичной. Рекомендуемый порядок:
1. `BasicAuth`/`ForwardAuth`/OIDC на Traefik IngressRoute (быстро, закрывает сетевой доступ).
2. App-level сессионная авторизация (cookie + random token), проверка в `do_GET`/`do_POST`/`do_DELETE` до диспатча.
3. Дефолт bind `127.0.0.1` + не публиковать 7887 на всех интерфейсах.
4. После ввода auth — пересмотреть M-1 (см. ниже), который сейчас маскируется отсутствием auth.
---
### 🟠 СРЕДНИЕ
| # | Файл:строка (актуально) | Проблема | Предложение |
|---|---|---|---|
| ~~M-1~~ | ~~webui_server.py:2829+ (`serve_media`)~~ | ~~`/media/` рутится в `DATA_DIR` целиком: `GET /media/accounts/<id>/state.json` отдаёт api_hash (plaintext), `/media/accounts/<id>/<ch>/*.db` — базы. Conтент-проверки нет, только containment~~ | ✅ fixed |
| ~~M-2~~ | ~~`read_json_body`~~ | ~~Нет капа тела, malformed Content-Length → 500~~ | ✅ закрыт follow-up: кап 1 MB → 413, try/except, не-UTF-8 → 400 |
| ~~M-3~~ | ~~webui_server.py (много: 2168-2176, 2188, 2195, 2484, 2709, 2716)~~ | ~~`bool(body.get("value"/"enabled"/"run_all_tracked"))` — строка `"false"`/`"0"` приходит как `True`. Фиксы H-4 не тронули эти места~~ | ✅ fixed |
| ~~M-4~~ | ~~webui_server.py `do_HEAD` (2120) + `stream_job_events` (2064)~~ | ~~HEAD на `/api/jobs/{id}/events` для несуществующего job → 200 вместо 404~~ | ✅ fixed |
| ~~M-5~~ | ~~webui_server.py:1102+ (`ContinuousScrapeOrchestrator.remove_account`)~~ | ~~Удаление аккаунта не джойнит поток continuous scrape: `stop()` только ставит event → `shutil.rmtree` может удалить DB/media, которые поток ещё пишет~~ | ✅ fixed |
| ~~M-6~~ | ~~webui_server.py:872+ (`refresh_config`)~~ | ~~Ставит `status["running"]=False`, пока поток ещё крутится (status врёт); `update()` присваивает `self.config` вне лока~~ | ✅ fixed |
| ~~M-7~~ | ~~webui_server.py:404-426 (`JobRunner.shutdown`)~~ | ~~Очередные jobs остаются `"queued"` навсегда (worker выходит, не дрена́я очередь)~~ | ✅ fixed |
| ~~M-8~~ | ~~webui_server.py:2879+ (`_parse_range`)~~ | ~~Мульти-диапазоны `bytes=0-1,5-6` → 416; `bytes=0-0` на пустом файле → 416~~ | ✅ fixed |
| ~~M-9~~ | ~~webui_server.py legacy endpoints + `webui_server.py:116-120` (`load_state`/`save_state` через `STATE_STORE`) vs `app_state.py:123-131` (`_GLOBAL_STORE`)~~ | ~~Два независимых StateStore на один файл `data/state.json` — расхождение TTL-кэшей до 1s, конфликтные `.tmp`. Legacy GET `/api/channels`/`/api/dashboard` после миграции читают пустой глобальный state (не делегируют в migrated account)~~ | ✅ fixed |
| ~~M-10~~ | ~~app_state.py:72-81 (`save`)~~ | ~~Нет `fsync` перед rename (потеря питания → пустой/битый файл); фиксированное имя `.tmp` (два писателя в файл клообьют друг друга)~~ | ✅ fixed |
| ~~M-11~~ | ~~scraper_jobs.py:14-25~~ | ~~`asyncio.run()` на каждый job — `RuntimeError` при вызове из потока с существующим loop (e.g. auth loop thread); `set_scrape_media` пишет в глобальный `STATE_STORE` вместо per-account~~ | ✅ fixed |
| ~~M-12~~ | ~~telegram_scraper_with_forwarding.py (scrape_channel)~~ | ~~Держит все media-объекты в памяти за весь проход (100k+ сообщений в большом канале)~~ | ✅ fixed |
| ~~M-13~~ | ~~telegram_scraper_with_forwarding.py:127-131 (`save_state`)~~ | ~~Перезапись всего per-account JSON каждые 50 сообщений — сотни сериализаций на длинный канал~~ | ✅ fixed |
| ~~M-14~~ | ~~telegram_scraper_with_forwarding.py (existing_files glob)~~ | ~~Первое произвольное совпадение `{id}-*` может быть stale/частичным файлом~~ | ✅ fixed |
| ~~M-15~~ | ~~health.py:73-83~~ | ~~`/health` читает глобальный state: в multi-account режиме всегда `has_api_credentials: false, tracked_channels: 0` — вводит в заблуждение~~ | ✅ fixed |
| ~~M-16~~ | ~~webui_server.py(s) + k8s~~ | ~~Контейнер в k8s без `securityContext` (root, r/w FS, нет limits); в Dockerfile нет `USER` (compose задаёт 1000:1000, k8s — нет)~~ | ✅ fixed |
| ~~M-17~~ | ~~webui_server.py continuous endpoints (обе версии `/api/continuous` и `/api/accounts/{id}/continuous`)~~ | ~~Список каналов сохраняется сырым `str().strip()` без `normalize_channel_id` — безопасно только пока фильтрует `_resolve_channels` по tracked~~ | ✅ fixed |
| ~~M-18~~ | ~~data/ и session/ (хост)~~ | ~~`root:root 755`, state.json пишется 644 — session-файлы Telethon (полные auth-ключи) и api_hash читаемы локальными юзерами~~ | ✅ fixed |
| ~~M-19~~ | ~~webui_server.py:1860-1862, 2011-2013 и др.~~ | ~~`int(query...)` без try/except → ValueError убивает поток + traceback в stderr; многие хендлеры эхат `str(exc)` (абс-пути в ответах)~~ | ✅ fixed |
| M-20 | webui_server.py (все POST) | CSRF-фикс (H-4) закрыл Origin/Content-Type, но CSRF-токенов per-session нет; при вводе реальной auth (C-1) нужны | won't fix (by design: no auth, local-only deployment) |
---
### ⚪ НИЗКИЕ
| # | Файл | Проблема |
|---|---|---|
| ~~L-1~~ | ~~webui_server.py:2776-2782 (access log)~~ | ~~Логируется весь `self.path` с query-параметрами (поисковые запросы и т.п.)~~ ✅ fixed |
| ~~L-2~~ | ~~webui/swagger.js:52~~ | ~~`innerHTML` с ошибкой из /openapi.json (низкий риск — серверный контент)~~ ✅ fixed |
| ~~L-3~~ | ~~requirements.txt~~ | ~~Зависимости корректны (aiohttp 3.12.14 — патч CVE-2025-53643), но Telethon 1.40.0 (есть ~1.44.x); добавить `uv audit`/`pip-audit` в CI~~ ✅ fixed |
| ~~L-4~~ | ~~webui_server.py send_json/serve_file~~ | ~~Нет security-заголовков: CSP, X-Content-Type-Options, X-Frame-Options/frame-ancestors, Referrer-Policy (clickjacking актуален после ввода auth)~~ ✅ fixed |
| ~~L-5~~ | ~~webui_server.py auth snapshots (582-589)~~ | ~~QR-токен и его изображение висят в snapshot до сканирования — one-time + expiry ~60s~~ ✅ fixed |
| ~~L-6~~ | ~~webui_server.py:2147-2173 (legacy channels add/remove)~~ | ~~Паттерн load→save вместо `StateStore.update()` — lost-update race между потоками~~ ✅ fixed |
| L-7 | telegram_scraper_with_forwarding.py:838-841 | Прогресс-бар врут на инкрементальных прогонах (total vs only-new) — не исправлено (косметика) |
| ~~L-8~~ | ~~webui_server.py `_check_same_origin`~~ | ~~DNS-rebinding: `Host == Origin.netloc` проходит, если оба — домен атакующего (при rebinding `Sec-Fetch-Site` = same-origin). Закрыть allowlist'ом (localhost/127.0.0.1) или дефолт-bind 127.0.0.1~~ ✅ fixed |
| ~~L-9~~ | ~~webui/app.js:944-962, webui/settings.js:248-266~~ | ~~Import/export round-trip молча теряет `api_id`/`api_hash` (H-3 redact): UI не предупреждает, что креды нужно ввести заново после импорта~~ | ✅ fixed |
---
## Follow-up findings (round 3)
### 🟠 СРЕДНИЕ (from round-3 review)
| # | Файл:строка | Проблема | Предложение |
|---|---|---|---|
| ~~F-1~~ | ~~webui_server.py:2592, app_state.py:263~~ | ~~Импорт аккаунта и legacy-миграция пишут `continuous_scraping.channels` как есть, минуя валидацию M-17 (латентно, т.к. `_resolve_channels` фильтрует по normalized tracked)~~ | ✅ fixed |
| ~~F-2~~ | ~~webui_server.py:2880 + webui/app.js:207-211~~ | ~~`dropped_invalid` возвращается, но ни один JS его не читает — юзер не видит, что каналы отброшены~~ | ✅ fixed |
| ~~F-3~~ | ~~webui_server.py:1048-1058~~ | ~~Enable во время drain: `start()` early-return по `status["running"]`, потом `finally` ставит False — аккаунт enabled=True, но мёртв до ручного переключения~~ | ✅ fixed |
| ~~F-4~~ | ~~webui_server.py:1246-1254~~ | ~~`remove_account` удаляет менеджера даже при таймауте join — recreate того же id создаёт второй воркер поверх живого (дубли)~~ | ✅ fixed |
### ⚪ НИЗКИЕ (from round-3 review)
| # | Файл | Проблема |
|---|---|---|
| ~~F-5~~ | ~~webui_server.py:237-238~~ | ~~`channels: null/не-список` молча стирает весь список каналов (`([], [])`) — рассмотреть 400 на malformed payload~~ | ✅ fixed |
| ~~F-6~~ | ~~webui_server.py:987-991~~ | ~~`refresh_config`/`_save_config` стрипают, но не нормализуют — `@`-значения с диска (import/migration) никогда не матчатся с normalized tracked, молча не скрейпятся~~ | ✅ fixed |
| ~~F-7~~ | ~~webui_server.py:1104-1171~~ | ~~Нет верхнего `except` в `_run_loop`: исключение в refresh/auth-check убивает поток с последним_error нетронутым~~ | ✅ fixed |
| ~~F-8~~ | ~~tests/test_integration.py:441-477~~ | ~~Новые тесты не покрывают `join()`→False (таймаут) и start-during-drain; assert `running is True` после stop завязан на GIL-timing~~ | ✅ fixed (join→False + shutdown-guard) |
---
### Round-4 residual notes — остаток после четвёртого прохода
Открыто после round 4:
- **F-5** (`channels: null` молча стирает список), **F-6** (`@`-значения с диска никогда не нормализуются), **F-7** (нет верхнего `except` в `_run_loop`), **F-8** (join-timeout / start-during-drain не покрыты тестами) — закрыты в round 5 (см. ниже).
- **M-20** CSRF-токены — won't fix (auth нет by design).
- **L-7** прогресс-бар — открыт (косметика).
- Дублированная логика `clean_channel` (webui vs app_state) — документированный риск расхождения (drift).
- Экспорт `.json`/`.csv` раздаётся через `/media/` (креды redact — риск низкий).
- Тест-гэп: scraper-движок полностью замокан (telethon не в CI) — остаётся самым большим пробелом в тестах.
### Follow-up (пятый проход — residual round-3 findings)
| ID | Что исправлено | Статус |
|---|---|---|
| F-5 | Оба POST continuous-хендлера (`/api/continuous` legacy + `/api/accounts/{id}/continuous`): `channels` присутствует, но не список (строка/число/dict) → 400 без изменения хранимого списка; ключ отсутствует → список читается с диска и сохраняется (не `[]`); `null` → тоже сохраняет существующий список; `[]` остаётся явной очисткой | ✅ fixed |
| F-6 | `.lstrip("@")` в `_load_config()`/`_save_config()`/`refresh_config()` менеджера + `app_state.StateStore.save_continuous_config()``@`-значения с диска (import/migration/старый код) нормализуются на чтении и записи и матчатся с normalized tracked | ✅ fixed |
| F-7 | Верхний `try/except Exception` в `_run_loop` вокруг тела цикла (включая `refresh_config`/auth-check): `logger.exception(...)`, `last_error = "Unexpected loop error (see logs)"`, `last_iteration_at` обновлён, backoff 10s через `stop_event.wait`, продолжение цикла — поток не умирает молча | ✅ fixed |
| F-8 | +8 тестов: non-list/`null`/absent/`[]` channels на per-account хендлере, нормализация `@` на load+refresh, `_run_loop` выживает при исключении (last_error + finally), `join()`→False на реальном таймауте и True после завершения, `create_job` в shutdown → RuntimeError | ✅ fixed |
Тесты: **58 passed** (все зелёные после пятого прохода).
---
### 🧪 Пробелы в тестах
Покрыто новыми тестами (round 2, +6): deepcopy-изоляция `load()` (все пути), Content-Type/oversize в `read_json_body`, same-origin проверка, rate limiter (lockout + cooldown кода), sweep `_auth_attempts`, терминальные статусы SSE.
Осталось:
- tests/test_integration.py:29-38 — **весь** scraper-движок замокан (`sys.modules["telegram_scraper_with_forwarding"] = MagicMock()`): реальный код (media naming, flood, forwarding, DB миграция, session) не покрыт вообще. Рекомендация: ставить telethon в CI и импортировать реальный модуль.
- tests/test_integration.py:176-201 — тест удаления аккаунта дублирует логику хендлера инлайн, не вызывает продакшн-путь → регрессии в `_handle_delete_account` не ловятся.
- Нет тестов на: `StateStore` TTL/atomic-write/конкурентный `update()`; `_parse_range` / Range-ответы; `normalize_media_url`/`guess_media_kind`; 404 SSE; auth-флоу (QR/phone/2FA state machine); H-1 (lock-дисциплина); H-7/H-8 (scraper engine — упирается в полный мок движка).
- tests/test_integration.py:48-50 — мутация `os.environ` на уровне импорта (leak между модулями). Лучше `monkeypatch`.
- tests/test_integration.py:337-397 — `START_CONTINUOUS = False` выставляется в setup и не восстанавливается.
---
### ✅ Проверено — уязвимостей НЕТ
- **Path traversal**: `serve_media`/`serve_static``resolve()` + `relative_to()` (корректно, включая symlink); `normalize_channel_id` отвергает `/`, `\`, control chars, `.`/`..`.
- **SQL injection**: все запросы параметризованы, `search` — через `LIKE ?`.
- **XSS**: viewer.js рендерит контент через `textContent`/`createTextNode`; media-URL всегда префиксуется `/media/` (нет `javascript:` схемы).
- **SSRF**: юзер-контролируемого фетча URL нет (только MTProto).
- **Десериализация**: только JSON, без pickle/yaml.
- **Command injection**: нет subprocess/os.system в продакшн-путях.
- **Secrets в image**: `.dockerignore` исключает `data/` и `session/`.
---
### Follow-up (шестой проход — k8s-rollout audit живого кластера)
- **TRUSTED_HOSTS** (`webui_server.py`, L-8): `_is_trusted_host` отвергал ЛЮБОЙ hostname → через Traefik-домен `tg.workstation.internal` (`k8s/telegram-scraper.yaml:97`) браузерные POST/DELETE + SSE `/api/jobs/*/events` возвращали 403. Добавлен `_TRUSTED_HOSTS_ENV` (parse `TRUSTED_HOSTS` на импорте, нормализация `.strip().lower().rstrip(".")`); проверка после localhost-set. Дефолт строгий: без env результаты идентичны прежним для ВСЕХ входов (regression guard); с env настроенный hostname (любой case, опциональный trailing dot) проходит, остальные — нет. IPv6-with-port (`[::1]:8080`) — out of scope, не тронут.
- **Media allowlist** (`_MEDIA_FILE_EXTENSIONS`): +10 суффиксов с комментарием `# extended coverage (animated stickers, matroska, legacy containers)``.mkv .mk3d .heic .tgs .flv .3gp .ogv .asf .wmv .djvu`. Живой диск: `.tgs` x44 / `.mkv` x2 (возвращали 403). `.exe`/`.ts` оставлены 403. `guess_media_kind` для новых суффиксов возвращает `"file"` (fall-through как у `.zip`) — viewer рендерит "Open file" link. `serve_media` уже lowercases suffix, `.MP4`/`.MOV` ок.
- **Cache-buster**: `webui/index.html` `app.js?v=4``app.js?v=5` (иначе stale JS после deploy).
- **k8s-манифест**: `containers[0].env: TRUSTED_HOSTS="tg.workstation.internal"` (после `tty: true`); image/probes/securityContext/resources/PVCs/IngressRoute не тронуты; YAML проверен `yaml.safe_load`.
- **Live-cluster факты**: local-path PVCs, state v2 `accounts=[default, forust]`, смешанное владение 1000/root → обязательный `chown -R 1000:1000 /app/data /app/session` ДО первого старта нового пода.
Тесты: **58 → 64 passed** (+3 TRUSTED_HOSTS env, +3 extended media).
+25 -32
View File
@@ -5,20 +5,21 @@ import shutil
import tempfile import tempfile
import threading import threading
import time import time
from collections.abc import Callable
from copy import deepcopy from copy import deepcopy
from pathlib import Path from pathlib import Path
from typing import Any, Callable, Dict, List, Optional from typing import Any
logger = logging.getLogger(__name__) logger = logging.getLogger(__name__)
# ── defaults ────────────────────────────────────────────────────────── # ── defaults ──────────────────────────────────────────────────────────
GLOBAL_DEFAULTS: Dict[str, Any] = { GLOBAL_DEFAULTS: dict[str, Any] = {
"accounts": [], "accounts": [],
"version": 2, "version": 2,
} }
ACCOUNT_DEFAULTS: Dict[str, Any] = { ACCOUNT_DEFAULTS: dict[str, Any] = {
"label": "", "label": "",
"api_id": None, "api_id": None,
"api_hash": None, "api_hash": None,
@@ -40,15 +41,15 @@ ACCOUNT_DEFAULTS: Dict[str, Any] = {
class StateStore: class StateStore:
"""Thread-safe JSON state store with TTL cache.""" """Thread-safe JSON state store with TTL cache."""
def __init__(self, path: Path, defaults: Optional[Dict[str, Any]] = None): def __init__(self, path: Path, defaults: dict[str, Any] | None = None):
self.path = path self.path = path
self.defaults = defaults or {} self.defaults = defaults or {}
self.lock = threading.RLock() self.lock = threading.RLock()
self._cache: Optional[Dict[str, Any]] = None self._cache: dict[str, Any] | None = None
self._cache_time: float = 0 self._cache_time: float = 0
self._cache_ttl: float = 1.0 self._cache_ttl: float = 1.0
def load(self) -> Dict[str, Any]: def load(self) -> dict[str, Any]:
with self.lock: with self.lock:
now = time.time() now = time.time()
if self._cache is not None and (now - self._cache_time) < self._cache_ttl: if self._cache is not None and (now - self._cache_time) < self._cache_ttl:
@@ -60,7 +61,7 @@ class StateStore:
return deepcopy(result) return deepcopy(result)
try: try:
with self.path.open("r", encoding="utf-8") as handle: with self.path.open("r", encoding="utf-8") as handle:
state: Dict[str, Any] = json.load(handle) state: dict[str, Any] = json.load(handle)
except (json.JSONDecodeError, OSError): except (json.JSONDecodeError, OSError):
result = deepcopy(self.defaults) result = deepcopy(self.defaults)
self._cache = result self._cache = result
@@ -71,7 +72,7 @@ class StateStore:
self._cache_time = now self._cache_time = now
return deepcopy(result) return deepcopy(result)
def save(self, state: Dict[str, Any]) -> None: def save(self, state: dict[str, Any]) -> None:
with self.lock: with self.lock:
self.path.parent.mkdir(parents=True, exist_ok=True, mode=0o700) self.path.parent.mkdir(parents=True, exist_ok=True, mode=0o700)
merged = self._merge_defaults(state) merged = self._merge_defaults(state)
@@ -141,19 +142,19 @@ class StateStore:
except OSError: except OSError:
pass pass
def update(self, mutator: Callable[[Dict[str, Any]], None]) -> Dict[str, Any]: def update(self, mutator: Callable[[dict[str, Any]], None]) -> dict[str, Any]:
with self.lock: with self.lock:
state = self.load() state = self.load()
mutator(state) mutator(state)
self.save(state) self.save(state)
return state return state
def continuous_config(self) -> Dict[str, Any]: def continuous_config(self) -> dict[str, Any]:
state = self.load() state = self.load()
return deepcopy(state.get("continuous_scraping") or ACCOUNT_DEFAULTS["continuous_scraping"]) return deepcopy(state.get("continuous_scraping") or ACCOUNT_DEFAULTS["continuous_scraping"])
def save_continuous_config(self, config: Dict[str, Any]) -> Dict[str, Any]: def save_continuous_config(self, config: dict[str, Any]) -> dict[str, Any]:
def mutate(state: Dict[str, Any]) -> None: def mutate(state: dict[str, Any]) -> None:
state["continuous_scraping"] = { state["continuous_scraping"] = {
"enabled": bool(config.get("enabled", True)), "enabled": bool(config.get("enabled", True)),
"interval_minutes": max(1, int(config.get("interval_minutes", 1) or 1)), "interval_minutes": max(1, int(config.get("interval_minutes", 1) or 1)),
@@ -167,7 +168,7 @@ class StateStore:
return self.update(mutate)["continuous_scraping"] return self.update(mutate)["continuous_scraping"]
def _merge_defaults(self, state: Dict[str, Any]) -> Dict[str, Any]: def _merge_defaults(self, state: dict[str, Any]) -> dict[str, Any]:
merged = deepcopy(self.defaults) merged = deepcopy(self.defaults)
for key, value in state.items(): for key, value in state.items():
if isinstance(value, dict) and isinstance(merged.get(key), dict): if isinstance(value, dict) and isinstance(merged.get(key), dict):
@@ -181,7 +182,7 @@ class StateStore:
# ── global state helpers ────────────────────────────────────────────── # ── global state helpers ──────────────────────────────────────────────
_GLOBAL_STORE: Optional[StateStore] = None _GLOBAL_STORE: StateStore | None = None
def get_global_store(data_dir: Path) -> StateStore: def get_global_store(data_dir: Path) -> StateStore:
@@ -203,15 +204,15 @@ def get_global_store(data_dir: Path) -> StateStore:
return _GLOBAL_STORE return _GLOBAL_STORE
def load_global(data_dir: Path) -> Dict[str, Any]: def load_global(data_dir: Path) -> dict[str, Any]:
return get_global_store(data_dir).load() return get_global_store(data_dir).load()
def save_global(data_dir: Path, state: Dict[str, Any]) -> None: def save_global(data_dir: Path, state: dict[str, Any]) -> None:
get_global_store(data_dir).save(state) get_global_store(data_dir).save(state)
def list_accounts(data_dir: Path) -> List[str]: def list_accounts(data_dir: Path) -> list[str]:
return list(load_global(data_dir).get("accounts", [])) return list(load_global(data_dir).get("accounts", []))
@@ -221,7 +222,7 @@ def account_exists(data_dir: Path, account_id: str) -> bool:
# ── per-account state helpers ───────────────────────────────────────── # ── per-account state helpers ─────────────────────────────────────────
_ACCOUNT_STORES: Dict[str, StateStore] = {} _ACCOUNT_STORES: dict[str, StateStore] = {}
_account_stores_lock = threading.Lock() _account_stores_lock = threading.Lock()
@@ -238,11 +239,11 @@ def get_account_store(data_dir: Path, account_id: str) -> StateStore:
return _ACCOUNT_STORES[key] return _ACCOUNT_STORES[key]
def load_account(data_dir: Path, account_id: str) -> Dict[str, Any]: def load_account(data_dir: Path, account_id: str) -> dict[str, Any]:
return get_account_store(data_dir, account_id).load() return get_account_store(data_dir, account_id).load()
def save_account(data_dir: Path, account_id: str, state: Dict[str, Any]) -> None: def save_account(data_dir: Path, account_id: str, state: dict[str, Any]) -> None:
get_account_store(data_dir, account_id).save(state) get_account_store(data_dir, account_id).save(state)
@@ -268,24 +269,16 @@ def _is_valid_channel_id(channel_id: str) -> bool:
which would be circular) so it can be shared by migration. which would be circular) so it can be shared by migration.
""" """
channel_id = str(channel_id or "").strip() channel_id = str(channel_id or "").strip()
if ( return not (not channel_id or "/" in channel_id or "\\" in channel_id or channel_id in {".", ".."} or any(ord(ch) < 32 for ch in channel_id))
not channel_id
or "/" in channel_id
or "\\" in channel_id
or channel_id in {".", ".."}
or any(ord(ch) < 32 for ch in channel_id)
):
return False
return True
def _clean_continuous_channels(channels: Any) -> List[str]: def _clean_continuous_channels(channels: Any) -> list[str]:
"""Normalize/drop invalid continuous-scraping channel entries during """Normalize/drop invalid continuous-scraping channel entries during
migration. Mirrors webui_server.clean_continuous_channels: strips a migration. Mirrors webui_server.clean_continuous_channels: strips a
leading ``@``, keeps numbers/names, and drops unsafe entries so they can leading ``@``, keeps numbers/names, and drops unsafe entries so they can
never become a path-traversal vector. never become a path-traversal vector.
""" """
cleaned: List[str] = [] cleaned: list[str] = []
if not isinstance(channels, list): if not isinstance(channels, list):
return cleaned return cleaned
for item in channels: for item in channels:
@@ -342,7 +335,7 @@ def migrate_legacy_state(data_dir: Path, session_dir: Path) -> bool:
try: try:
with state_path.open("r", encoding="utf-8") as f: with state_path.open("r", encoding="utf-8") as f:
raw: Dict[str, Any] = json.load(f) raw: dict[str, Any] = json.load(f)
except (json.JSONDecodeError, OSError): except (json.JSONDecodeError, OSError):
return False return False
+10 -10
View File
@@ -1,6 +1,6 @@
import sqlite3 import sqlite3
from pathlib import Path from pathlib import Path
from typing import Any, Dict, List, Optional from typing import Any
from app_state import StateStore, load_account from app_state import StateStore, load_account
@@ -9,10 +9,10 @@ def health_payload(
data_dir: Path, data_dir: Path,
session_dir: Path, session_dir: Path,
state_store: StateStore, state_store: StateStore,
continuous_snapshot: Dict[str, Any], continuous_snapshot: dict[str, Any],
job_queue_size: int, job_queue_size: int,
account_ids: Optional[List[str]] = None, account_ids: list[str] | None = None,
) -> Dict[str, Any]: ) -> dict[str, Any]:
""" """
Return a health-check payload for the application. Return a health-check payload for the application.
@@ -32,7 +32,7 @@ def health_payload(
# Per-account health # Per-account health
if account_ids: if account_ids:
account_checks: Dict[str, Any] = {} account_checks: dict[str, Any] = {}
for acc_id in account_ids: for acc_id in account_ids:
acc_dir = data_dir / "accounts" / acc_id acc_dir = data_dir / "accounts" / acc_id
session_file = session_dir / f"{acc_id}.session" session_file = session_dir / f"{acc_id}.session"
@@ -55,10 +55,10 @@ def health_payload(
return {"ok": ok, "status": "ok" if ok else "degraded", "checks": checks} return {"ok": ok, "status": "ok" if ok else "degraded", "checks": checks}
def _dir_check(path: Path, writable: bool = False) -> Dict[str, Any]: def _dir_check(path: Path, writable: bool = False) -> dict[str, Any]:
path.mkdir(parents=True, exist_ok=True) path.mkdir(parents=True, exist_ok=True)
ok = path.exists() and path.is_dir() ok = path.exists() and path.is_dir()
payload: Dict[str, Any] = {"ok": ok, "path": str(path)} payload: dict[str, Any] = {"ok": ok, "path": str(path)}
if writable: if writable:
probe = path / ".healthcheck" probe = path / ".healthcheck"
try: try:
@@ -70,7 +70,7 @@ def _dir_check(path: Path, writable: bool = False) -> Dict[str, Any]:
return payload return payload
def _state_check(state_store: StateStore) -> Dict[str, Any]: def _state_check(state_store: StateStore) -> dict[str, Any]:
try: try:
state = state_store.load() state = state_store.load()
data_dir = state_store.path.parent data_dir = state_store.path.parent
@@ -105,7 +105,7 @@ def _state_check(state_store: StateStore) -> Dict[str, Any]:
return {"ok": False, "path": str(state_store.path), "error": str(exc)} return {"ok": False, "path": str(state_store.path), "error": str(exc)}
def _sqlite_check() -> Dict[str, Any]: def _sqlite_check() -> dict[str, Any]:
try: try:
conn = sqlite3.connect(":memory:") conn = sqlite3.connect(":memory:")
conn.execute("SELECT 1") conn.execute("SELECT 1")
@@ -115,7 +115,7 @@ def _sqlite_check() -> Dict[str, Any]:
return {"ok": False, "error": str(exc)} return {"ok": False, "error": str(exc)}
def _continuous_check(snapshot: Dict[str, Any]) -> Dict[str, Any]: def _continuous_check(snapshot: dict[str, Any]) -> dict[str, Any]:
running_accounts = snapshot.get("running_accounts") running_accounts = snapshot.get("running_accounts")
if running_accounts is None: if running_accounts is None:
running_accounts = snapshot running_accounts = snapshot
+1 -1
View File
@@ -43,7 +43,7 @@ spec:
tty: true tty: true
env: env:
- name: TRUSTED_HOSTS - name: TRUSTED_HOSTS
value: "tg.workstation.internal" value: 'tg.workstation.internal'
ports: ports:
- containerPort: 8080 - containerPort: 8080
resources: resources:
+7 -7
View File
@@ -2,7 +2,7 @@ import asyncio
import logging import logging
import threading import threading
from pathlib import Path from pathlib import Path
from typing import Any, Dict, List, Optional from typing import Any
from app_state import StateStore from app_state import StateStore
@@ -46,7 +46,7 @@ class ScraperJobService:
def __init__(self, state_store: StateStore): def __init__(self, state_store: StateStore):
self.state_store = state_store self.state_store = state_store
def run(self, job_type: str, payload: Dict[str, Any]) -> None: def run(self, job_type: str, payload: dict[str, Any]) -> None:
if job_type == "set_scrape_media": if job_type == "set_scrape_media":
# The webui handler already persists the scrape_media setting to # The webui handler already persists the scrape_media setting to
# the per-account (or legacy) store before enqueueing this job. # the per-account (or legacy) store before enqueueing this job.
@@ -61,9 +61,9 @@ class ScraperJobService:
# "asyncio.run() cannot be called from a running event loop". # "asyncio.run() cannot be called from a running event loop".
_run_in_new_loop(lambda: self._run_async(job_type, payload)) _run_in_new_loop(lambda: self._run_async(job_type, payload))
async def _run_async(self, job_type: str, payload: Dict[str, Any]) -> None: async def _run_async(self, job_type: str, payload: dict[str, Any]) -> None:
# Extract account_id from payload, default to None (legacy) # Extract account_id from payload, default to None (legacy)
account_id: Optional[str] = payload.get("account_id") account_id: str | None = payload.get("account_id")
ScraperClass = self._import_scraper_class() ScraperClass = self._import_scraper_class()
scraper = ScraperClass(account_id=account_id, base_dir=BASE_DIR) scraper = ScraperClass(account_id=account_id, base_dir=BASE_DIR)
@@ -109,18 +109,18 @@ class ScraperJobService:
if scraper.client: if scraper.client:
await scraper.client.disconnect() await scraper.client.disconnect()
async def _scrape_channels(self, scraper, channels: List[str]) -> None: async def _scrape_channels(self, scraper, channels: list[str]) -> None:
"""Scrape all channels resiliently: a single channel failure does not """Scrape all channels resiliently: a single channel failure does not
abort the rest. Each channel's offset is persisted even on failure abort the rest. Each channel's offset is persisted even on failure
(see scrape_channel's finally block), so partial progress is retained. (see scrape_channel's finally block), so partial progress is retained.
If *every* channel fails, raise so the job is marked failed. If *every* channel fails, raise so the job is marked failed.
""" """
failed: List[str] = [] failed: list[str] = []
for channel_id in channels: for channel_id in channels:
offset = int(scraper.state.get("channels", {}).get(channel_id, 0) or 0) offset = int(scraper.state.get("channels", {}).get(channel_id, 0) or 0)
try: try:
ok = await scraper.scrape_channel(channel_id, offset) ok = await scraper.scrape_channel(channel_id, offset)
except Exception: # noqa: BLE001 - scrape_channel re-raises some errors except Exception:
logger.exception("Scrape of channel %s raised", channel_id) logger.exception("Scrape of channel %s raised", channel_id)
failed.append(channel_id) failed.append(channel_id)
continue continue
+1 -4
View File
@@ -5,7 +5,6 @@ import sys
import time import time
import urllib.request import urllib.request
PORT = os.environ.get("TELEGRAM_SCRAPER_SMOKE_PORT", "18080") PORT = os.environ.get("TELEGRAM_SCRAPER_SMOKE_PORT", "18080")
BASE_URL = f"http://127.0.0.1:{PORT}" BASE_URL = f"http://127.0.0.1:{PORT}"
ENDPOINTS = [ ENDPOINTS = [
@@ -61,9 +60,7 @@ def main() -> int:
if status != 200: if status != 200:
raise RuntimeError(f"{endpoint} returned HTTP {status}") raise RuntimeError(f"{endpoint} returned HTTP {status}")
if ( if (
endpoint.endswith(".json") endpoint.endswith(".json") or endpoint.startswith(("/api", "/health"))
or endpoint.startswith("/api")
or endpoint.startswith("/health")
): ):
json.loads(body.decode("utf-8")) json.loads(body.decode("utf-8"))
print(f"ok {endpoint}") print(f"ok {endpoint}")
+28 -27
View File
@@ -1,27 +1,28 @@
import asyncio
import csv
import json
import os import os
import sqlite3 import sqlite3
import json
import csv
import asyncio
import time
import sys import sys
import time
import warnings import warnings
from dataclasses import dataclass from dataclasses import dataclass
from typing import Dict, List, Optional, Any
from pathlib import Path
from io import StringIO from io import StringIO
from pathlib import Path
from typing import Any
import qrcode
from telethon import TelegramClient from telethon import TelegramClient
from telethon.errors import FloodWaitError, SessionPasswordNeededError
from telethon.tl.types import ( from telethon.tl.types import (
MessageMediaPhoto,
MessageMediaDocument,
MessageMediaWebPage,
User,
PeerChannel,
Channel, Channel,
Chat, Chat,
MessageMediaDocument,
MessageMediaPhoto,
MessageMediaWebPage,
PeerChannel,
User,
) )
from telethon.errors import FloodWaitError, SessionPasswordNeededError
import qrcode
warnings.filterwarnings( warnings.filterwarnings(
"ignore", message="Using async sessions support is an experimental feature" "ignore", message="Using async sessions support is an experimental feature"
@@ -47,17 +48,17 @@ class MessageData:
message_id: int message_id: int
date: str date: str
sender_id: int sender_id: int
first_name: Optional[str] first_name: str | None
last_name: Optional[str] last_name: str | None
username: Optional[str] username: str | None
message: str message: str
media_type: Optional[str] media_type: str | None
media_path: Optional[str] media_path: str | None
reply_to: Optional[int] reply_to: int | None
post_author: Optional[str] post_author: str | None
views: Optional[int] views: int | None
forwards: Optional[int] forwards: int | None
reactions: Optional[str] reactions: str | None
class OptimizedTelegramScraper: class OptimizedTelegramScraper:
@@ -71,7 +72,7 @@ class OptimizedTelegramScraper:
self.state_save_interval = 50 self.state_save_interval = 50
self.db_connections = {} self.db_connections = {}
def load_state(self) -> Dict[str, Any]: def load_state(self) -> dict[str, Any]:
if os.path.exists(self.STATE_FILE): if os.path.exists(self.STATE_FILE):
try: try:
with open(self.STATE_FILE, "r") as f: with open(self.STATE_FILE, "r") as f:
@@ -148,7 +149,7 @@ class OptimizedTelegramScraper:
conn.close() conn.close()
self.db_connections.clear() self.db_connections.clear()
def batch_insert_messages(self, channel: str, messages: List[MessageData]): def batch_insert_messages(self, channel: str, messages: list[MessageData]):
if not messages: if not messages:
return return
@@ -183,7 +184,7 @@ class OptimizedTelegramScraper:
) )
conn.commit() conn.commit()
async def download_media(self, channel: str, message) -> Optional[str]: async def download_media(self, channel: str, message) -> str | None:
if not message.media or not self.state["scrape_media"]: if not message.media or not self.state["scrape_media"]:
return None return None
@@ -701,7 +702,7 @@ class OptimizedTelegramScraper:
async for dialog in self.client.iter_dialogs(): async for dialog in self.client.iter_dialogs():
entity = dialog.entity entity = dialog.entity
if dialog.id != 777000 and ( if dialog.id != 777000 and (
isinstance(entity, Channel) or isinstance(entity, Chat) isinstance(entity, (Channel, Chat))
): ):
channel_type = ( channel_type = (
"Channel" "Channel"
+31 -29
View File
@@ -1,27 +1,29 @@
import sqlite3
import json
import csv
import asyncio import asyncio
import csv
import json
import logging import logging
import time import sqlite3
import sys import sys
import time
import warnings import warnings
from dataclasses import dataclass from dataclasses import dataclass
from typing import Dict, List, Optional, Any
from pathlib import Path
from io import StringIO from io import StringIO
from pathlib import Path
from typing import Any
import qrcode
from telethon import TelegramClient, events from telethon import TelegramClient, events
from telethon.errors import FloodWaitError, SessionPasswordNeededError
from telethon.tl.types import ( from telethon.tl.types import (
MessageMediaPhoto,
MessageMediaDocument,
MessageMediaWebPage,
User,
PeerChannel,
Channel, Channel,
Chat, Chat,
MessageMediaDocument,
MessageMediaPhoto,
MessageMediaWebPage,
PeerChannel,
User,
) )
from telethon.errors import FloodWaitError, SessionPasswordNeededError
import qrcode
from app_state import ( from app_state import (
StateStore, StateStore,
account_session_path, account_session_path,
@@ -55,17 +57,17 @@ class MessageData:
message_id: int message_id: int
date: str date: str
sender_id: int sender_id: int
first_name: Optional[str] first_name: str | None
last_name: Optional[str] last_name: str | None
username: Optional[str] username: str | None
message: str message: str
media_type: Optional[str] media_type: str | None
media_path: Optional[str] media_path: str | None
reply_to: Optional[int] reply_to: int | None
post_author: Optional[str] post_author: str | None
views: Optional[int] views: int | None
forwards: Optional[int] forwards: int | None
reactions: Optional[str] reactions: str | None
@dataclass @dataclass
@@ -94,7 +96,7 @@ def _ensure_session_wal(session_path: str) -> None:
class OptimizedTelegramScraper: class OptimizedTelegramScraper:
def __init__(self, account_id: Optional[str] = None, base_dir: Optional[Path] = None): def __init__(self, account_id: str | None = None, base_dir: Path | None = None):
self.account_id = account_id self.account_id = account_id
base_dir = base_dir or BASE_DIR base_dir = base_dir or BASE_DIR
self.BASE_DIR = base_dir self.BASE_DIR = base_dir
@@ -123,7 +125,7 @@ class OptimizedTelegramScraper:
self.db_connections = {} self.db_connections = {}
self.forwarding_handler = None self.forwarding_handler = None
def load_state(self) -> Dict[str, Any]: def load_state(self) -> dict[str, Any]:
return self.state_store.load() return self.state_store.load()
def save_state(self): def save_state(self):
@@ -148,7 +150,7 @@ class OptimizedTelegramScraper:
): ):
self.save_state() self.save_state()
def get_forwarding_rules(self) -> List[ForwardingRule]: def get_forwarding_rules(self) -> list[ForwardingRule]:
rules = [] rules = []
for rule_dict in self.state.get("forwarding_rules", []): for rule_dict in self.state.get("forwarding_rules", []):
rules.append( rules.append(
@@ -273,7 +275,7 @@ class OptimizedTelegramScraper:
conn.close() conn.close()
self.db_connections.clear() self.db_connections.clear()
def batch_insert_messages(self, channel: str, messages: List[MessageData]): def batch_insert_messages(self, channel: str, messages: list[MessageData]):
if not messages: if not messages:
return return
@@ -308,7 +310,7 @@ class OptimizedTelegramScraper:
) )
conn.commit() conn.commit()
async def download_media(self, channel: str, message) -> Optional[str]: async def download_media(self, channel: str, message) -> str | None:
if not message.media or not self.state["scrape_media"]: if not message.media or not self.state["scrape_media"]:
return None return None
@@ -402,7 +404,7 @@ class OptimizedTelegramScraper:
return False return False
async def forward_message( async def forward_message(
self, message, rule: ForwardingRule, source_channel_id: int = None, _retry: int = 0 self, message, rule: ForwardingRule, source_channel_id: int | None = None, _retry: int = 0
): ):
try: try:
dest_entity = await self._resolve_entity(rule.destination_channel) dest_entity = await self._resolve_entity(rule.destination_channel)
+16 -15
View File
@@ -23,7 +23,7 @@ import tempfile
import threading import threading
import time import time
from pathlib import Path from pathlib import Path
from typing import Any, Dict, List, Optional from typing import Any
from unittest.mock import MagicMock from unittest.mock import MagicMock
# ── Mock heavy dependencies before any webui_server import ───────────── # ── Mock heavy dependencies before any webui_server import ─────────────
@@ -51,8 +51,8 @@ os.environ["TELEGRAM_SCRAPER_PORT"] = "0"
os.environ["TELEGRAM_SCRAPER_START_CONTINUOUS"] = "0" os.environ["TELEGRAM_SCRAPER_START_CONTINUOUS"] = "0"
# Now safe to import app modules # Now safe to import app modules
import app_state # noqa: E402 import app_state
from app_state import ( # noqa: E402 from app_state import (
ACCOUNT_DEFAULTS, ACCOUNT_DEFAULTS,
account_data_dir, account_data_dir,
account_exists, account_exists,
@@ -62,7 +62,7 @@ from app_state import ( # noqa: E402
list_accounts, list_accounts,
load_account, load_account,
) )
from health import health_payload # noqa: E402 from health import health_payload
def _init_test_env(): def _init_test_env():
@@ -82,14 +82,14 @@ def make_account_id() -> str:
return f"test-{int(time.time() * 1000000)}" return f"test-{int(time.time() * 1000000)}"
def create_account(data_dir: Path, account_id: str, **overrides) -> Dict[str, Any]: def create_account(data_dir: Path, account_id: str, **overrides) -> dict[str, Any]:
store = get_account_store(data_dir, account_id) store = get_account_store(data_dir, account_id)
state = dict(ACCOUNT_DEFAULTS) state = dict(ACCOUNT_DEFAULTS)
state.update(overrides) state.update(overrides)
state["label"] = overrides.get("label", account_id) state["label"] = overrides.get("label", account_id)
store.save(state) store.save(state)
def mutate(g: Dict[str, Any]) -> None: def mutate(g: dict[str, Any]) -> None:
accounts = g.setdefault("accounts", []) accounts = g.setdefault("accounts", [])
if account_id not in accounts: if account_id not in accounts:
accounts.append(account_id) accounts.append(account_id)
@@ -98,7 +98,7 @@ def create_account(data_dir: Path, account_id: str, **overrides) -> Dict[str, An
return load_account(data_dir, account_id) return load_account(data_dir, account_id)
def create_channel_db(data_dir: Path, account_id: Optional[str], channel_id: str, messages: List[Dict[str, Any]]) -> None: def create_channel_db(data_dir: Path, account_id: str | None, channel_id: str, messages: list[dict[str, Any]]) -> None:
if account_id: if account_id:
db_dir = account_data_dir(data_dir, account_id) / channel_id db_dir = account_data_dir(data_dir, account_id) / channel_id
else: else:
@@ -378,7 +378,7 @@ class TestContinuousOrchestrator:
def test_does_not_auto_start_on_import(self): def test_does_not_auto_start_on_import(self):
"""Importing an account must not trigger continuous scraping.""" """Importing an account must not trigger continuous scraping."""
PerAccountContinuousScrapeManager, ContinuousScrapeOrchestrator, ws = self._import_orch_classes() _PerAccountContinuousScrapeManager, ContinuousScrapeOrchestrator, _ws = self._import_orch_classes()
self._setup_ws_data_dir() self._setup_ws_data_dir()
try: try:
orch = ContinuousScrapeOrchestrator() orch = ContinuousScrapeOrchestrator()
@@ -394,7 +394,7 @@ class TestContinuousOrchestrator:
def test_no_duplicate_workers_on_same_account(self): def test_no_duplicate_workers_on_same_account(self):
"""Multiple start() calls should not spawn duplicate threads.""" """Multiple start() calls should not spawn duplicate threads."""
PerAccountContinuousScrapeManager, _, ws = self._import_orch_classes() PerAccountContinuousScrapeManager, _, _ws = self._import_orch_classes()
self._setup_ws_data_dir() self._setup_ws_data_dir()
try: try:
aid = make_account_id() aid = make_account_id()
@@ -424,7 +424,7 @@ class TestContinuousOrchestrator:
def test_disabled_account_not_started(self): def test_disabled_account_not_started(self):
"""start_all() must not start accounts with enabled=False.""" """start_all() must not start accounts with enabled=False."""
_, ContinuousScrapeOrchestrator, ws = self._import_orch_classes() _, ContinuousScrapeOrchestrator, _ws = self._import_orch_classes()
self._setup_ws_data_dir() self._setup_ws_data_dir()
try: try:
orch = ContinuousScrapeOrchestrator() orch = ContinuousScrapeOrchestrator()
@@ -442,7 +442,7 @@ class TestContinuousOrchestrator:
"""stop()+join() must let status['running'] become False only after the """stop()+join() must let status['running'] become False only after the
worker thread truly exits, and join must be safe/idempotent on a worker thread truly exits, and join must be safe/idempotent on a
short-lived thread.""" short-lived thread."""
PerAccountContinuousScrapeManager, _, ws = self._import_orch_classes() PerAccountContinuousScrapeManager, _, _ws = self._import_orch_classes()
self._setup_ws_data_dir() self._setup_ws_data_dir()
try: try:
aid = make_account_id() aid = make_account_id()
@@ -479,7 +479,7 @@ class TestContinuousOrchestrator:
def test_refresh_config_disable_requests_stop_without_lying_about_running(self): def test_refresh_config_disable_requests_stop_without_lying_about_running(self):
"""refresh_config() disabling the account must request a stop but must """refresh_config() disabling the account must request a stop but must
NOT set status['running']=False (the worker thread owns that).""" NOT set status['running']=False (the worker thread owns that)."""
PerAccountContinuousScrapeManager, _, ws = self._import_orch_classes() PerAccountContinuousScrapeManager, _, _ws = self._import_orch_classes()
self._setup_ws_data_dir() self._setup_ws_data_dir()
try: try:
aid = make_account_id() aid = make_account_id()
@@ -825,7 +825,7 @@ class TestSecurityHardening:
assert handler.read_json_body() == {} assert handler.read_json_body() == {}
def test_check_same_origin_rejects_cross_origin(self): def test_check_same_origin_rejects_cross_origin(self):
ws = self._ws_module() self._ws_module()
# cross-origin Origin header -> rejected (403 response sent) # cross-origin Origin header -> rejected (403 response sent)
h = _make_ws_handler( h = _make_ws_handler(
@@ -1564,5 +1564,6 @@ def cleanup_test_data():
shutil.rmtree(str(TEST_TMP), ignore_errors=True) shutil.rmtree(str(TEST_TMP), ignore_errors=True)
import atexit # noqa: E402 import atexit
atexit.register(cleanup_test_data) # noqa: E402
atexit.register(cleanup_test_data)
+46 -42
View File
@@ -408,7 +408,8 @@ function renderJobs(accountId, jobs) {
root.innerHTML = ''; root.innerHTML = '';
if (!jobs.length) { if (!jobs.length) {
root.innerHTML = '<div class="empty-state"><p class="muted">No jobs yet. Start a scrape or export to see progress here.</p></div>'; root.innerHTML =
'<div class="empty-state"><p class="muted">No jobs yet. Start a scrape or export to see progress here.</p></div>';
return; return;
} }
@@ -947,55 +948,58 @@ async function main() {
const addAccountForm = document.getElementById('add-account-form'); const addAccountForm = document.getElementById('add-account-form');
if (addAccountForm) { if (addAccountForm) {
addAccountForm.addEventListener('submit', async (event) => { addAccountForm.addEventListener('submit', async (event) => {
event.preventDefault(); event.preventDefault();
const accountId = document.getElementById('add-account-id').value.trim(); const accountId = document.getElementById('add-account-id').value.trim();
const label = document.getElementById('add-account-label').value.trim(); const label = document.getElementById('add-account-label').value.trim();
const apiId = document.getElementById('add-account-api-id').value.trim(); const apiId = document.getElementById('add-account-api-id').value.trim();
const apiHash = document.getElementById('add-account-api-hash').value.trim(); const apiHash = document.getElementById('add-account-api-hash').value.trim();
if (!accountId) return; if (!accountId) return;
try { try {
await api('/api/accounts', { await api('/api/accounts', {
method: 'POST', method: 'POST',
body: JSON.stringify({ body: JSON.stringify({
account_id: accountId, account_id: accountId,
label: label || accountId, label: label || accountId,
api_id: apiId ? parseInt(apiId) : undefined, api_id: apiId ? parseInt(apiId) : undefined,
api_hash: apiHash || undefined, api_hash: apiHash || undefined,
}), }),
}); });
event.target.reset(); event.target.reset();
await loadAccounts(); await loadAccounts();
if (!state.activeAccount) { if (!state.activeAccount) {
switchAccount(accountId); switchAccount(accountId);
}
} catch (err) {
showToast('Failed to add account: ' + err.message, 'error');
} }
} catch (err) {
showToast('Failed to add account: ' + err.message, 'error');
}
}); });
} }
const importAccountFile = document.getElementById('import-account-file'); const importAccountFile = document.getElementById('import-account-file');
if (importAccountFile) { if (importAccountFile) {
importAccountFile.addEventListener('change', async (event) => { importAccountFile.addEventListener('change', async (event) => {
const file = event.currentTarget.files?.[0]; const file = event.currentTarget.files?.[0];
if (!file) return; if (!file) return;
try { try {
const payload = JSON.parse(await file.text()); const payload = JSON.parse(await file.text());
const accountId = payload.account_id || payload.id; const accountId = payload.account_id || payload.id;
if (!accountId) throw new Error('account_id is missing in import file'); if (!accountId) throw new Error('account_id is missing in import file');
await api('/api/accounts/import', { await api('/api/accounts/import', {
method: 'POST', method: 'POST',
body: JSON.stringify(payload), body: JSON.stringify(payload),
}); });
event.currentTarget.value = ''; event.currentTarget.value = '';
await loadAccounts(); await loadAccounts();
switchAccount(accountId); switchAccount(accountId);
showToast(`Imported ${accountId}.`, 'success'); showToast(`Imported ${accountId}.`, 'success');
showToast('Credentials (api_id/api_hash) are not exported for security — re-enter them in Settings if needed.', 'warn'); showToast(
} catch (err) { 'Credentials (api_id/api_hash) are not exported for security — re-enter them in Settings if needed.',
showToast(`Failed to import account: ${err.message}`, 'error'); 'warn',
} );
} catch (err) {
showToast(`Failed to import account: ${err.message}`, 'error');
}
}); });
} }
+2 -4
View File
@@ -170,14 +170,12 @@
name="channel_id" name="channel_id"
placeholder="ID or @username" placeholder="ID or @username"
aria-label="Channel ID or username" aria-label="Channel ID or username"
required required />
/>
<input <input
class="add-channel-name" class="add-channel-name"
name="name" name="name"
placeholder="Display name" placeholder="Display name"
aria-label="Channel display name" aria-label="Channel display name" />
/>
<button class="button primary" type="submit">Add</button> <button class="button primary" type="submit">Add</button>
</form> </form>
</div> </div>
+102 -101
View File
@@ -113,47 +113,49 @@
</div> </div>
</div> </div>
<div class="panel-body"> <div class="panel-body">
<form id="settings-add-account-form" class="stack-form"> <form id="settings-add-account-form" class="stack-form">
<label class="field"> <label class="field">
<span class="field-label">Account ID</span> <span class="field-label">Account ID</span>
<input <input
id="settings-add-account-id" id="settings-add-account-id"
name="account_id" name="account_id"
placeholder="Account ID (e.g. work)" placeholder="Account ID (e.g. work)"
aria-label="Account ID" aria-label="Account ID"
required required />
/> <span class="field-hint">Lowercase short id used in file names and URLs.</span>
<span class="field-hint">Lowercase short id used in file names and URLs.</span> </label>
<label class="field">
<span class="field-label">Display name</span>
<input
id="settings-add-account-label"
name="label"
placeholder="Display name"
aria-label="Display name" />
</label>
<label class="field">
<span class="field-label">API ID</span>
<input id="settings-add-account-api-id" name="api_id" placeholder="API ID" aria-label="API ID" />
<span class="field-hint">From my.telegram.org &gt; API development tools.</span>
</label>
<label class="field">
<span class="field-label">API Hash</span>
<input
id="settings-add-account-api-hash"
name="api_hash"
placeholder="API Hash"
aria-label="API Hash" />
<span class="field-hint">Paste the hash exactly as shown on my.telegram.org.</span>
</label>
<button class="button primary button-block" type="submit">Add Account</button>
</form>
<label class="import-account-row">
<span class="field-label">Import account settings JSON</span>
<span class="muted small"
>Choose a file previously exported from Settings. It must be account JSON, for example:</span
>
<span class="file-example">{"account_id": "work", "api_id": 12345, ...}</span>
<input id="settings-import-account-file" type="file" accept="application/json,.json" />
</label> </label>
<label class="field">
<span class="field-label">Display name</span>
<input
id="settings-add-account-label"
name="label"
placeholder="Display name"
aria-label="Display name"
/>
</label>
<label class="field">
<span class="field-label">API ID</span>
<input id="settings-add-account-api-id" name="api_id" placeholder="API ID" aria-label="API ID" />
<span class="field-hint">From my.telegram.org &gt; API development tools.</span>
</label>
<label class="field">
<span class="field-label">API Hash</span>
<input id="settings-add-account-api-hash" name="api_hash" placeholder="API Hash" aria-label="API Hash" />
<span class="field-hint">Paste the hash exactly as shown on my.telegram.org.</span>
</label>
<button class="button primary button-block" type="submit">Add Account</button>
</form>
<label class="import-account-row">
<span class="field-label">Import account settings JSON</span>
<span class="muted small"
>Choose a file previously exported from Settings. It must be account JSON, for example:</span
>
<span class="file-example">{"account_id": "work", "api_id": 12345, ...}</span>
<input id="settings-import-account-file" type="file" accept="application/json,.json" />
</label>
</div> </div>
</div> </div>
@@ -165,60 +167,59 @@
</div> </div>
</div> </div>
<div class="panel-body"> <div class="panel-body">
<p class="muted small" id="settings-credentials-hint"> <p class="muted small" id="settings-credentials-hint">
Select an account first, then save API credentials and complete login below. Select an account first, then save API credentials and complete login below.
</p> </p>
<form id="settings-credentials-form" class="stack-form"> <form id="settings-credentials-form" class="stack-form">
<label class="field"> <label class="field">
<span class="field-label">API ID</span> <span class="field-label">API ID</span>
<input id="settings-api-id-input" name="api_id" placeholder="API ID" aria-label="API ID" /> <input id="settings-api-id-input" name="api_id" placeholder="API ID" aria-label="API ID" />
<span class="field-hint">From my.telegram.org &gt; API development tools.</span> <span class="field-hint">From my.telegram.org &gt; API development tools.</span>
</label> </label>
<label class="field"> <label class="field">
<span class="field-label">API Hash</span> <span class="field-label">API Hash</span>
<input id="settings-api-hash-input" name="api_hash" placeholder="API Hash" aria-label="API Hash" /> <input id="settings-api-hash-input" name="api_hash" placeholder="API Hash" aria-label="API Hash" />
</label> </label>
<button class="button" type="submit">Save credentials</button> <button class="button" type="submit">Save credentials</button>
</form> </form>
<div class="auth-actions"> <div class="auth-actions">
<button class="button primary" id="settings-start-qr-btn" type="button">Start QR login</button> <button class="button primary" id="settings-start-qr-btn" type="button">Start QR login</button>
</div> </div>
<div id="settings-qr-wrap" class="qr-wrap hidden"> <div id="settings-qr-wrap" class="qr-wrap hidden">
<img id="settings-qr-image" alt="Telegram QR login" /> <img id="settings-qr-image" alt="Telegram QR login" />
<p class="muted small">Open Telegram -> Settings -> Devices -> Scan QR.</p> <p class="muted small">Open Telegram -> Settings -> Devices -> Scan QR.</p>
</div> </div>
<form id="settings-phone-form" class="stack-form"> <form id="settings-phone-form" class="stack-form">
<label class="field"> <label class="field">
<span class="field-label">Phone number</span> <span class="field-label">Phone number</span>
<input id="settings-phone-input" name="phone" placeholder="+1234567890" aria-label="Phone number" /> <input id="settings-phone-input" name="phone" placeholder="+1234567890" aria-label="Phone number" />
</label> </label>
<button class="button" type="submit">Send code</button> <button class="button" type="submit">Send code</button>
</form> </form>
<form id="settings-code-form" class="stack-form hidden"> <form id="settings-code-form" class="stack-form hidden">
<label class="field"> <label class="field">
<span class="field-label">Telegram code</span> <span class="field-label">Telegram code</span>
<input id="settings-code-input" name="code" placeholder="Telegram code" aria-label="Telegram code" /> <input id="settings-code-input" name="code" placeholder="Telegram code" aria-label="Telegram code" />
</label> </label>
<button class="button" type="submit">Confirm code</button> <button class="button" type="submit">Confirm code</button>
</form> </form>
<form id="settings-password-form" class="stack-form hidden"> <form id="settings-password-form" class="stack-form hidden">
<label class="field"> <label class="field">
<span class="field-label">2FA password</span> <span class="field-label">2FA password</span>
<input <input
id="settings-password-input" id="settings-password-input"
name="password" name="password"
type="password" type="password"
placeholder="2FA password" placeholder="2FA password"
aria-label="2FA password" aria-label="2FA password" />
/> </label>
</label> <button class="button" type="submit">Confirm password</button>
<button class="button" type="submit">Confirm password</button> </form>
</form>
</div> </div>
</div> </div>
@@ -230,18 +231,18 @@
</div> </div>
</div> </div>
<div class="panel-body"> <div class="panel-body">
<label class="toggle-row"> <label class="toggle-row">
<span>Download media</span> <span>Download media</span>
<span class="switch"> <span class="switch">
<input id="settings-scrape-media-toggle" type="checkbox" /> <input id="settings-scrape-media-toggle" type="checkbox" />
<span class="switch-slider"></span> <span class="switch-slider"></span>
</span> </span>
</label> </label>
<div class="settings-meta-list"> <div class="settings-meta-list">
<div>Tracked channels: <span id="settings-channel-count">0</span></div> <div>Tracked channels: <span id="settings-channel-count">0</span></div>
<div>Messages: <span id="settings-message-count">0</span></div> <div>Messages: <span id="settings-message-count">0</span></div>
<div>Media: <span id="settings-media-count">0</span></div> <div>Media: <span id="settings-media-count">0</span></div>
</div> </div>
</div> </div>
</div> </div>
</section> </section>
+7 -6
View File
@@ -207,10 +207,7 @@ function renderAccountData() {
setHealthState(document.getElementById('health-credentials'), Boolean(health.api_credentials)); setHealthState(document.getElementById('health-credentials'), Boolean(health.api_credentials));
document.getElementById('health-session').textContent = document.getElementById('health-session').textContent =
health.session_ready || isAccountAuthorized(auth) ? 'Ready' : 'Missing'; health.session_ready || isAccountAuthorized(auth) ? 'Ready' : 'Missing';
setHealthState( setHealthState(document.getElementById('health-session'), Boolean(health.session_ready || isAccountAuthorized(auth)));
document.getElementById('health-session'),
Boolean(health.session_ready || isAccountAuthorized(auth)),
);
document.getElementById('health-continuous').textContent = continuousStatus.running document.getElementById('health-continuous').textContent = continuousStatus.running
? 'Running' ? 'Running'
: continuousConfig.enabled : continuousConfig.enabled
@@ -298,7 +295,10 @@ async function importAccount(event) {
await loadAccounts(); await loadAccounts();
await loadAccount(accountId); await loadAccount(accountId);
showToast(`Imported ${accountId}.`, 'success'); showToast(`Imported ${accountId}.`, 'success');
showToast('Credentials (api_id/api_hash) are not exported for security — re-enter them in Settings if needed.', 'warn'); showToast(
'Credentials (api_id/api_hash) are not exported for security — re-enter them in Settings if needed.',
'warn',
);
} catch (err) { } catch (err) {
showToast(`Failed to import account: ${err.message}`, 'error'); showToast(`Failed to import account: ${err.message}`, 'error');
} }
@@ -380,7 +380,8 @@ async function exportAccount() {
async function deleteAccount() { async function deleteAccount() {
if (!settingsState.activeAccount) return; if (!settingsState.activeAccount) return;
const account = settingsState.accounts.find((item) => item.id === settingsState.activeAccount); const account = settingsState.accounts.find((item) => item.id === settingsState.activeAccount);
if (!(await confirmAction(`Remove account "${accountLabel(account)}"? All its account data will be deleted.`))) return; if (!(await confirmAction(`Remove account "${accountLabel(account)}"? All its account data will be deleted.`)))
return;
await api(`/api/accounts/${encodeURIComponent(settingsState.activeAccount)}`, { method: 'DELETE' }); await api(`/api/accounts/${encodeURIComponent(settingsState.activeAccount)}`, { method: 'DELETE' });
setActiveAccount(null); setActiveAccount(null);
await loadAccounts(); await loadAccounts();
+2 -1
View File
@@ -38,7 +38,8 @@
--gutter: 18px; --gutter: 18px;
--gap: 16px; --gap: 16px;
--font-ui: Inter, 'SF Pro Display', 'Helvetica Neue', -apple-system, BlinkMacSystemFont, 'Segoe UI', Arial, sans-serif; --font-ui:
Inter, 'SF Pro Display', 'Helvetica Neue', -apple-system, BlinkMacSystemFont, 'Segoe UI', Arial, sans-serif;
--font-mono: 'IBM Plex Mono', 'SFMono-Regular', Consolas, monospace; --font-mono: 'IBM Plex Mono', 'SFMono-Regular', Consolas, monospace;
/* Legacy aliases (single system — all map to opaque tokens) */ /* Legacy aliases (single system — all map to opaque tokens) */
+119 -135
View File
@@ -16,31 +16,31 @@ import traceback
import urllib.parse import urllib.parse
import uuid import uuid
from dataclasses import dataclass, field from dataclasses import dataclass, field
from datetime import datetime, timezone from datetime import UTC, datetime
from http import HTTPStatus from http import HTTPStatus
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
from pathlib import Path from pathlib import Path
from typing import Any, Dict, List, Optional, Tuple from typing import Any
import qrcode import qrcode
import qrcode.image.svg import qrcode.image.svg
from telethon import TelegramClient
from telethon.errors import SessionPasswordNeededError
from app_state import ( from app_state import (
StateStore, StateStore,
account_exists,
account_data_dir, account_data_dir,
account_exists,
account_session_path, account_session_path,
get_account_store, get_account_store,
get_global_store, get_global_store,
load_account,
list_accounts, list_accounts,
load_account,
) )
from health import health_payload from health import health_payload
from scraper_jobs import ScraperJobService from scraper_jobs import ScraperJobService
from telethon import TelegramClient
from telethon.errors import SessionPasswordNeededError
from telegram_scraper_with_forwarding import _ensure_session_wal from telegram_scraper_with_forwarding import _ensure_session_wal
BASE_DIR = Path(__file__).resolve().parent BASE_DIR = Path(__file__).resolve().parent
DATA_DIR = BASE_DIR / "data" DATA_DIR = BASE_DIR / "data"
WEBUI_DIR = BASE_DIR / "webui" WEBUI_DIR = BASE_DIR / "webui"
@@ -88,7 +88,7 @@ AUTH_LOCKOUT_SECONDS = 60
AUTH_CODE_REQUEST_COOLDOWN_SECONDS = 30 AUTH_CODE_REQUEST_COOLDOWN_SECONDS = 30
# Cap the in-memory auth tracking dict; beyond this we sweep expired entries. # Cap the in-memory auth tracking dict; beyond this we sweep expired entries.
AUTH_ATTEMPTS_MAX_ENTRIES = 10_000 AUTH_ATTEMPTS_MAX_ENTRIES = 10_000
_auth_attempts: Dict[Tuple[str, str], Dict[str, Any]] = {} _auth_attempts: dict[tuple[str, str], dict[str, Any]] = {}
_auth_attempts_lock = threading.Lock() _auth_attempts_lock = threading.Lock()
# ── M-3: boolean coercion helper ───────────────────────────────────────── # ── M-3: boolean coercion helper ─────────────────────────────────────────
@@ -116,7 +116,7 @@ def parse_bool(value: Any, default: bool = False) -> bool:
# ── L-8: trusted-host check ────────────────────────────────────────────── # ── L-8: trusted-host check ──────────────────────────────────────────────
import ipaddress # noqa: E402 import ipaddress
# Comma-separated extra trusted hostnames for browser state-mutating # Comma-separated extra trusted hostnames for browser state-mutating
# requests when the panel is served through a proxy domain. # requests when the panel is served through a proxy domain.
@@ -211,9 +211,7 @@ def _check_auth_code_cooldown(ip: str, account_id: str) -> bool:
if not entry: if not entry:
return True return True
cooldown_until = entry.get("cooldown_until") cooldown_until = entry.get("cooldown_until")
if cooldown_until and now < cooldown_until: return not (cooldown_until and now < cooldown_until)
return False
return True
def _record_auth_failure(ip: str, account_id: str) -> None: def _record_auth_failure(ip: str, account_id: str) -> None:
@@ -248,18 +246,18 @@ def _record_auth_code_request(ip: str, account_id: str) -> None:
def utc_now_iso() -> str: def utc_now_iso() -> str:
return datetime.now(timezone.utc).isoformat() return datetime.now(UTC).isoformat()
def load_state() -> Dict[str, Any]: def load_state() -> dict[str, Any]:
return STATE_STORE.load() return STATE_STORE.load()
def save_state(state: Dict[str, Any]) -> None: def save_state(state: dict[str, Any]) -> None:
STATE_STORE.save(state) STATE_STORE.save(state)
def guess_media_kind(media_path: Optional[str], media_type: Optional[str]) -> Optional[str]: def guess_media_kind(media_path: str | None, media_type: str | None) -> str | None:
if not media_path: if not media_path:
return None return None
suffix = Path(media_path).suffix.lower() suffix = Path(media_path).suffix.lower()
@@ -274,7 +272,7 @@ def guess_media_kind(media_path: Optional[str], media_type: Optional[str]) -> Op
return "file" return "file"
def normalize_media_url(media_path: Optional[str]) -> Optional[str]: def normalize_media_url(media_path: str | None) -> str | None:
if not media_path: if not media_path:
return None return None
path = Path(media_path) path = Path(media_path)
@@ -304,7 +302,7 @@ def normalize_channel_id(value: Any) -> str:
return channel_id return channel_id
def clean_continuous_channels(channels: Any) -> Tuple[List[str], List[str]]: def clean_continuous_channels(channels: Any) -> tuple[list[str], list[str]]:
"""Validate/normalize a continuous-scrape channel list at ingest. """Validate/normalize a continuous-scrape channel list at ingest.
Returns ``(cleaned, dropped)`` where ``cleaned`` holds the normalized, Returns ``(cleaned, dropped)`` where ``cleaned`` holds the normalized,
@@ -313,8 +311,8 @@ def clean_continuous_channels(channels: Any) -> Tuple[List[str], List[str]]:
dropped entries must NOT be persisted so they cannot become a path dropped entries must NOT be persisted so they cannot become a path
traversal vector for any future direct read. traversal vector for any future direct read.
""" """
cleaned: List[str] = [] cleaned: list[str] = []
dropped: List[str] = [] dropped: list[str] = []
if not isinstance(channels, list): if not isinstance(channels, list):
return cleaned, dropped return cleaned, dropped
for item in channels: for item in channels:
@@ -326,7 +324,7 @@ def clean_continuous_channels(channels: Any) -> Tuple[List[str], List[str]]:
return cleaned, dropped return cleaned, dropped
def export_account_state(state: Dict[str, Any]) -> Dict[str, Any]: def export_account_state(state: dict[str, Any]) -> dict[str, Any]:
"""Return a sanitized copy of an account state for export. """Return a sanitized copy of an account state for export.
Credentials (api_hash, api_id) are always redacted - they are never Credentials (api_hash, api_id) are always redacted - they are never
@@ -342,15 +340,15 @@ def export_account_state(state: Dict[str, Any]) -> Dict[str, Any]:
return exported return exported
def channel_db_path(account_id: Optional[str], channel_id: str) -> Path: def channel_db_path(account_id: str | None, channel_id: str) -> Path:
if account_id: if account_id:
return account_data_dir(DATA_DIR, account_id) / channel_id / f"{channel_id}.db" return account_data_dir(DATA_DIR, account_id) / channel_id / f"{channel_id}.db"
return DATA_DIR / channel_id / f"{channel_id}.db" return DATA_DIR / channel_id / f"{channel_id}.db"
def database_summary(account_id: Optional[str], channel_id: str) -> Dict[str, Any]: def database_summary(account_id: str | None, channel_id: str) -> dict[str, Any]:
db_path = channel_db_path(account_id, channel_id) db_path = channel_db_path(account_id, channel_id)
summary: Dict[str, Any] = { summary: dict[str, Any] = {
"message_count": 0, "message_count": 0,
"last_date": None, "last_date": None,
"first_date": None, "first_date": None,
@@ -398,16 +396,16 @@ def database_summary(account_id: Optional[str], channel_id: str) -> Dict[str, An
conn.close() conn.close()
def channel_display_name(state: Dict[str, Any], channel_id: str) -> str: def channel_display_name(state: dict[str, Any], channel_id: str) -> str:
return state.get("channel_names", {}).get(channel_id) or channel_id return state.get("channel_names", {}).get(channel_id) or channel_id
def list_channels_snapshot(account_id: Optional[str] = None) -> List[Dict[str, Any]]: def list_channels_snapshot(account_id: str | None = None) -> list[dict[str, Any]]:
if account_id: if account_id:
state = load_account(DATA_DIR, account_id) state = load_account(DATA_DIR, account_id)
else: else:
state = load_state() state = load_state()
channels: List[Dict[str, Any]] = [] channels: list[dict[str, Any]] = []
for channel_id, last_message_id in state.get("channels", {}).items(): for channel_id, last_message_id in state.get("channels", {}).items():
summary = database_summary(account_id, channel_id) summary = database_summary(account_id, channel_id)
channels.append( channels.append(
@@ -425,7 +423,7 @@ def list_channels_snapshot(account_id: Optional[str] = None) -> List[Dict[str, A
return channels return channels
def account_state_for_legacy() -> Optional[Dict[str, Any]]: def account_state_for_legacy() -> dict[str, Any] | None:
legacy = load_state() legacy = load_state()
if legacy.get("api_id") and legacy.get("api_hash"): if legacy.get("api_id") and legacy.get("api_hash"):
return legacy return legacy
@@ -433,12 +431,12 @@ def account_state_for_legacy() -> Optional[Dict[str, Any]]:
def load_messages( def load_messages(
account_id: Optional[str], account_id: str | None,
channel_id: str, channel_id: str,
limit: int = 120, limit: int = 120,
before_message_id: Optional[int] = None, before_message_id: int | None = None,
search: Optional[str] = None, search: str | None = None,
) -> List[Dict[str, Any]]: ) -> list[dict[str, Any]]:
db_path = channel_db_path(account_id, channel_id) db_path = channel_db_path(account_id, channel_id)
if not db_path.exists(): if not db_path.exists():
return [] return []
@@ -450,8 +448,8 @@ def load_messages(
"media_type, media_path, reply_to, post_author, views, forwards, reactions " "media_type, media_path, reply_to, post_author, views, forwards, reactions "
"FROM messages " "FROM messages "
) )
params: List[Any] = [] params: list[Any] = []
where: List[str] = [] where: list[str] = []
if before_message_id is not None: if before_message_id is not None:
where.append("message_id < ?") where.append("message_id < ?")
params.append(before_message_id) params.append(before_message_id)
@@ -466,7 +464,7 @@ def load_messages(
finally: finally:
conn.close() conn.close()
messages: List[Dict[str, Any]] = [] messages: list[dict[str, Any]] = []
for row in reversed(rows): for row in reversed(rows):
sender_name = " ".join( sender_name = " ".join(
part for part in [row["first_name"], row["last_name"]] if part part for part in [row["first_name"], row["last_name"]] if part
@@ -505,7 +503,7 @@ def load_messages(
f"FROM messages WHERE message_id IN ({placeholders})", f"FROM messages WHERE message_id IN ({placeholders})",
reply_ids, reply_ids,
).fetchall() ).fetchall()
reply_map: Dict[int, Dict[str, str]] = {} reply_map: dict[int, dict[str, str]] = {}
for r in rows2: for r in rows2:
r_sender = " ".join( r_sender = " ".join(
part for part in [r["first_name"], r["last_name"]] if part part for part in [r["first_name"], r["last_name"]] if part
@@ -533,16 +531,16 @@ class Job:
job_id: str job_id: str
job_type: str job_type: str
title: str title: str
payload: Dict[str, Any] payload: dict[str, Any]
status: str = "queued" status: str = "queued"
created_at: str = field(default_factory=utc_now_iso) created_at: str = field(default_factory=utc_now_iso)
started_at: Optional[str] = None started_at: str | None = None
finished_at: Optional[str] = None finished_at: str | None = None
logs: str = "" logs: str = ""
error: Optional[str] = None error: str | None = None
account_id: Optional[str] = None account_id: str | None = None
def to_dict(self) -> Dict[str, Any]: def to_dict(self) -> dict[str, Any]:
return { return {
"job_id": self.job_id, "job_id": self.job_id,
"job_type": self.job_type, "job_type": self.job_type,
@@ -562,14 +560,14 @@ class Job:
class JobRunner: class JobRunner:
def __init__(self) -> None: def __init__(self) -> None:
self.jobs: Dict[str, Job] = {} self.jobs: dict[str, Job] = {}
self.job_order: List[str] = [] self.job_order: list[str] = []
self.queue: "queue.Queue[Job]" = queue.Queue() self.queue: queue.Queue[Job] = queue.Queue()
self.lock = threading.Lock() self.lock = threading.Lock()
self._shutdown_flag = False self._shutdown_flag = False
# Active SSE event streams: stream_id -> started_at, plus any stream # Active SSE event streams: stream_id -> started_at, plus any stream
# ids that have been revoked (told to close) because the cap was hit. # ids that have been revoked (told to close) because the cap was hit.
self._stream_started: Dict[str, float] = {} self._stream_started: dict[str, float] = {}
self._stream_revoked: set = set() self._stream_revoked: set = set()
self.worker = threading.Thread(target=self._run, daemon=True) self.worker = threading.Thread(target=self._run, daemon=True)
self.worker.start() self.worker.start()
@@ -594,7 +592,7 @@ class JobRunner:
with self.lock: with self.lock:
return stream_id in self._stream_revoked return stream_id in self._stream_revoked
def create_job(self, job_type: str, title: str, payload: Dict[str, Any]) -> Job: def create_job(self, job_type: str, title: str, payload: dict[str, Any]) -> Job:
if self._shutdown_flag: if self._shutdown_flag:
raise RuntimeError("Server is shutting down, cannot create new jobs") raise RuntimeError("Server is shutting down, cannot create new jobs")
account_id = payload.get("account_id") account_id = payload.get("account_id")
@@ -626,14 +624,14 @@ class JobRunner:
self.queue.put(job) self.queue.put(job)
return job return job
def recent_jobs(self, account_id: Optional[str] = None) -> List[Dict[str, Any]]: def recent_jobs(self, account_id: str | None = None) -> list[dict[str, Any]]:
with self.lock: with self.lock:
result = [self.jobs[job_id].to_dict() for job_id in self.job_order] result = [self.jobs[job_id].to_dict() for job_id in self.job_order]
if account_id is not None: if account_id is not None:
result = [j for j in result if j.get("account_id") == account_id] result = [j for j in result if j.get("account_id") == account_id]
return result return result
def active_jobs_by_account(self) -> Dict[str, Dict[str, Any]]: def active_jobs_by_account(self) -> dict[str, dict[str, Any]]:
with self.lock: with self.lock:
active = {} active = {}
for job_id in self.job_order: for job_id in self.job_order:
@@ -642,7 +640,7 @@ class JobRunner:
active[job.account_id] = job.to_dict() active[job.account_id] = job.to_dict()
return active return active
def get_job(self, job_id: str) -> Optional[Dict[str, Any]]: def get_job(self, job_id: str) -> dict[str, Any] | None:
with self.lock: with self.lock:
job = self.jobs.get(job_id) job = self.jobs.get(job_id)
return job.to_dict() if job else None return job.to_dict() if job else None
@@ -731,8 +729,8 @@ class TelegramAuthManager:
self.loop = asyncio.new_event_loop() self.loop = asyncio.new_event_loop()
self.thread = threading.Thread(target=self._run_loop, daemon=True) self.thread = threading.Thread(target=self._run_loop, daemon=True)
self.thread.start() self.thread.start()
self.clients: Dict[str, TelegramClient] = {} self.clients: dict[str, TelegramClient] = {}
self.auth_data: Dict[str, Dict[str, Any]] = {} self.auth_data: dict[str, dict[str, Any]] = {}
def _run_loop(self) -> None: def _run_loop(self) -> None:
asyncio.set_event_loop(self.loop) asyncio.set_event_loop(self.loop)
@@ -742,7 +740,7 @@ class TelegramAuthManager:
future = asyncio.run_coroutine_threadsafe(coro, self.loop) future = asyncio.run_coroutine_threadsafe(coro, self.loop)
return future.result() return future.result()
def _get_auth_data(self, account_id: str) -> Dict[str, Any]: def _get_auth_data(self, account_id: str) -> dict[str, Any]:
with self.lock: with self.lock:
if account_id not in self.auth_data: if account_id not in self.auth_data:
self.auth_data[account_id] = { self.auth_data[account_id] = {
@@ -808,7 +806,7 @@ class TelegramAuthManager:
data["user_id"] = me.id data["user_id"] = me.id
return client return client
def auth_state(self, account_id: str) -> Dict[str, Any]: def auth_state(self, account_id: str) -> dict[str, Any]:
with self.lock: with self.lock:
data = self._get_auth_data(account_id) data = self._get_auth_data(account_id)
snapshot = dict(data) snapshot = dict(data)
@@ -833,14 +831,14 @@ class TelegramAuthManager:
snapshot["auth_status"] = auth_status_for(account_id) snapshot["auth_status"] = auth_status_for(account_id)
return snapshot return snapshot
def all_auth_states(self) -> Dict[str, Dict[str, Any]]: def all_auth_states(self) -> dict[str, dict[str, Any]]:
result: Dict[str, Dict[str, Any]] = {} result: dict[str, dict[str, Any]] = {}
for account_id in list_accounts(DATA_DIR): for account_id in list_accounts(DATA_DIR):
result[account_id] = self.auth_state(account_id) result[account_id] = self.auth_state(account_id)
return result return result
def save_credentials(self, account_id: str, api_id: int, api_hash: str) -> Dict[str, Any]: def save_credentials(self, account_id: str, api_id: int, api_hash: str) -> dict[str, Any]:
def mutate(state: Dict[str, Any]) -> None: def mutate(state: dict[str, Any]) -> None:
state["api_id"] = int(api_id) state["api_id"] = int(api_id)
state["api_hash"] = api_hash.strip() state["api_hash"] = api_hash.strip()
@@ -854,7 +852,7 @@ class TelegramAuthManager:
) )
return self.auth_state(account_id) return self.auth_state(account_id)
async def _start_qr_login(self, account_id: str) -> Dict[str, Any]: async def _start_qr_login(self, account_id: str) -> dict[str, Any]:
client = await self._get_client(account_id) client = await self._get_client(account_id)
if await client.is_user_authorized(): if await client.is_user_authorized():
self._set_state( self._set_state(
@@ -923,10 +921,10 @@ class TelegramAuthManager:
qr_created_at=None, qr_created_at=None,
) )
def start_qr_login(self, account_id: str) -> Dict[str, Any]: def start_qr_login(self, account_id: str) -> dict[str, Any]:
return self._run(self._start_qr_login(account_id)) return self._run(self._start_qr_login(account_id))
async def _request_phone_code(self, account_id: str, phone: str) -> Dict[str, Any]: async def _request_phone_code(self, account_id: str, phone: str) -> dict[str, Any]:
client = await self._get_client(account_id) client = await self._get_client(account_id)
if await client.is_user_authorized(): if await client.is_user_authorized():
self._set_state( self._set_state(
@@ -947,10 +945,10 @@ class TelegramAuthManager:
) )
return self.auth_state(account_id) return self.auth_state(account_id)
def request_phone_code(self, account_id: str, phone: str) -> Dict[str, Any]: def request_phone_code(self, account_id: str, phone: str) -> dict[str, Any]:
return self._run(self._request_phone_code(account_id, phone)) return self._run(self._request_phone_code(account_id, phone))
async def _submit_phone_code(self, account_id: str, code: str) -> Dict[str, Any]: async def _submit_phone_code(self, account_id: str, code: str) -> dict[str, Any]:
client = await self._get_client(account_id) client = await self._get_client(account_id)
with self.lock: with self.lock:
data = self._get_auth_data(account_id) data = self._get_auth_data(account_id)
@@ -982,10 +980,10 @@ class TelegramAuthManager:
) )
return self.auth_state(account_id) return self.auth_state(account_id)
def submit_phone_code(self, account_id: str, code: str) -> Dict[str, Any]: def submit_phone_code(self, account_id: str, code: str) -> dict[str, Any]:
return self._run(self._submit_phone_code(account_id, code)) return self._run(self._submit_phone_code(account_id, code))
async def _submit_password(self, account_id: str, password: str) -> Dict[str, Any]: async def _submit_password(self, account_id: str, password: str) -> dict[str, Any]:
client = await self._get_client(account_id) client = await self._get_client(account_id)
await client.sign_in(password=password) await client.sign_in(password=password)
self._set_state( self._set_state(
@@ -999,7 +997,7 @@ class TelegramAuthManager:
) )
return self.auth_state(account_id) return self.auth_state(account_id)
def submit_password(self, account_id: str, password: str) -> Dict[str, Any]: def submit_password(self, account_id: str, password: str) -> dict[str, Any]:
return self._run(self._submit_password(account_id, password)) return self._run(self._submit_password(account_id, password))
def delete_account(self, account_id: str) -> None: def delete_account(self, account_id: str) -> None:
@@ -1049,14 +1047,14 @@ class PerAccountContinuousScrapeManager:
def __init__(self, account_id: str) -> None: def __init__(self, account_id: str) -> None:
self.account_id = account_id self.account_id = account_id
self.lock = threading.RLock() self.lock = threading.RLock()
self.thread: Optional[threading.Thread] = None self.thread: threading.Thread | None = None
self.stop_event = threading.Event() self.stop_event = threading.Event()
# F-4: set on remove_account when its thread could not be joined in # F-4: set on remove_account when its thread could not be joined in
# time; the manager is kept as a tombstone so it is reused (after the # time; the manager is kept as a tombstone so it is reused (after the
# drain) instead of spawning a duplicate worker on account re-add. # drain) instead of spawning a duplicate worker on account re-add.
self._removing = False self._removing = False
self.config: Dict[str, Any] = self._load_config() self.config: dict[str, Any] = self._load_config()
self.status: Dict[str, Any] = { self.status: dict[str, Any] = {
"running": False, "running": False,
"last_started_at": None, "last_started_at": None,
"last_finished_at": None, "last_finished_at": None,
@@ -1066,7 +1064,7 @@ class PerAccountContinuousScrapeManager:
"log_entries": [], "log_entries": [],
} }
def _load_config(self) -> Dict[str, Any]: def _load_config(self) -> dict[str, Any]:
acc_state = load_account(DATA_DIR, self.account_id) acc_state = load_account(DATA_DIR, self.account_id)
cfg = dict(acc_state.get("continuous_scraping", { cfg = dict(acc_state.get("continuous_scraping", {
"enabled": True, "enabled": True,
@@ -1085,7 +1083,7 @@ class PerAccountContinuousScrapeManager:
def _save_config(self) -> None: def _save_config(self) -> None:
store = get_account_store(DATA_DIR, self.account_id) store = get_account_store(DATA_DIR, self.account_id)
def mutate(state: Dict[str, Any]) -> None: def mutate(state: dict[str, Any]) -> None:
state["continuous_scraping"] = { state["continuous_scraping"] = {
"enabled": bool(self.config.get("enabled", True)), "enabled": bool(self.config.get("enabled", True)),
"interval_minutes": max(1, int(self.config.get("interval_minutes", 1) or 1)), "interval_minutes": max(1, int(self.config.get("interval_minutes", 1) or 1)),
@@ -1140,7 +1138,7 @@ class PerAccountContinuousScrapeManager:
if len(log_entries) > 300: if len(log_entries) > 300:
del log_entries[:-300] del log_entries[:-300]
def snapshot(self) -> Dict[str, Any]: def snapshot(self) -> dict[str, Any]:
with self.lock: with self.lock:
return { return {
"config": dict(self.config), "config": dict(self.config),
@@ -1155,9 +1153,9 @@ class PerAccountContinuousScrapeManager:
self, self,
enabled: bool, enabled: bool,
interval_minutes: int, interval_minutes: int,
channels: List[str], channels: list[str],
run_all_tracked: bool, run_all_tracked: bool,
) -> Dict[str, Any]: ) -> dict[str, Any]:
interval_minutes = max(1, int(interval_minutes)) interval_minutes = max(1, int(interval_minutes))
with self.lock: with self.lock:
self.config = { self.config = {
@@ -1242,7 +1240,7 @@ class PerAccountContinuousScrapeManager:
thread.join(timeout=timeout) thread.join(timeout=timeout)
return not thread.is_alive() return not thread.is_alive()
def _resolve_channels(self) -> List[str]: def _resolve_channels(self) -> list[str]:
acc_state = load_account(DATA_DIR, self.account_id) acc_state = load_account(DATA_DIR, self.account_id)
with self.lock: with self.lock:
run_all_tracked = self.config.get("run_all_tracked", True) run_all_tracked = self.config.get("run_all_tracked", True)
@@ -1344,7 +1342,7 @@ class PerAccountContinuousScrapeManager:
class ContinuousScrapeOrchestrator: class ContinuousScrapeOrchestrator:
def __init__(self) -> None: def __init__(self) -> None:
self.lock = threading.Lock() self.lock = threading.Lock()
self.managers: Dict[str, PerAccountContinuousScrapeManager] = {} self.managers: dict[str, PerAccountContinuousScrapeManager] = {}
def _get_or_create(self, account_id: str) -> PerAccountContinuousScrapeManager: def _get_or_create(self, account_id: str) -> PerAccountContinuousScrapeManager:
with self.lock: with self.lock:
@@ -1371,12 +1369,12 @@ class ContinuousScrapeOrchestrator:
if mgr: if mgr:
mgr.stop() mgr.stop()
def snapshot_for(self, account_id: str) -> Dict[str, Any]: def snapshot_for(self, account_id: str) -> dict[str, Any]:
mgr = self._get_or_create(account_id) mgr = self._get_or_create(account_id)
return mgr.snapshot() return mgr.snapshot()
def snapshot(self) -> Dict[str, Dict[str, Any]]: def snapshot(self) -> dict[str, dict[str, Any]]:
result: Dict[str, Dict[str, Any]] = {} result: dict[str, dict[str, Any]] = {}
with self.lock: with self.lock:
for account_id, mgr in self.managers.items(): for account_id, mgr in self.managers.items():
result[account_id] = dict(mgr.snapshot()) result[account_id] = dict(mgr.snapshot())
@@ -1396,9 +1394,9 @@ class ContinuousScrapeOrchestrator:
account_id: str, account_id: str,
enabled: bool, enabled: bool,
interval_minutes: int, interval_minutes: int,
channels: List[str], channels: list[str],
run_all_tracked: bool, run_all_tracked: bool,
) -> Dict[str, Any]: ) -> dict[str, Any]:
mgr = self._get_or_create(account_id) mgr = self._get_or_create(account_id)
result = mgr.update(enabled, interval_minutes, channels, run_all_tracked) result = mgr.update(enabled, interval_minutes, channels, run_all_tracked)
# After explicit update, sync state to disk is already done by mgr.update # After explicit update, sync state to disk is already done by mgr.update
@@ -1456,18 +1454,18 @@ def import_scraper_class():
return OptimizedTelegramScraper return OptimizedTelegramScraper
def run_job(job_type: str, payload: Dict[str, Any]) -> None: def run_job(job_type: str, payload: dict[str, Any]) -> None:
SCRAPER_JOBS.run(job_type, payload) SCRAPER_JOBS.run(job_type, payload)
def auth_status_for(account_id: Optional[str] = None) -> Dict[str, Any]: def auth_status_for(account_id: str | None = None) -> dict[str, Any]:
if account_id: if account_id:
acc_state = load_account(DATA_DIR, account_id) acc_state = load_account(DATA_DIR, account_id)
session_file = account_session_path(SESSION_DIR, account_id) session_file = account_session_path(SESSION_DIR, account_id)
else: else:
acc_state = load_state() acc_state = load_state()
session_file = str(SESSION_DIR / "session.session") session_file = str(SESSION_DIR / "session.session")
status: Dict[str, Any] = { status: dict[str, Any] = {
"has_api_credentials": bool(acc_state.get("api_id") and acc_state.get("api_hash")), "has_api_credentials": bool(acc_state.get("api_id") and acc_state.get("api_hash")),
"telethon_available": False, "telethon_available": False,
"session_ready": False, "session_ready": False,
@@ -1494,7 +1492,7 @@ def auth_status_for(account_id: Optional[str] = None) -> Dict[str, Any]:
return status return status
def account_health_summary(account_id: str, job_runner: Optional[JobRunner] = None) -> Dict[str, Any]: def account_health_summary(account_id: str, job_runner: JobRunner | None = None) -> dict[str, Any]:
acc_state = load_account(DATA_DIR, account_id) acc_state = load_account(DATA_DIR, account_id)
acc_dir = account_data_dir(DATA_DIR, account_id) acc_dir = account_data_dir(DATA_DIR, account_id)
session_file = Path(account_session_path(SESSION_DIR, account_id)) session_file = Path(account_session_path(SESSION_DIR, account_id))
@@ -1515,11 +1513,11 @@ def account_health_summary(account_id: str, job_runner: Optional[JobRunner] = No
} }
def auth_status() -> Dict[str, Any]: def auth_status() -> dict[str, Any]:
return auth_status_for(account_id=None) return auth_status_for(account_id=None)
def dashboard_payload(job_runner: JobRunner, account_id: Optional[str] = None) -> Dict[str, Any]: def dashboard_payload(job_runner: JobRunner, account_id: str | None = None) -> dict[str, Any]:
if account_id: if account_id:
acc_state = load_account(DATA_DIR, account_id) acc_state = load_account(DATA_DIR, account_id)
state = { state = {
@@ -1547,7 +1545,7 @@ def dashboard_payload(job_runner: JobRunner, account_id: Optional[str] = None) -
} }
def openapi_payload() -> Dict[str, Any]: def openapi_payload() -> dict[str, Any]:
json_response = { json_response = {
"200": { "200": {
"description": "JSON response", "description": "JSON response",
@@ -1562,7 +1560,7 @@ def openapi_payload() -> Dict[str, Any]:
} }
error_response = {"400": {"description": "Bad request"}} error_response = {"400": {"description": "Bad request"}}
def json_body(properties: Dict[str, Any], required: Optional[List[str]] = None): def json_body(properties: dict[str, Any], required: list[str] | None = None):
return { return {
"required": True, "required": True,
"content": { "content": {
@@ -2335,7 +2333,7 @@ class TelegramScraperRequestHandler(BaseHTTPRequestHandler):
return self.send_error_json(HTTPStatus.NOT_FOUND, "Not found") return self.send_error_json(HTTPStatus.NOT_FOUND, "Not found")
def _legacy_auth_snapshot(self) -> Dict[str, Any]: def _legacy_auth_snapshot(self) -> dict[str, Any]:
state = load_state() state = load_state()
return { return {
"phase": "unknown", "phase": "unknown",
@@ -2369,7 +2367,7 @@ class TelegramScraperRequestHandler(BaseHTTPRequestHandler):
}) })
return self.send_json({"accounts": accounts}) return self.send_json({"accounts": accounts})
def _handle_get_account(self, path: str, query: Dict[str, List[str]]) -> None: def _handle_get_account(self, path: str, query: dict[str, list[str]]) -> None:
rest = path[len("/api/accounts/"):] rest = path[len("/api/accounts/"):]
parts = rest.split("/") parts = rest.split("/")
account_id = urllib.parse.unquote(parts[0]) account_id = urllib.parse.unquote(parts[0])
@@ -2446,7 +2444,7 @@ class TelegramScraperRequestHandler(BaseHTTPRequestHandler):
payload = json.dumps(job, ensure_ascii=False) payload = json.dumps(job, ensure_ascii=False)
if payload != last_payload: if payload != last_payload:
try: try:
self.wfile.write(f"data: {payload}\n\n".encode("utf-8")) self.wfile.write(f"data: {payload}\n\n".encode())
self.wfile.flush() self.wfile.flush()
except (BrokenPipeError, ConnectionResetError, OSError): except (BrokenPipeError, ConnectionResetError, OSError):
break break
@@ -2461,7 +2459,7 @@ class TelegramScraperRequestHandler(BaseHTTPRequestHandler):
return self.send_json(list_channels_snapshot(account_id)) return self.send_json(list_channels_snapshot(account_id))
def _handle_get_account_channel_messages( def _handle_get_account_channel_messages(
self, account_id: str, channel_id: str, query: Dict[str, List[str]] self, account_id: str, channel_id: str, query: dict[str, list[str]]
) -> None: ) -> None:
try: try:
limit = max(1, min(int(query.get("limit", ["120"])[0]), 300)) limit = max(1, min(int(query.get("limit", ["120"])[0]), 300))
@@ -2511,21 +2509,7 @@ class TelegramScraperRequestHandler(BaseHTTPRequestHandler):
relative = path.removeprefix("/media/") relative = path.removeprefix("/media/")
return self.serve_media(relative, head_only=True) return self.serve_media(relative, head_only=True)
if ( if (
path path in {"/api/dashboard", "/api/auth", "/api/continuous", "/api/jobs", "/api/channels", "/api/accounts", "/health", "/health/continuous", "/openapi.json"} or path.startswith(("/api/jobs/", "/api/accounts/")) or path.startswith("/api/channels/") and path.endswith("/messages")
in {
"/api/dashboard",
"/api/auth",
"/api/continuous",
"/api/jobs",
"/api/channels",
"/api/accounts",
"/health",
"/health/continuous",
"/openapi.json",
}
or path.startswith("/api/jobs/")
or (path.startswith("/api/channels/") and path.endswith("/messages"))
or path.startswith("/api/accounts/")
): ):
# M-4: for the job-events route, mirror GET's behavior and verify # M-4: for the job-events route, mirror GET's behavior and verify
# the job actually exists before responding with 200. # the job actually exists before responding with 200.
@@ -2562,7 +2546,7 @@ class TelegramScraperRequestHandler(BaseHTTPRequestHandler):
# No legacy account: persist to the (global) legacy store directly. # No legacy account: persist to the (global) legacy store directly.
# The ScraperJobService set_scrape_media job is a passthrough, so # The ScraperJobService set_scrape_media job is a passthrough, so
# enqueueing it alone would silently drop the setting. # enqueueing it alone would silently drop the setting.
def _media_mutate(global_state: Dict[str, Any]) -> None: def _media_mutate(global_state: dict[str, Any]) -> None:
global_state["scrape_media"] = value global_state["scrape_media"] = value
STATE_STORE.update(_media_mutate) STATE_STORE.update(_media_mutate)
return self.send_json({"ok": True, "scrape_media": value}) return self.send_json({"ok": True, "scrape_media": value})
@@ -2655,7 +2639,7 @@ class TelegramScraperRequestHandler(BaseHTTPRequestHandler):
if self.app.legacy_account_id: if self.app.legacy_account_id:
return self._handle_account_channel_add(self.app.legacy_account_id, channel_id, body.get("name")) return self._handle_account_channel_add(self.app.legacy_account_id, channel_id, body.get("name"))
# L-6: use the atomic mutator to avoid lost-update races. # L-6: use the atomic mutator to avoid lost-update races.
def _add_mutate(state: Dict[str, Any]) -> None: def _add_mutate(state: dict[str, Any]) -> None:
if channel_id not in state.get("channels", {}): if channel_id not in state.get("channels", {}):
state.setdefault("channels", {})[channel_id] = 0 state.setdefault("channels", {})[channel_id] = 0
if body.get("name"): if body.get("name"):
@@ -2672,7 +2656,7 @@ class TelegramScraperRequestHandler(BaseHTTPRequestHandler):
return self._handle_account_channel_remove(self.app.legacy_account_id, channel_id) return self._handle_account_channel_remove(self.app.legacy_account_id, channel_id)
# L-6: use the atomic mutator to avoid lost-update races. # L-6: use the atomic mutator to avoid lost-update races.
existed = [False] existed = [False]
def _remove_mutate(state: Dict[str, Any]) -> None: def _remove_mutate(state: dict[str, Any]) -> None:
chans = state.get("channels", {}) chans = state.get("channels", {})
if channel_id in chans: if channel_id in chans:
existed[0] = True existed[0] = True
@@ -2782,7 +2766,7 @@ class TelegramScraperRequestHandler(BaseHTTPRequestHandler):
return self.send_error_json(HTTPStatus.NOT_FOUND, "Not found") return self.send_error_json(HTTPStatus.NOT_FOUND, "Not found")
def _handle_post_accounts_create(self, body: Dict[str, Any]) -> None: def _handle_post_accounts_create(self, body: dict[str, Any]) -> None:
account_id = str(body.get("account_id", "")).strip() account_id = str(body.get("account_id", "")).strip()
if not account_id: if not account_id:
return self.send_error_json(HTTPStatus.BAD_REQUEST, "account_id is required") return self.send_error_json(HTTPStatus.BAD_REQUEST, "account_id is required")
@@ -2797,7 +2781,7 @@ class TelegramScraperRequestHandler(BaseHTTPRequestHandler):
api_id = body.get("api_id") api_id = body.get("api_id")
api_hash = str(body.get("api_hash", "")).strip() api_hash = str(body.get("api_hash", "")).strip()
def mutate_global(state: Dict[str, Any]) -> None: def mutate_global(state: dict[str, Any]) -> None:
accounts = state.setdefault("accounts", []) accounts = state.setdefault("accounts", [])
if account_id not in accounts: if account_id not in accounts:
accounts.append(account_id) accounts.append(account_id)
@@ -2824,7 +2808,7 @@ class TelegramScraperRequestHandler(BaseHTTPRequestHandler):
self.app.continuous_orchestrator.add_account(account_id, auto_start=False) self.app.continuous_orchestrator.add_account(account_id, auto_start=False)
return self.send_json({"ok": True, "account_id": account_id}) return self.send_json({"ok": True, "account_id": account_id})
def _handle_post_accounts_import(self, body: Dict[str, Any]) -> None: def _handle_post_accounts_import(self, body: dict[str, Any]) -> None:
account_id = str(body.get("account_id") or body.get("id") or "").strip() account_id = str(body.get("account_id") or body.get("id") or "").strip()
state = body.get("state") state = body.get("state")
if not account_id: if not account_id:
@@ -2863,7 +2847,7 @@ class TelegramScraperRequestHandler(BaseHTTPRequestHandler):
}, },
} }
def mutate_global(global_state: Dict[str, Any]) -> None: def mutate_global(global_state: dict[str, Any]) -> None:
accounts = global_state.setdefault("accounts", []) accounts = global_state.setdefault("accounts", [])
if account_id not in accounts: if account_id not in accounts:
accounts.append(account_id) accounts.append(account_id)
@@ -2874,7 +2858,7 @@ class TelegramScraperRequestHandler(BaseHTTPRequestHandler):
self.app.continuous_orchestrator.refresh_account(account_id) self.app.continuous_orchestrator.refresh_account(account_id)
return self.send_json({"ok": True, "account_id": account_id}) return self.send_json({"ok": True, "account_id": account_id})
def _handle_post_account(self, path: str, body: Dict[str, Any]) -> None: def _handle_post_account(self, path: str, body: dict[str, Any]) -> None:
rest = path[len("/api/accounts/"):] rest = path[len("/api/accounts/"):]
parts = rest.split("/") parts = rest.split("/")
account_id = urllib.parse.unquote(parts[0]) account_id = urllib.parse.unquote(parts[0])
@@ -2922,7 +2906,7 @@ class TelegramScraperRequestHandler(BaseHTTPRequestHandler):
return False return False
return True return True
def _handle_post_account_auth_credentials(self, account_id: str, body: Dict[str, Any]) -> None: def _handle_post_account_auth_credentials(self, account_id: str, body: dict[str, Any]) -> None:
api_id = body.get("api_id") api_id = body.get("api_id")
api_hash = str(body.get("api_hash", "")).strip() api_hash = str(body.get("api_hash", "")).strip()
if not api_id or not api_hash: if not api_id or not api_hash:
@@ -2940,7 +2924,7 @@ class TelegramScraperRequestHandler(BaseHTTPRequestHandler):
return self.send_bad_request_from_exc(exc, fallback="Failed to start QR login") return self.send_bad_request_from_exc(exc, fallback="Failed to start QR login")
return self.send_json(payload) return self.send_json(payload)
def _handle_post_account_auth_phone_request(self, account_id: str, body: Dict[str, Any]) -> None: def _handle_post_account_auth_phone_request(self, account_id: str, body: dict[str, Any]) -> None:
ip = self.client_address[0] if self.client_address else "" ip = self.client_address[0] if self.client_address else ""
if not _check_auth_throttle(ip, account_id): if not _check_auth_throttle(ip, account_id):
return self.send_error_json( return self.send_error_json(
@@ -2963,7 +2947,7 @@ class TelegramScraperRequestHandler(BaseHTTPRequestHandler):
_record_auth_code_request(ip, account_id) _record_auth_code_request(ip, account_id)
return self.send_json(payload) return self.send_json(payload)
def _handle_post_account_auth_phone_submit(self, account_id: str, body: Dict[str, Any]) -> None: def _handle_post_account_auth_phone_submit(self, account_id: str, body: dict[str, Any]) -> None:
ip = self.client_address[0] if self.client_address else "" ip = self.client_address[0] if self.client_address else ""
if not _check_auth_throttle(ip, account_id): if not _check_auth_throttle(ip, account_id):
return self.send_error_json( return self.send_error_json(
@@ -2981,7 +2965,7 @@ class TelegramScraperRequestHandler(BaseHTTPRequestHandler):
_record_auth_success(ip, account_id) _record_auth_success(ip, account_id)
return self.send_json(payload) return self.send_json(payload)
def _handle_post_account_auth_password(self, account_id: str, body: Dict[str, Any]) -> None: def _handle_post_account_auth_password(self, account_id: str, body: dict[str, Any]) -> None:
ip = self.client_address[0] if self.client_address else "" ip = self.client_address[0] if self.client_address else ""
if not _check_auth_throttle(ip, account_id): if not _check_auth_throttle(ip, account_id):
return self.send_error_json( return self.send_error_json(
@@ -2999,10 +2983,10 @@ class TelegramScraperRequestHandler(BaseHTTPRequestHandler):
_record_auth_success(ip, account_id) _record_auth_success(ip, account_id)
return self.send_json(payload) return self.send_json(payload)
def _clean_imported_channels(self, channels: Any) -> Dict[str, Any]: def _clean_imported_channels(self, channels: Any) -> dict[str, Any]:
if not isinstance(channels, dict): if not isinstance(channels, dict):
return {} return {}
cleaned: Dict[str, Any] = {} cleaned: dict[str, Any] = {}
for raw_channel_id, last_message_id in channels.items(): for raw_channel_id, last_message_id in channels.items():
try: try:
channel_id = normalize_channel_id(raw_channel_id) channel_id = normalize_channel_id(raw_channel_id)
@@ -3011,10 +2995,10 @@ class TelegramScraperRequestHandler(BaseHTTPRequestHandler):
cleaned[channel_id] = last_message_id cleaned[channel_id] = last_message_id
return cleaned return cleaned
def _clean_imported_channel_names(self, channel_names: Any) -> Dict[str, str]: def _clean_imported_channel_names(self, channel_names: Any) -> dict[str, str]:
if not isinstance(channel_names, dict): if not isinstance(channel_names, dict):
return {} return {}
cleaned: Dict[str, str] = {} cleaned: dict[str, str] = {}
for raw_channel_id, name in channel_names.items(): for raw_channel_id, name in channel_names.items():
try: try:
channel_id = normalize_channel_id(raw_channel_id) channel_id = normalize_channel_id(raw_channel_id)
@@ -3023,13 +3007,13 @@ class TelegramScraperRequestHandler(BaseHTTPRequestHandler):
cleaned[channel_id] = str(name).strip() cleaned[channel_id] = str(name).strip()
return cleaned return cleaned
def _handle_account_channel_add(self, account_id: str, channel_id: Any, name: Optional[str]) -> None: def _handle_account_channel_add(self, account_id: str, channel_id: Any, name: str | None) -> None:
try: try:
channel_id = normalize_channel_id(channel_id) channel_id = normalize_channel_id(channel_id)
except ValueError as exc: except ValueError as exc:
return self.send_error_json(HTTPStatus.BAD_REQUEST, str(exc)) return self.send_error_json(HTTPStatus.BAD_REQUEST, str(exc))
store = get_account_store(DATA_DIR, account_id) store = get_account_store(DATA_DIR, account_id)
def mutate(state: Dict[str, Any]) -> None: def mutate(state: dict[str, Any]) -> None:
if channel_id not in state.setdefault("channels", {}): if channel_id not in state.setdefault("channels", {}):
state["channels"][channel_id] = 0 state["channels"][channel_id] = 0
if name: if name:
@@ -3045,7 +3029,7 @@ class TelegramScraperRequestHandler(BaseHTTPRequestHandler):
return self.send_error_json(HTTPStatus.BAD_REQUEST, str(exc)) return self.send_error_json(HTTPStatus.BAD_REQUEST, str(exc))
store = get_account_store(DATA_DIR, account_id) store = get_account_store(DATA_DIR, account_id)
existed = [False] existed = [False]
def mutate(state: Dict[str, Any]) -> None: def mutate(state: dict[str, Any]) -> None:
chans = state.get("channels", {}) chans = state.get("channels", {})
if channel_id in chans: if channel_id in chans:
existed[0] = True existed[0] = True
@@ -3055,7 +3039,7 @@ class TelegramScraperRequestHandler(BaseHTTPRequestHandler):
self.app.continuous_orchestrator.refresh_account(account_id) self.app.continuous_orchestrator.refresh_account(account_id)
return self.send_json({"ok": existed[0], "channel_id": channel_id}) return self.send_json({"ok": existed[0], "channel_id": channel_id})
def _handle_account_job_scrape(self, account_id: str, body: Dict[str, Any]) -> None: def _handle_account_job_scrape(self, account_id: str, body: dict[str, Any]) -> None:
channel_id = body.get("channel_id") channel_id = body.get("channel_id")
if channel_id: if channel_id:
try: try:
@@ -3083,7 +3067,7 @@ class TelegramScraperRequestHandler(BaseHTTPRequestHandler):
) )
return self.send_json(job.to_dict(), status=HTTPStatus.ACCEPTED) return self.send_json(job.to_dict(), status=HTTPStatus.ACCEPTED)
def _handle_account_job_export_channel(self, account_id: str, body: Dict[str, Any]) -> None: def _handle_account_job_export_channel(self, account_id: str, body: dict[str, Any]) -> None:
try: try:
channel_id = normalize_channel_id(body.get("channel_id")) channel_id = normalize_channel_id(body.get("channel_id"))
except ValueError as exc: except ValueError as exc:
@@ -3095,7 +3079,7 @@ class TelegramScraperRequestHandler(BaseHTTPRequestHandler):
) )
return self.send_json(job.to_dict(), status=HTTPStatus.ACCEPTED) return self.send_json(job.to_dict(), status=HTTPStatus.ACCEPTED)
def _handle_account_job_rescrape_media(self, account_id: str, body: Dict[str, Any]) -> None: def _handle_account_job_rescrape_media(self, account_id: str, body: dict[str, Any]) -> None:
try: try:
channel_id = normalize_channel_id(body.get("channel_id")) channel_id = normalize_channel_id(body.get("channel_id"))
except ValueError as exc: except ValueError as exc:
@@ -3107,7 +3091,7 @@ class TelegramScraperRequestHandler(BaseHTTPRequestHandler):
) )
return self.send_json(job.to_dict(), status=HTTPStatus.ACCEPTED) return self.send_json(job.to_dict(), status=HTTPStatus.ACCEPTED)
def _handle_account_job_fix_media(self, account_id: str, body: Dict[str, Any]) -> None: def _handle_account_job_fix_media(self, account_id: str, body: dict[str, Any]) -> None:
try: try:
channel_id = normalize_channel_id(body.get("channel_id")) channel_id = normalize_channel_id(body.get("channel_id"))
except ValueError as exc: except ValueError as exc:
@@ -3129,7 +3113,7 @@ class TelegramScraperRequestHandler(BaseHTTPRequestHandler):
def _handle_account_settings_media(self, account_id: str, value: bool) -> None: def _handle_account_settings_media(self, account_id: str, value: bool) -> None:
store = get_account_store(DATA_DIR, account_id) store = get_account_store(DATA_DIR, account_id)
def mutate(state: Dict[str, Any]) -> None: def mutate(state: dict[str, Any]) -> None:
state["scrape_media"] = value state["scrape_media"] = value
store.update(mutate) store.update(mutate)
job = self.app.job_runner.create_job( job = self.app.job_runner.create_job(
@@ -3139,7 +3123,7 @@ class TelegramScraperRequestHandler(BaseHTTPRequestHandler):
) )
return self.send_json(job.to_dict(), status=HTTPStatus.ACCEPTED) return self.send_json(job.to_dict(), status=HTTPStatus.ACCEPTED)
def _handle_account_continuous(self, account_id: str, body: Dict[str, Any]) -> None: def _handle_account_continuous(self, account_id: str, body: dict[str, Any]) -> None:
try: try:
enabled = parse_bool(body.get("enabled")) enabled = parse_bool(body.get("enabled"))
interval_minutes = int(body.get("interval_minutes", 1)) interval_minutes = int(body.get("interval_minutes", 1))
@@ -3199,7 +3183,7 @@ class TelegramScraperRequestHandler(BaseHTTPRequestHandler):
self.app.continuous_orchestrator.remove_account(account_id) self.app.continuous_orchestrator.remove_account(account_id)
self.app.auth_manager.delete_account(account_id) self.app.auth_manager.delete_account(account_id)
def mutate_global(state: Dict[str, Any]) -> None: def mutate_global(state: dict[str, Any]) -> None:
accounts = state.setdefault("accounts", []) accounts = state.setdefault("accounts", [])
if account_id in accounts: if account_id in accounts:
accounts.remove(account_id) accounts.remove(account_id)
@@ -3456,7 +3440,7 @@ class TelegramScraperWebServer(ThreadingHTTPServer):
self.job_runner = JobRunner() self.job_runner = JobRunner()
self.auth_manager = TelegramAuthManager() self.auth_manager = TelegramAuthManager()
self.continuous_orchestrator = ContinuousScrapeOrchestrator() self.continuous_orchestrator = ContinuousScrapeOrchestrator()
self.legacy_account_id: Optional[str] = None self.legacy_account_id: str | None = None
self._detect_legacy_account() self._detect_legacy_account()
if START_CONTINUOUS: if START_CONTINUOUS: