Back to skill

Security audit

TG-Crawler — Telegram 舆情采集工具

Security checks for vulnerabilities and agentic risk

Overview

The Telegram crawler is mostly purpose-aligned, but it needs Review because it includes broad permissions, unsafe credential handling, third-party query disclosure, automatic private-group joining, and risky cleanup/proxy instructions.

Install only if you are comfortable using a dedicated Telegram account and carefully scoped credentials. Disable third-party bots for sensitive searches, avoid private invite-link targets unless authorized, do not pass 2FA passwords on the command line, review any purge media directory before confirming, and do not run the proxy setup as written without unique credentials and firewall restrictions.

Vulnerability Patterns
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
Findings (7)

T09 · Insecure Skill Coding Practices

Error
Location
references/proxy-pool-setup.md:72
Finding

Persistent Internet-Facing SOCKS5 Proxy Uses Publicly Documented Credentials

Content
View full analysis
/etc/danted.conf << 'DANTE_EOF' internal: eth0 port = 1080 external: eth0 socksmethod: username client pass { from: 0.0.0.0/0 to: 0.0.0.0/0 log: connect disconnect error } socks pass { from: 0.0.0.0/0 to: 0.0.0.0/0 command: bind connect udpassociate log: connect disconnect error } DANTE_EOF echo "=== Creating authentication user ===" useradd -r -s /bin/false dante_proxy 2>/dev/null || true echo "dante_proxy:ChangeThisPassword123" | chpasswd systemctl restart danted systemctl enable danted if systemctl is-active --quiet danted; then echo "Dante is running" else echo "Dante failed to start" fi echo "" echo "SOCKS5 proxy information:" echo " Address: $(curl -s ifconfig.me)" echo " Port: 1080" echo " User: dante_proxy" echo " Password: ChangeThisPassword123" ``` ### Technical Analysis The root-level deployment script exposes Dante on `eth0:1080`, accepts clients from `0.0.0.0/0`, permits arbitrary destinations, and installs a fixed password published directly in the repository. It then enables the service through `systemctl`, causing the proxy to survive reboots. Service persistence is functionally consistent with operating an always-on proxy and is not evidence of a hidden backdoor by itself. However, the persistent service is deployed before mandatory source restrictions are enforced. The later firewall guidance only recommends source-IP filtering and therefore does not safely constrain the configuration shown above. The use of a known password also means that network reachability effectively provides access to the proxy. Printing the credentials to the terminal can additionally expose them through terminal logs or operational records. ### At ...[truncated 923 chars]
Remediation
View remediation

other

Warning
Location
scripts/channel_discoverer.py:211
Finding

Search Keywords Are Disclosed to Third-Party Telegram Bots by Default

Content
View full analysis
list[DiscoveredChannel]: """ Strategy 2: invoke a search bot. Send keywords to the bot and parse t.me links from its response. """ channels = [] seen_links = set() for kw in keywords: try: await self.client.send_message(bot_username, kw) await asyncio.sleep(3) replies = await self.client.get_messages(bot_username, limit=5) ``` ### Technical Analysis The default discovery workflow sends every supplied search keyword as a Telegram message to the third-party accounts `xbso1` and `jisou`. These messages originate from the user's authenticated Telegram account. The transmitted data can include brands, investigation subjects, suspected criminal activity, or other sensitive research terms. The bot operators can associate those queries with the user's Telegram identity, account metadata, timestamps, and IP-related information available to Telegram. The project implements official Telegram search separately, and the CLI exposes a `--no-bots` option. Consequently, third-party disclosure is not strictly necessary for the crawler's core functionality and should not be the privacy-impacting default. ### Attack Path 1. A user starts `discover` or `hybrid` mode without `--no-bots`. 2. The crawler first performs official Telegram search. 3. It then sends each user-provided keyword to both configured search bots. 4. The bot operators receive and m ...[truncated 541 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
README.md:85
Finding

OTP and Telegram 2FA Secrets Are Handled Through Unsafe Files and Command-Line Arguments

Content
View full analysis
/tmp/tg_code.txt python3 main.py --mode discover --keywords "game cheats" --code-file /tmp/tg_code.txt ``` ```text If the account has two-factor authentication enabled, pass --password your_2FA_password. ``` ```python # 2FA password: command line takes precedence over environment tg_password = args.password or pwd_env or None ``` ```python def read_code(): if is_failover: import os as _os dedicated = f"code_acc{acct}.txt" if _os.path.exists(dedicated): with open(dedicated) as f: return f.read().strip() if args.code: return args.code raise RuntimeError(...) if args.code: return args.code if args.code_file: try: with open(args.code_file, 'r') as f: return f.read().strip() except Exception as e: logging.error(f"Failed to read verification-code file: {e}") raise kwargs = {"phone": _build_phone(acct), "code_callback": read_code} if tg_password: kwargs["password"] = tg_password await client.start(**kwargs) ``` ### Technical Analysis The documentation recommends creating a predictably named plaintext OTP file in `/tmp` using ordinary shell redirection. It does not set restrictive permissions, securely create the file, or delete it after use. The project also supports and documents passing the Telegram 2FA password through `--password`. Command-line arguments may be visible to other local users through process inspection and are commonly retained in shell history, terminal logs, automation logs, and diagnostic output. The implementation reads code files but does not val ...[truncated 1097 chars]
Remediation
View remediation

T08 · Insecure Dependencies

Warning
Location
scripts/requirements.txt:1
Finding

Dependencies Are Installed Without Exact Version or Integrity Pinning

Content
View full analysis
=1.36.0 pyyaml>=6.0 aiosqlite>=0.20.0 python-dotenv>=1.0.0 ``` The installation documentation instructs users to execute: ```bash pip install -r requirements.txt ``` ### Technical Analysis Every dependency uses an open-ended lower bound. A future installation can therefore resolve to versions that did not exist when the project was audited. The repository contains no lock file, package hashes, upper bounds, or reproducible installation metadata. Python package installation may execute package build hooks or other installation-time logic. If a dependency maintainer account or package index is compromised, or if a future release introduces unsafe behavior, users can receive changed executable code without any modification to this repository. No typosquatted package was identified in the current manifest; the issue is the absence of reproducible and integrity-checked dependency selection. ### Attack Path 1. An attacker compromises an upstream package, maintainer account, or distribution channel. 2. A malicious version is published with a version satisfying the listed `>=` constraint. 3. A user runs the documented `pip install -r requirements.txt` command. 4. The resolver selects the new malicious version. 5. Malicious installation or runtime code executes with the privileges of the user running `pip`. ### Impact Assessment A compromised dependency can obtain the same privileges as the crawler process. This may include reading Telegram API credentials, proxy credentials, session files, collected messages, and other files accessible to the current user, as well as making outbound network connections. The issue does not independently provide root access unless installation is performed with elevated privileges. ]]>
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/database.py:471
Finding

Attacker-Controlled Telegram Content Is Exported to CSV Without Formula Neutralization

Content
View full analysis
int: import csv with open(path, 'w', newline='', encoding='utf-8-sig') as f: writer = csv.writer(f) headers = ["id", "msg_id", "chat_id", "chat_title", "chat_username", "sender_id", "sender_username", "sender_name", "text", "media_type", "msg_date", "matched_keywords", "collected_at"] writer.writerow(headers) count = 0 for r in rows: writer.writerow([ r["id"], r["msg_id"], r["chat_id"], r["chat_title"], r["chat_username"], r["sender_id"], r["sender_username"], r["sender_name"], r["text"], r["media_type"], r["msg_date"], r["matched_keywords"], r["collected_at"] ]) count += 1 ``` ### Technical Analysis Telegram message text, channel titles, usernames, and sender names are attacker-controlled or externally controlled data. These values are written directly into CSV cells. The Python CSV writer correctly quotes CSV syntax but does not prevent spreadsheet applications from interpreting values beginning with `=`, `+`, `-`, or `@` as formulas. A malicious message can therefore become an active spreadsheet expression when an analyst opens the exported file. The exact result depends on the spreadsheet client and its security settings, but formula execution may trigger external network requests, disclose file-derived data, mislead the analyst, or invoke dangerous legacy functionality. ### Attack Path 1. An attacker posts a keyword-matching Telegram message whose text begins with a spreadsheet formula marker. 2. The crawler stores the message in SQLite. 3. An analyst exports the collected records in CSV format. 4. The ...[truncated 628 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
scripts/main.py:1092
Finding

Recursive Media Purge Accepts an Arbitrary Unvalidated Directory

Content
View full analysis
Remediation
View remediation

T05 · Unauthorized Access and Privilege Escalation

Warning
Location
SKILL.md:5
Finding

Skill Manifest Grants Tools Beyond the Documented Operational Requirements

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (74)

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · README.md (reported line 38)May include surrounding context.

bash
cd tg-crawler/config
cp .env.example .env

编辑 .env 文件,填入你的凭证:

Tp4

High
Category
MCP Tool Poisoning
Confidence
90% confidence
Finding

The declared description presents a broader Telegram intelligence collection tool with four operating modes and automatic data retention cleanup. The supplied code chunk only covers the discovery portion: finding channels/groups using Telegram search and search bots, then deduplicating and sorting results. There is no evidence here of message backfill, live monitoring, storage management, or cleanup of expired data. Additionally, the code relies on external search bots (xbso1/jisou), which is a meaningful behavior not reflected in the declared description. While this code is consistent with the 'discover' part of the description, the overall declared purpose materially overstates what this specific chunk does.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
90% confidence
Finding

The supplied code does not implement the declared tool’s primary operational behaviors of Telegram data collection, discovery/backfill/monitoring modes, or retention-based cleanup. Instead, it focuses on configuration management: loading target definitions from YAML, merging keyword rules, deriving identifiers, listing available profiles, and writing newly discovered targets into config files with backup/rollback and heuristic categorization. While config loading is plausibly a supporting component of such a tool, this chunk also performs undeclared capabilities—most notably modifying local configuration files and auto-classifying channels by industry. Because the actual behavior is materially narrower and partly different from the declared description, this code chunk is a mismatch.

Content

No source excerpt is available for this finding.

Credential Access

High
Category
Privilege Escalation
Confidence
91% confidence
Finding

The skill is designed to consume Telegram API credentials from a .env file and also supports verification-code and password flows, which constitutes credential access and use. In combination with allowed tools such as exec, read, write, and process, this meaningfully expands the risk of secret exposure, credential misuse, and unauthorized account actions if the skill is triggered improperly or if outputs/logs are not tightly controlled.

Content

Scanner excerpt · SKILL.md (reported line 76)May include surrounding context.

--keywords "生态词1,生态词2"
--backfill-keywords "品牌名1,品牌名2"
--targets <行业profile>
--env ../config/.env

text

各行业两步法完整命令示例 → `references/industry-playbook.md`。

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

This configuration explicitly includes a private invite-only Telegram group and even labels it as a high-risk operation. Joining private groups goes beyond passive public collection and can require account actions, expose operator identity, and create legal, policy, or entrapment-style risk if the tool accesses restricted communities without strong authorization controls.

Content

No source excerpt is available for this finding.

Lp1

High
Category
MCP Least Privilege
Confidence
75% confidence
Finding

The skill uses 'env' capability that is not listed in its permissions. This may indicate deceptive intent or missing permission declarations.

Content

No source excerpt is available for this finding.

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · README.md (reported line 34)May include surrounding context.

md
python main.py --backfill --backfill-limit 500

    # 指定配置路径
    python main.py --targets ../data/targets.yaml --env .env

环境变量 (.env):
    TG_API_ID     - Telegram API ID (从 my.telegram.org 获取)

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · README.md (reported line 205)May include surrounding context.

md
python main.py --backfill --backfill-limit 500

    # 指定配置路径
    python main.py --targets ../data/targets.yaml --env .env

环境变量 (.env):
    TG_API_ID     - Telegram API ID (从 my.telegram.org 获取)

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · references/industry-playbook.md (reported line 30)May include surrounding context.

md
python main.py --backfill --backfill-limit 500

    # 指定配置路径
    python main.py --targets ../data/targets.yaml --env .env

环境变量 (.env):
    TG_API_ID     - Telegram API ID (从 my.telegram.org 获取)

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · references/industry-playbook.md (reported line 40)May include surrounding context.

md
python main.py --backfill --backfill-limit 500

    # 指定配置路径
    python main.py --targets ../data/targets.yaml --env .env

环境变量 (.env):
    TG_API_ID     - Telegram API ID (从 my.telegram.org 获取)

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · references/industry-playbook.md (reported line 47)May include surrounding context.

md
python main.py --backfill --backfill-limit 500

    # 指定配置路径
    python main.py --targets ../data/targets.yaml --env .env

环境变量 (.env):
    TG_API_ID     - Telegram API ID (从 my.telegram.org 获取)

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · references/industry-playbook.md (reported line 76)May include surrounding context.

md
python main.py --backfill --backfill-limit 500

    # 指定配置路径
    python main.py --targets ../data/targets.yaml --env .env

环境变量 (.env):
    TG_API_ID     - Telegram API ID (从 my.telegram.org 获取)

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · references/industry-playbook.md (reported line 83)May include surrounding context.

md
python main.py --backfill --backfill-limit 500

    # 指定配置路径
    python main.py --targets ../data/targets.yaml --env .env

环境变量 (.env):
    TG_API_ID     - Telegram API ID (从 my.telegram.org 获取)

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · references/industry-playbook.md (reported line 93)May include surrounding context.

md
python main.py --backfill --backfill-limit 500

    # 指定配置路径
    python main.py --targets ../data/targets.yaml --env .env

环境变量 (.env):
    TG_API_ID     - Telegram API ID (从 my.telegram.org 获取)

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · references/industry-playbook.md (reported line 100)May include surrounding context.

md
python main.py --backfill --backfill-limit 500

    # 指定配置路径
    python main.py --targets ../data/targets.yaml --env .env

环境变量 (.env):
    TG_API_ID     - Telegram API ID (从 my.telegram.org 获取)

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · references/industry-playbook.md (reported line 110)May include surrounding context.

md
python main.py --backfill --backfill-limit 500

    # 指定配置路径
    python main.py --targets ../data/targets.yaml --env .env

环境变量 (.env):
    TG_API_ID     - Telegram API ID (从 my.telegram.org 获取)

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · references/industry-playbook.md (reported line 117)May include surrounding context.

md
python main.py --backfill --backfill-limit 500

    # 指定配置路径
    python main.py --targets ../data/targets.yaml --env .env

环境变量 (.env):
    TG_API_ID     - Telegram API ID (从 my.telegram.org 获取)

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · references/proxy-pool-setup.md (reported line 122)May include surrounding context.

md
python main.py --backfill --backfill-limit 500

    # 指定配置路径
    python main.py --targets ../data/targets.yaml --env .env

环境变量 (.env):
    TG_API_ID     - Telegram API ID (从 my.telegram.org 获取)

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · references/tg-crawler-architecture.md (reported line 21)May include surrounding context.

md
python main.py --backfill --backfill-limit 500

    # 指定配置路径
    python main.py --targets ../data/targets.yaml --env .env

环境变量 (.env):
    TG_API_ID     - Telegram API ID (从 my.telegram.org 获取)

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/main.py (reported line 15)May include surrounding context.

python
python main.py --backfill --backfill-limit 500

    # 指定配置路径
    python main.py --targets ../data/targets.yaml --env .env

环境变量 (.env):
    TG_API_ID     - Telegram API ID (从 my.telegram.org 获取)

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/main.py (reported line 673)May include surrounding context.

python
python main.py --backfill --backfill-limit 500

    # 指定配置路径
    python main.py --targets ../data/targets.yaml --env .env

环境变量 (.env):
    TG_API_ID     - Telegram API ID (从 my.telegram.org 获取)

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/main.py (reported line 713)May include surrounding context.

python
python main.py --backfill --backfill-limit 500

    # 指定配置路径
    python main.py --targets ../data/targets.yaml --env .env

环境变量 (.env):
    TG_API_ID     - Telegram API ID (从 my.telegram.org 获取)

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/main.py (reported line 816)May include surrounding context.

python
python main.py --backfill --backfill-limit 500

    # 指定配置路径
    python main.py --targets ../data/targets.yaml --env .env

环境变量 (.env):
    TG_API_ID     - Telegram API ID (从 my.telegram.org 获取)

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/main.py (reported line 821)May include surrounding context.

python
python main.py --backfill --backfill-limit 500

    # 指定配置路径
    python main.py --targets ../data/targets.yaml --env .env

环境变量 (.env):
    TG_API_ID     - Telegram API ID (从 my.telegram.org 获取)

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/main.py (reported line 683)May include surrounding context.

python
help="目标配置文件路径 或 行业 profile (gaming/retail/all)。逗号分隔可组合多个,如 gaming,retail"
    )
    parser.add_argument(
        "--env", default="../config/.env",
        help="环境变量文件路径 (默认: ../config/.env)"
    )
    parser.add_argument(

Static analysis

No suspicious patterns detected.