Back to skill

Security audit

Web TTS Speaker

Security checks for vulnerabilities and agentic risk

Overview

The skill does what it claims, but it can automatically fetch arbitrary URLs and send extracted content to an external TTS service without enough scoping or consent.

Review before installing. Use this only in an environment where arbitrary URL fetching is acceptable, avoid private/internal URLs or sensitive text, and prefer adding confirmation plus public-web URL validation before deployment. Pin dependencies and remove text snippets from machine-readable output if logs or shared channels may contain sensitive data.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T09 · Insecure Skill Coding Practices

Error
Location
cli.py:77
Finding

Unrestricted URL Fetching Enables Server-Side Request Forgery

Content
View full analysis

Vulnerability Details

File Location: cli.py, lines 77-86; invocation at lines 269-271
Vulnerability Type: Server-Side Request Forgery
Risk Level: High

Vulnerable Code

python
def extract_web_text(url):
    """提取网页正文"""
    import requests
    from bs4 import BeautifulSoup

    headers = {"User-Agent": "Mozilla/5.0"}
    r = requests.get(url, timeout=15, headers=headers)
    r.encoding = r.apparent_encoding

    soup = BeautifulSoup(r.text, "html.parser")
    for tag in soup(["script", "style", "nav", "footer", "header", "aside"]):
        tag.decompose()

    text = soup.get_text(separator="\n", strip=True)
    lines = [l.strip() for l in text.split("\n") if len(l.strip()) > 15]
    return "\n".join(lines)

The user-controlled URL reaches this function through:

python
if args.url:
    print(f"\U0001f4d6 读取网页:{args.url}")
    text = extract_web_text(args.url)

Technical Analysis

The application passes a user-controlled URL directly to requests.get() without validating the URL scheme, destination hostname, resolved IP address, port, or redirect chain. Python Requests also follows redirects by default.

Consequently, a URL may target loopback addresses, private network ranges, link-local services, cloud instance metadata endpoints, or internal hostnames accessible from the Agent's execution environment. An initially public URL could also redirect to one of these destinations.

This behavior exceeds the minimum network privileges required for reading public webpages because no boundary distinguishes public web content from internal network resources. Extracted response text is subsequently supplied to Edge TTS, creating a potential path by which readable internal content is transmitted to an external TTS provider.

Attack Path

  1. An attacker asks the Agent to read a URL controlled by the attacker or directly supplies an internal-service URL.
  2. The Ski ...[truncated 1277 chars]
Remediation
View remediation

Remediation Suggestions

  • Accept only explicitly supported http and https URLs.
  • Reject URLs containing credentials, ambiguous host syntax, unsupported ports, or malformed hostnames.
  • Resolve the destination hostname before connecting and reject every address in loopback, private, link-local, multicast, unspecified, and reserved ranges for both IPv4 and IPv6.
  • Disable automatic redirects or validate the scheme, hostname, and resolved address again at every redirect hop.
  • Apply the same destination policy immediately before connection to reduce DNS-rebinding risk.
  • Consider an allowlist of approved public domains when the operating environment permits it.
  • Route webpage retrieval through a restricted proxy with egress controls rather than granting the Skill unrestricted network access.
  • Impose response-size and content-type limits before parsing or forwarding content.
  • Require explicit user confirmation before fetching a URL rather than automatically treating every standalone URL as a TTS request.
  • Clearly disclose that extracted text is sent to Edge TTS, and do not submit content obtained from internal or sensitive sources to the external provider.

T08 · Insecure Dependencies

Warning
Location
requirements.txt:1
Finding

Unbounded and Unverified Third-Party Dependency Installation

Content
View full analysis

Vulnerability Details

File Location: requirements.txt, lines 1-3; installation declaration in SKILL.md, lines 8-11
Vulnerability Type: Non-reproducible dependency resolution and supply-chain exposure
Risk Level: Medium

Vulnerable Code

text
edge-tts>=6.0.0
beautifulsoup4>=4.12.0
requests>=2.31.0

The Skill installation metadata directs pip to consume this file:

yaml
install:
  - id: python
    kind: pip
    package: "-r requirements.txt"

Technical Analysis

All dependencies use open-ended minimum-version constraints. A future installation may therefore resolve to package versions that did not exist and were not reviewed when the Skill was published. No lock file or package hashes are present to verify the exact artifacts being installed.

The package names are consistent with the declared TTS and webpage-processing functionality, and the audit found no evidence that they are deliberate typosquats or currently malicious. The security issue is that installation results are not reproducible and implicitly trust all future compatible releases selected by pip and the configured package index.

Python dependency installation and import both execute third-party code. If a permitted future release or package-index artifact is compromised, that code would execute with the same operating-system permissions as the process installing or running the Skill.

Attack Path

  1. An attacker compromises a permitted dependency release, its maintainer account, or the package-index delivery path.
  2. A user installs or reinstalls the Skill after the compromised version becomes the newest version satisfying the open-ended constraint.
  3. Pip resolves and installs the compromised artifact because no exact version or integrity hash prevents it.
  4. Malicious package code executes during installation or when cli.py imports the dependency.
  5. The package receives the filesystem, netw ...[truncated 639 chars]
Remediation
View remediation

Remediation Suggestions

  • Pin each dependency to an exact version that has been reviewed and tested.
  • Generate a lock file containing transitive dependencies rather than constraining only direct dependencies.
  • Require cryptographic hashes for every resolved package artifact, for example through a hash-locked requirements file and pip's --require-hashes mode.
  • Install only from a trusted, explicitly configured package index.
  • Review dependency release notes and security advisories before updating pinned versions.
  • Use automated dependency scanning while retaining a review gate for lock-file changes.
  • Perform installation in an isolated virtual environment or container with minimal filesystem and network privileges.
  • Separate dependency installation privileges from the Agent runtime account wherever possible.
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
Findings (23)

Tainted flow: 'cmd' from requests.get (line 162, network input) → subprocess.run (code execution)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

External input (network, user) flows to a code execution sink. This enables remote code execution or command injection.

Content

Scanner excerpt · cli.py (reported line 168)May include surrounding context.

python
*cfg["params"],
            out_path,
        ]
        subprocess.run(cmd, capture_output=True, text=True, check=True)
        dur = get_duration(out_path)
        size = os.path.getsize(out_path)
        results[ch] = {

Tainted flow: 'cmd' from requests.get (line 162, network input) → subprocess.run (code execution)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

External input (network, user) flows to a code execution sink. This enables remote code execution or command injection.

Content

Scanner excerpt · cli.py (reported line 198)May include surrounding context.

python
*cfg["params"],
            out_path,
        ]
        subprocess.run(cmd, capture_output=True, text=True, check=True)
        dur = get_duration(out_path)
        size = os.path.getsize(out_path)
        results[ch] = {

Tainted flow: 'merged_wav' from requests.get (line 190, network input) → subprocess.run (code execution)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

External input (network, user) flows to a code execution sink. This enables remote code execution or command injection.

Content

Scanner excerpt · cli.py (reported line 226)May include surrounding context.

python
f.write(f"file '{wav_path}'\n")

        merged_wav = os.path.join(out_dir, f"_merged_{os.getpid()}.wav")
        subprocess.run(
            [FFMPEG, "-y", "-f", "concat", "-safe", "0", "-i", concat_list, "-c", "copy", merged_wav],
            capture_output=True, text=True, check=True,
        )

Hidden Instructions

High
Category
Prompt Injection
Confidence
60% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · SKILL.md (reported line 1)May include surrounding context.

md
---
name: web-tts-speaker
version: 3.2.1
description: "网页一键朗读:URL/文本 → TTS语音 → 多渠道自动匹配(飞书语音条 / 微信/TG/Discord MP3)"

Vague Triggers

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The trigger conditions are overly broad: the skill auto-activates on any URL and on common phrases like '读一下' or '帮我读'. This can cause unintended execution, fetch attacker-controlled URLs, and send user content to external processing without clear confirmation, increasing the chance of surprise data exfiltration or misuse.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The natural-language instructions, descriptions, and examples are entirely in Chinese, which effectively forces a specific language for users reading and operating the skill. The policy allows locale constraints only when users are given a choice or when the restriction is clearly documented and justified, neither of which appears here.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The skill does not warn users that supplied text/URLs will be retrieved over the network, processed by external TTS-related components, and converted into audio files for delivery across messaging channels. This undermines informed consent and can expose sensitive content if users provide private text or internal-only URLs.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The docstring materially understates what the script does: it fetches arbitrary URLs and sends text to a cloud-backed TTS service. Misleading operational claims can cause unsafe deployment decisions, especially in agent contexts where operators may assume there is no outbound data transfer beyond local conversion.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill sends user-provided text or extracted webpage content to external network services without explicit warning or consent messaging. In an agent setting, this can disclose sensitive internal data, private URLs, or copyrighted material to third parties unexpectedly.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · cli.py (reported line 134)May include surrounding context.

python
def get_duration(audio_path):
    """用 ffprobe 获取音频时长(秒)"""
    try:
        result = subprocess.run(
            [
                FFMPEG.replace("ffmpeg.exe", "ffprobe.exe"),
                "-v", "quiet",

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · cli.py (reported line 168)May include surrounding context.

python
*cfg["params"],
            out_path,
        ]
        subprocess.run(cmd, capture_output=True, text=True, check=True)
        dur = get_duration(out_path)
        size = os.path.getsize(out_path)
        results[ch] = {

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · cli.py (reported line 198)May include surrounding context.

python
*cfg["params"],
            out_path,
        ]
        subprocess.run(cmd, capture_output=True, text=True, check=True)
        dur = get_duration(out_path)
        size = os.path.getsize(out_path)
        results[ch] = {

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
73% confidence
Finding

The concat manifest is generated by writing raw temporary file paths into ffmpeg's concat file using single-quoted entries without escaping embedded quote characters. On platforms where temp or output paths can contain apostrophes, this can corrupt the concat manifest and potentially cause ffmpeg to interpret unintended paths or directives, making the media-processing step less trustworthy.

Content

Scanner excerpt · cli.py (reported line 226)May include surrounding context.

python
f.write(f"file '{wav_path}'\n")

        merged_wav = os.path.join(out_dir, f"_merged_{os.getpid()}.wav")
        subprocess.run(
            [FFMPEG, "-y", "-f", "concat", "-safe", "0", "-i", concat_list, "-c", "copy", merged_wav],
            capture_output=True, text=True, check=True,
        )

Ssd 3

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The agent-readable markers echo a snippet of the input text into stdout/loggable output. If the source text contains secrets, personal data, or internal content, this creates an avoidable secondary disclosure channel across logs, downstream agents, or the wrong messaging channel.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The single-channel output path has the same leakage issue: it includes part of the user/web content in structured marker output intended for downstream automation. In multi-tenant or logged environments, this can expose sensitive text beyond the original requester's context.

Content

No source excerpt is available for this finding.

Ae4

Medium
Category
analysis-evasion
Confidence
80% confidence
Finding

Suspicious Unicode normalization or mixed-script content

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The manifest description and tags are entirely Chinese-centric, including Chinese channel/use-case wording and a Chinese voice example, but the file does not indicate that the skill is region-specific or that users can choose another language/locale. This creates a natural-language policy concern because the skill appears to assume a fixed language context without opt-in or justification.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
83% confidence
Finding

This markdown file states that the skill extracts text from webpages, generates speech, and outputs files, but it does not include any warning about contacting external services or writing audio artifacts to disk. Under the markdown criteria for SQP-2, user-facing documentation should disclose behaviors that may affect privacy or user data handling.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
89% confidence
Finding

The default voice is hard-coded to zh-CN-XiaoxiaoNeural, which imposes a specific language/locale by default. Although --voice can override it, the skill description and default behavior do not explicitly present language choice or justify the locale restriction as region-specific.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
90% confidence
Finding

Using edge-tts>=6.0.0 leaves the installed version open-ended, which creates non-reproducible environments and avoidable supply-chain exposure. If a later release introduces a vulnerability or malicious code, downstream users may ingest it automatically during installation.

Content

Scanner excerpt · requirements.txt (reported line 1)May include surrounding context.

text
edge-tts>=6.0.0
beautifulsoup4>=4.12.0
requests>=2.31.0

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
92% confidence
Finding

Using beautifulsoup4>=4.12.0 allows future releases to be installed without review, reducing reproducibility and increasing supply-chain risk if a bad or incompatible version is later published. While not an immediate exploit, unpinned dependencies make builds non-deterministic and can unexpectedly pull vulnerable or malicious releases.

Content

Scanner excerpt · requirements.txt (reported line 2)May include surrounding context.

text
edge-tts>=6.0.0
beautifulsoup4>=4.12.0
requests>=2.31.0

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
96% confidence
Finding

Using requests>=2.31.0 permits installation of any later version and does not guarantee that a known-safe release is used in all environments. This weakens build integrity and can expose consumers to vulnerable or malicious upstream versions if dependency resolution changes.

Content

Scanner excerpt · requirements.txt (reported line 3)May include surrounding context.

text
edge-tts>=6.0.0
beautifulsoup4>=4.12.0
requests>=2.31.0

Unverifiable Dependency: requests has 16 known advisory(ies) (CVE-2014-1830 (Exposure of Sensitive Information to an Unauthorized Actor in Requests); CVE-2024-47081 (Requests vulnerable to .netrc credentials leak via malicious URLs); CVE-2024-35195 (Requests `Session` object does not verify requests after making first request wi) +13 more), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
94% confidence
Finding

The manifest includes requests without an exact version pin, and requests has multiple known advisories across versions. Because the allowed range is broad, it is impossible to verify from this file alone whether installations will avoid affected releases, leaving consumers at risk of resolving to a vulnerable version.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.