Back to skill

Security audit

qqbot-stt

Security checks for vulnerabilities and agentic risk

Overview

This appears to be a real local speech-to-text skill, but it needs review because it handles voice data through weakly hardened services and mutable installs.

Install only if you are comfortable running a local service that processes voice messages. Keep it bound to 127.0.0.1 unless you add authentication and firewall controls, add upload size and rate limits, fix temporary-file cleanup and mktemp usage, avoid logging transcript contents, protect OpenClaw credentials, and pin reviewed dependency and installer versions before using it with untrusted users.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (4)

T09 · Insecure Skill Coding Practices

Error
Location
main.py:25
Finding

Unauthenticated Network Service Permits Unbounded Resource Consumption

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
server.py:70
Finding

Original Audio Uploads Are Not Deleted After Successful Format Conversion

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
transcribe.py:23
Finding

Race-Prone Temporary Filename Can Enable Local File Overwrite

Content
View full analysis
str: """Convert audio to a 16 kHz mono WAV using ffmpeg.""" wav_path = tempfile.mktemp(suffix=".wav") try: result = subprocess.run( ["ffmpeg", "-y", "-i", input_path, "-ar", "16000", "-ac", "1", "-f", "wav", wav_path], capture_output=True, timeout=60 ) if result.returncode == 0 and os.path.exists(wav_path): return wav_path except (FileNotFoundError, subprocess.TimeoutExpired): pass return "" ``` ### Technical Analysis `tempfile.mktemp()` generates a candidate pathname but does not atomically create and reserve the file. A time-of-check/time-of-use window exists between generating the path and FFmpeg opening it. A local attacker who can observe or predict the temporary filename may create that path first, potentially as a symbolic link to another file. Because FFmpeg is invoked with `-y`, it is instructed to overwrite an existing output without prompting. If the operating system and FFmpeg follow the attacker-created link, the target can be overwritten with WAV output. The subprocess argument list prevents shell command injection, but it does not mitigate the temporary-file race. ### Attack Path 1. A victim runs `transcribe.py` on a file requiring conversion. 2. `tempfile.mktemp()` selects a path but does not create it. 3. A local attacker identifies the selected path during the race window. 4. The attacker creates the path or a symbolic link at that location, pointing to a file writable by the victim account. 5. FFmpeg opens the attacker-controlled path with overwrite behavior enabled. 6. The linked target may be truncated or overwritten with generated audio data. Exploitation requires local access and the ability to win the r ...[truncated 436 chars]
Remediation
View remediation

T08 · Insecure Dependencies

Warning
Location
requirements.txt:1
Finding

Mutable and Unpinned Dependencies Create Supply-Chain Exposure

Content
View full analysis
Remediation
View remediation
uvicorn== python-multipart== mlx-qwen3-asr== ``` 2. Generate a lockfile that includes transitive dependencies. 3. Use hashes with `pip --require-hashes` to verify downloaded artifacts. 4. Replace `npx clawhub@latest` with a reviewed, exact version. 5. Document the trusted package indexes and avoid untrusted extra indexes. 6. Run dependency vulnerability and provenance checks in CI. 7. Perform installation in an isolated virtual environment under an unprivileged account. 8. Regularly update pinned versions through a controlled review and testing process rather than resolving mutable releases during deployment. ]]>
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Rogue AgentSelf-Modification, Session Persistence
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (23)

Possible Typosquatting: 'uvicorn' resembles popular package 'gunicorn'

High
Category
Supply Chain
Confidence
70% confidence
Finding

Package name closely resembles a popular package, suggesting possible typosquatting. Attackers publish malicious packages with similar names to trick developers into installing them.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
92% confidence
Finding

The README instructs users to run npx clawhub@latest install local-stt, which fetches and executes the latest remote package without pinning a version. If the upstream package is compromised or a breaking/malicious release is published, users may execute unreviewed code during installation.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The documentation at L203 says the CLI mode works for Telegram and not QQBot, framing the tutorial around a channel limitation. But the later architectural explanation at L387-L397 and summary at L524 say tools.media.audio is framework-level and automatically shared across channels, which directly contradicts the earlier claim about channel applicability.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · README.md (reported line 248)May include surrounding context.

创建文件 ~/.openclaw/scripts/qwen3_asr_cli.py:

bash
mkdir -p ~/.openclaw/scripts
nano ~/.openclaw/scripts/qwen3_asr_cli.py

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · README.md (reported line 248)May include surrounding context.

创建文件 ~/.openclaw/scripts/qwen3_asr_cli.py:

bash
mkdir -p ~/.openclaw/scripts
nano ~/.openclaw/scripts/qwen3_asr_cli.py

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

In the comparison table, L494 states the CLI method applies to '仅 Telegram' only. However, the surrounding documentation at L387-L397 and L524 describes tools.media.audio as a shared framework-level configuration used automatically by all channels, making the table's applicability statement contradictory rather than merely incomplete.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
91% confidence
Finding

The documentation instructs users to run npx clawhub@latest install local-stt, which pulls and executes the latest remote package version without pinning. This creates a supply-chain risk: if the package is compromised or a breaking/malicious release is published, users may execute attacker-controlled code during installation.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The skill processes user voice messages and transcribes them, but the documentation does not warn operators about privacy, consent, retention, or handling of potentially sensitive audio content. In a QQ bot context, this can lead to unintentional collection or processing of personal data without adequate notice or controls.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The module documentation presents this as a 'Local STT API Server' wrapping qwen-asr in an OpenAI-compatible format. However, the actual transcription logic does not implement STT itself; it invokes another script at a hard-coded path using subprocess, meaning the server is primarily a wrapper/launcher rather than the STT implementation. This is a meaningful intent/documentation mismatch, not just an omitted detail.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · main.py (reported line 49)May include surrounding context.

python
try:
        # 调用 qwen-asr
        result = subprocess.run(
            [
                sys.executable,  # 使用当前 Python 解释器
                "-m", "uv", "run",

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · server.py (reported line 39)May include surrounding context.

python
def convert_to_wav(src: str) -> str | None:
    wav = src + ".wav"
    try:
        r = subprocess.run(
            ["ffmpeg", "-y", "-i", src, "-ar", "16000", "-ac", "1", "-f", "wav", wav],
            capture_output=True, timeout=60)
        if r.returncode == 0 and os.path.exists(wav):

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The server logs the first 80 characters of the transcription text, which may contain sensitive user-provided speech content. Although logging exists, there is no user-facing warning in the code comments, docstring, or endpoint description that uploaded audio will be transcribed and partially recorded in logs.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
89% confidence
Finding

The script invokes an external binary (ffmpeg) on user-supplied files, which expands the attack surface to parser bugs in ffmpeg and allows resource-consumption or malformed-media attacks. In addition, it uses tempfile.mktemp(), which is race-prone and can let a local attacker pre-create or redirect the output path, potentially causing unintended file overwrite or symlink abuse.

Content

Scanner excerpt · transcribe.py (reported line 24)May include surrounding context.

python
"""用 ffmpeg 转换为 16kHz 单声道 WAV"""
    wav_path = tempfile.mktemp(suffix=".wav")
    try:
        result = subprocess.run(
            ["ffmpeg", "-y", "-i", input_path,
             "-ar", "16000", "-ac", "1", "-f", "wav", wav_path],
            capture_output=True, timeout=60

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

The guide shows API credentials being placed directly into a config file but does not warn about file permissions, secret storage, or avoiding accidental disclosure. Users may leave long-lived secrets in plaintext configs that can be exposed through backups, logs, screenshots, or weak local permissions.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
85% confidence
Finding

The configuration example includes apiKey and clientSecret fields but does not warn readers to protect secrets or avoid committing them to files and repositories. This can encourage insecure secret handling and accidental credential exposure, especially when users copy the example verbatim into config files.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
97% confidence
Finding

The dependency list leaves FastAPI unpinned, which makes builds non-reproducible and can silently introduce vulnerable or breaking versions over time. In a web-facing stack, this increases supply-chain risk because deployment behavior depends on whatever version is resolved at install time.

Content

Scanner excerpt · requirements.txt (reported line 1)May include surrounding context.

text
fastapi
uvicorn
python-multipart

Unverifiable Dependency: fastapi has 3 known advisory(ies) (CVE-2021-32677 (Cross-Site Request Forgery (CSRF) in FastAPI); CVE-2021-32677 (FastAPI is a web framework for building APIs with Python 3.6+ based on standard ); CVE-2024-24762 (FastAPI is a web framework for building APIs with Python 3.8+ based on standard )), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
88% confidence
Finding

FastAPI has known advisories, and because no version is specified, there is no way to verify whether the deployed package is affected. In a web application framework, even low-severity uncertainty is a real supply-chain weakness because vulnerable framework versions can expose the entire API surface.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
97% confidence
Finding

The dependency list leaves Uvicorn unpinned, so installations may resolve to different versions across environments and over time. For an internet-exposed ASGI server, that uncertainty can result in pulling versions with known security flaws or incompatible behavior.

Content

Scanner excerpt · requirements.txt (reported line 2)May include surrounding context.

text
fastapi
uvicorn
python-multipart

Unverifiable Dependency: uvicorn has 4 known advisory(ies) (CVE-2020-7694 (Log injection in uvicorn); CVE-2020-7695 (HTTP response splitting in uvicorn); CVE-2020-7694 (This affects all versions of package uvicorn. The request logger provided by the) +1 more), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
90% confidence
Finding

Uvicorn has multiple known advisories, and the unpinned requirement prevents validation that the installed version is safe. Since Uvicorn is the HTTP server entry point, affected versions could enable request handling or logging issues that are directly reachable by remote clients.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
98% confidence
Finding

python-multipart is unpinned despite a history of multipart parsing vulnerabilities, making it impossible to determine whether installs are affected by known DoS issues. Because multipart parsers often process attacker-controlled request bodies, an unsafe resolved version could expose the service to resource-exhaustion attacks.

Content

Scanner excerpt · requirements.txt (reported line 3)May include surrounding context.

text
fastapi
uvicorn
python-multipart

Unverifiable Dependency: python-multipart has 16 known advisory(ies) (CVE-2024-24762 (python-multipart vulnerable to Content-Type Header ReDoS); CVE-2024-53981 (Denial of service (DoS) via deformation `multipart/form-data` boundary); CVE-2026-53539 (python-multipart: Quadratic-time querystring parsing with semicolon separators c) +13 more), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
96% confidence
Finding

python-multipart has numerous advisories, including denial-of-service classes relevant to multipart/form-data parsing, and the manifest does not pin a safe version. In this stack, that is more dangerous because FastAPI commonly uses python-multipart for user-supplied form and file upload parsing, making vulnerable code plausibly reachable from external requests.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
93% confidence
Finding

The module docstring and user-facing usage/output text are entirely in Chinese, and the CLI description/help strings later in the file are also fixed to Chinese. Under the language/locale policy, hard-coding a single language without opt-in or choice is a natural-language policy issue unless clearly justified.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
90% confidence
Finding

The parser description and argument help strings are user-facing natural-language content, and they are fixed to one language with no user opt-in. This continues the same locale restriction in operational CLI interactions and can violate language-choice policy if not justified.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.