T09 · Insecure Skill Coding Practices
- Location
main.py:25- Finding
Unauthenticated Network Service Permits Unbounded Resource Consumption
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This appears to be a real local speech-to-text skill, but it needs review because it handles voice data through weakly hardened services and mutable installs.
Install only if you are comfortable running a local service that processes voice messages. Keep it bound to 127.0.0.1 unless you add authentication and firewall controls, add upload size and rate limits, fix temporary-file cleanup and mktemp usage, avoid logging transcript contents, protect OpenClaw credentials, and pin reviewed dependency and installer versions before using it with untrusted users.
main.py:25Unauthenticated Network Service Permits Unbounded Resource Consumption
server.py:70Original Audio Uploads Are Not Deleted After Successful Format Conversion
transcribe.py:23Race-Prone Temporary Filename Can Enable Local File Overwrite
requirements.txt:1Mutable and Unpinned Dependencies Create Supply-Chain Exposure
Package name closely resembles a popular package, suggesting possible typosquatting. Attackers publish malicious packages with similar names to trick developers into installing them.
The README instructs users to run npx clawhub@latest install local-stt, which fetches and executes the latest remote package without pinning a version. If the upstream package is compromised or a breaking/malicious release is published, users may execute unreviewed code during installation.
The documentation at L203 says the CLI mode works for Telegram and not QQBot, framing the tutorial around a channel limitation. But the later architectural explanation at L387-L397 and summary at L524 say tools.media.audio is framework-level and automatically shared across channels, which directly contradicts the earlier claim about channel applicability.
Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.
创建文件 ~/.openclaw/scripts/qwen3_asr_cli.py:
mkdir -p ~/.openclaw/scripts
nano ~/.openclaw/scripts/qwen3_asr_cli.py
Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.
创建文件 ~/.openclaw/scripts/qwen3_asr_cli.py:
mkdir -p ~/.openclaw/scripts
nano ~/.openclaw/scripts/qwen3_asr_cli.py
In the comparison table, L494 states the CLI method applies to '仅 Telegram' only. However, the surrounding documentation at L387-L397 and L524 describes tools.media.audio as a shared framework-level configuration used automatically by all channels, making the table's applicability statement contradictory rather than merely incomplete.
The documentation instructs users to run npx clawhub@latest install local-stt, which pulls and executes the latest remote package version without pinning. This creates a supply-chain risk: if the package is compromised or a breaking/malicious release is published, users may execute attacker-controlled code during installation.
The skill processes user voice messages and transcribes them, but the documentation does not warn operators about privacy, consent, retention, or handling of potentially sensitive audio content. In a QQ bot context, this can lead to unintentional collection or processing of personal data without adequate notice or controls.
The module documentation presents this as a 'Local STT API Server' wrapping qwen-asr in an OpenAI-compatible format. However, the actual transcription logic does not implement STT itself; it invokes another script at a hard-coded path using subprocess, meaning the server is primarily a wrapper/launcher rather than the STT implementation. This is a meaningful intent/documentation mismatch, not just an omitted detail.
subprocess module calls execute external commands. Without careful input validation, this enables command injection.
try:
# 调用 qwen-asr
result = subprocess.run(
[
sys.executable, # 使用当前 Python 解释器
"-m", "uv", "run",
subprocess module calls execute external commands. Without careful input validation, this enables command injection.
def convert_to_wav(src: str) -> str | None:
wav = src + ".wav"
try:
r = subprocess.run(
["ffmpeg", "-y", "-i", src, "-ar", "16000", "-ac", "1", "-f", "wav", wav],
capture_output=True, timeout=60)
if r.returncode == 0 and os.path.exists(wav):
The server logs the first 80 characters of the transcription text, which may contain sensitive user-provided speech content. Although logging exists, there is no user-facing warning in the code comments, docstring, or endpoint description that uploaded audio will be transcribed and partially recorded in logs.
The script invokes an external binary (ffmpeg) on user-supplied files, which expands the attack surface to parser bugs in ffmpeg and allows resource-consumption or malformed-media attacks. In addition, it uses tempfile.mktemp(), which is race-prone and can let a local attacker pre-create or redirect the output path, potentially causing unintended file overwrite or symlink abuse.
"""用 ffmpeg 转换为 16kHz 单声道 WAV"""
wav_path = tempfile.mktemp(suffix=".wav")
try:
result = subprocess.run(
["ffmpeg", "-y", "-i", input_path,
"-ar", "16000", "-ac", "1", "-f", "wav", wav_path],
capture_output=True, timeout=60
The guide shows API credentials being placed directly into a config file but does not warn about file permissions, secret storage, or avoiding accidental disclosure. Users may leave long-lived secrets in plaintext configs that can be exposed through backups, logs, screenshots, or weak local permissions.
The configuration example includes apiKey and clientSecret fields but does not warn readers to protect secrets or avoid committing them to files and repositories. This can encourage insecure secret handling and accidental credential exposure, especially when users copy the example verbatim into config files.
The dependency list leaves FastAPI unpinned, which makes builds non-reproducible and can silently introduce vulnerable or breaking versions over time. In a web-facing stack, this increases supply-chain risk because deployment behavior depends on whatever version is resolved at install time.
fastapi
uvicorn
python-multipart
FastAPI has known advisories, and because no version is specified, there is no way to verify whether the deployed package is affected. In a web application framework, even low-severity uncertainty is a real supply-chain weakness because vulnerable framework versions can expose the entire API surface.
The dependency list leaves Uvicorn unpinned, so installations may resolve to different versions across environments and over time. For an internet-exposed ASGI server, that uncertainty can result in pulling versions with known security flaws or incompatible behavior.
fastapi
uvicorn
python-multipart
Uvicorn has multiple known advisories, and the unpinned requirement prevents validation that the installed version is safe. Since Uvicorn is the HTTP server entry point, affected versions could enable request handling or logging issues that are directly reachable by remote clients.
python-multipart is unpinned despite a history of multipart parsing vulnerabilities, making it impossible to determine whether installs are affected by known DoS issues. Because multipart parsers often process attacker-controlled request bodies, an unsafe resolved version could expose the service to resource-exhaustion attacks.
fastapi
uvicorn
python-multipart
python-multipart has numerous advisories, including denial-of-service classes relevant to multipart/form-data parsing, and the manifest does not pin a safe version. In this stack, that is more dangerous because FastAPI commonly uses python-multipart for user-supplied form and file upload parsing, making vulnerable code plausibly reachable from external requests.
The module docstring and user-facing usage/output text are entirely in Chinese, and the CLI description/help strings later in the file are also fixed to Chinese. Under the language/locale policy, hard-coding a single language without opt-in or choice is a natural-language policy issue unless clearly justified.
The parser description and argument help strings are user-facing natural-language content, and they are fixed to one language with no user opt-in. This continues the same locale restriction in operational CLI interactions and can violate language-choice policy if not justified.
No suspicious patterns detected.