T09 · Insecure Skill Coding Practices
- Location
scripts/tts.py:53- Finding
Unrestricted Download of an API-Controlled URL
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
The skill appears to be a real Zvukogram TTS helper, but its download and audio-merge helpers use broader, weakly validated network and local-file behavior than users are clearly told.
Install only if you are comfortable sending TTS text, account email, and API token to Zvukogram. Avoid using it for secrets or regulated data. Prefer environment variables or a protected config file, avoid GET URLs with credentials, and harden or review the downloader and merge helper before using it in automated or shared environments.
scripts/tts.py:53Unrestricted Download of an API-Controlled URL
scripts/merge.py:18Predictable Temporary File and FFmpeg Concat-List Injection
references/API.md:23API Documentation Encourages Credentials in URL Query Parameters
Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.
req.add_header("Content-Type", "application/x-www-form-urlencoded")
try:
with urllib.request.urlopen(req, timeout=30) as response:
result = json.loads(response.read().decode())
return result.get("balans")
except Exception as e:
The declared description presents a feature-rich TTS generation skill, but the actual code only checks the user's Zvukogram account balance. This is a materially different primary purpose: account/account-credit inspection versus speech synthesis. While both relate to the same service, the implemented capability is undeclared and the advertised TTS features are absent from this code chunk.
The declared description centers on a full TTS capability using the Zvukogram API, including SSML and various speech-generation features. The provided code only concatenates input audio files using ffmpeg by writing a temporary concat list and invoking the ffmpeg binary. While audio fragment merging is mentioned in the description, this code implements only that narrow helper function and none of the primary declared TTS behaviors. Therefore the code chunk does not accurately represent the declared purpose and is a material mismatch.
The core purpose broadly matches text-to-speech via the Zvukogram API, including basic speed control and audio generation. However, the description substantially overstates the implemented functionality. The code only accepts plain text or file input, passes basic fields (token, email, voice, text, format, speed) to a single API endpoint, and downloads the resulting audio. There is no code handling SSML parsing/validation, stress marks, English transcription, fragment merging, rich SSML references, or podcast-specific patterns. Additionally, the script reads credentials from a local config file or environment variables and writes audio to a local output file, which are operational behaviors not reflected in the declared permissions. This is therefore a description-behavior mismatch due to materially missing declared capabilities, even though the primary TTS purpose is aligned.
The pipeline trusts the API response field .file and feeds it directly into xargs curl, causing a second network request to an attacker-controlled or compromised URL if the service response is manipulated. This creates an unsafe chaining pattern that can be abused for unexpected outbound requests, retrieval of malicious content, or SSRF-like behavior from the user’s environment.
-d "token=$TOKEN" -d "email=$EMAIL" \
-d "voice=Алена" -d "speed=1.2" \
-d "text=Доброе утро! Новости ИИ на $(date +%d.%m.%Y)" \
-d "format=mp3" | jq -r '.file' | xargs curl -s -L -o /tmp/p1.mp3
# ... more parts ...
The skill declares capabilities that imply access to environment variables, filesystem writes, network calls, and shell execution, but it does not declare any explicit tool scope or allowed-tools boundary. That creates unnecessary privilege ambiguity and can let a caller invoke a skill with broader access than users or reviewers would reasonably expect.
Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.
---
name: zvukogram
description: Text-to-Speech via Zvukogram API with SSML support. Use when you need to generate speech from text, create podcasts, voice notifications, or work with audio. Supports speed control, stress marks, English word transcription, audio fragment merging, rich SSML references, and podcast-oriented TTS patterns.
metadata:
openclaw:
requires:
The documentation tells users to provide API credentials and send text to Zvukogram, but it does not warn that user-provided content and account identifiers are transmitted to a third-party service. This can cause unintended disclosure of sensitive prompts, names, notifications, or other private text to an external provider.
The documentation prescribes Russian stress marks and transliteration patterns for English words, and the listed voices are exclusively Russian-language voices. There is no explicit statement that the skill is Russian-locale-specific or that users can opt into another language/locale, which may violate language/locale policy expectations.
The example sends API credentials and user-provided text to a third-party service but does not warn that the content leaves the local environment. Users may unknowingly transmit sensitive or regulated text to an external provider, creating privacy, compliance, and credential-handling risk.
The manifest describes a text-to-speech skill using the Zvukogram API with audio fragment merging, but this example introduces local process execution via ffmpeg through Python subprocess. Invoking external binaries is a materially broader capability than API-based TTS generation and is not explicitly justified in the manifest text for this skill file.
The automation example posts credentials and generated content to a remote API in a backgroundable workflow without any warning or consent checkpoint. That increases the chance of routine or bulk exfiltration of sensitive text to an external service and normalizes embedding credentials in scripts.
This command explicitly transmits credentials and text to an external endpoint. In the context of a TTS skill, such transmission is expected functionally, but it is still security-relevant because users may provide confidential content or mishandle tokens when copying the example.
OUTPUT="podcast_${DATE}.mp3"
# Generate parts
curl -s -X POST "https://zvukogram.com/index.php?r=api/text" \
-d "token=$TOKEN" -d "email=$EMAIL" \
-d "voice=Алена" -d "speed=1.2" \
-d "text=Доброе утро! Новости ИИ на $(date +%d.%m.%Y)" \
The file defines a single mandatory pronunciation/transcription scheme using Cyrillic renderings throughout, such as 'OpenAI → Оупен Эй Ай' and the tip rules for English letters and syllables. Because it does not offer a language/locale option or explain that this is a justified region-specific resource, it creates a natural-language locale policy concern under the requirement to avoid forcing a specific language without user opt-in.
The guide states it is for making names, brands, and acronyms sound right in Russian TTS, which imposes a specific language/locale behavior. Because the file does not indicate user opt-in, alternatives, or that this is strictly a region-specific tool, it appears to force a locale-specific output mode.
This code includes natural-language descriptions and user-visible messages exclusively in Russian, which effectively forces a specific language on users. The file does not provide any opt-in, fallback, or documented justification for the locale constraint.
This code file contains natural-language strings and descriptions exclusively in Russian, including the module docstring and later CLI/help output. The stated policy flags language or locale constraints when a skill forces a specific language without user opt-in, and this file does not provide any alternative language or justification for the restriction.
subprocess module calls execute external commands. Without careful input validation, this enables command injection.
]
try:
result = subprocess.run(cmd, capture_output=True, text=True)
list_file.unlink() # Удаляем временный файл
if result.returncode == 0:
The manifest describes a text-to-speech skill centered on the Zvukogram API and SSML features, with audio fragment merging as a functional outcome. This implementation performs local process execution of ffmpeg, which is a materially different capability than calling a TTS API and is not explicitly justified in the manifest text.
The script transmits user-supplied text along with account identifiers (token and email) to a third-party TTS service, but it does not clearly warn the user that potentially sensitive content will leave the local environment. In a skill context, users may assume a local transformation, so this creates a real privacy and data-handling risk, especially if prompts, secrets, or personal data are passed as text.
This markdown file documents that every request must include a token and email and allows sending them via HTTP requests, but it does not warn users that the skill will transmit account credentials/identifiers to a third-party API. Under the markdown-specific warning criterion, this is a privacy-relevant behavior that should be disclosed explicitly.
Although the text says the endpoint deletes generated project/audio from the server, the document does not present this as a caution or warning to users about destructive behavior. For markdown files, destructive actions that can affect user data should be clearly signposted so users understand the risk before using the skill.
The examples consistently prescribe Russian text and Russian voice selections such as Алена and Андрей, with no indication that users may choose another language or locale. This can be a natural-language locale policy issue if organizational guidance requires user choice rather than a fixed language assumption.
The file presents examples, aliases, and explanatory content primarily in Russian Cyrillic without stating that the skill is region-specific or that users may choose another language. This can violate language/locale policy expectations when a skill implicitly forces one locale without opt-in.
This markdown content mixes English headings with Russian voice names and Russian-oriented style descriptors, but it does not tell users that the skill or reference is limited to Russian-language voices. Under the policy rule, forcing or assuming a specific language/locale without opt-in can be a natural-language policy issue when the constraint is not explicitly justified.
No suspicious patterns detected.