Back to skill

Security audit

smart-tts

Security checks for vulnerabilities and agentic risk

Overview

This skill is a straightforward cloud text-to-speech helper that uses a DashScope API key and writes generated audio files as described.

Install only if you are comfortable using DashScope for speech synthesis. Text you synthesize is sent to that service, so do not submit secrets, private records, or confidential business content. Check the output path first because existing WAV files with the same names may be overwritten.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (10)

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
70% confidence
Finding

Without declared permissions the skill's intent is opaque and cannot be validated.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The manifest description and the entire user-facing markdown content are presented in Chinese, with no indication that users may choose another language or that the skill is intentionally limited to a Chinese-speaking audience. This can violate language/locale policy when a skill forces a specific language without user opt-in or documented justification.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The script sends provided text to an external DashScope TTS service, which creates a privacy and data-handling risk if users supply sensitive content without realizing it leaves the local environment. Because the tool also accepts arbitrary CLI input or JSON-loaded text, confidential operational or personal data could be transmitted to a third party.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The module docstring is written entirely in Chinese and presents the skill as a general-purpose TTS tool, while the default sample text is also Chinese-only. Because the file does not offer a user language choice or explain that the skill is intentionally limited to Chinese, this appears to force a specific language/locale without opt-in.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
85% confidence
Finding

The manifest describes a smart text-to-speech skill that retries different models/voices until synthesis succeeds. While calling a remote TTS service is expected, directly reading credentials from process environment introduces a credential-access capability that is not mentioned in the stated purpose and is broader than 'automatic model/voice retry'.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

When no arguments are provided, the script defaults to generating the Chinese phrase "你好,这是智能语音合成测试". For a generic TTS skill, hardcoding a single language as the default behavior can violate language/locale policy unless the limitation is clearly documented or the user is given a choice.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

The markdown states that speech is automatically saved and names a default output file, but it does not explicitly warn users that running the skill will write or overwrite data on disk. For a skill that performs file output, a brief user-facing disclosure would help set expectations about where artifacts are stored and the potential for replacing an existing file.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
78% confidence
Finding

The file's natural-language title, comments, and default spoken text are entirely in Chinese, and the default behavior appears tailored to a single language without any visible opt-in or language-selection mechanism. Under the policy, forcing a specific language without user choice can be a locale/language policy issue unless clearly justified.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Low
Category
Not specified by scanner
Confidence
79% confidence
Finding

The manifest describes an intelligent speech synthesis skill that retries models/voices until success. While networked TTS itself is expected, directly reading credentials from the runtime environment is a separate capability that is not stated in the manifest and is called out as context-inappropriate in this audit model.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
83% confidence
Finding

This code performs file writes by saving synthesized audio to output_path, but the function itself provides no prior warning, confirmation, or explanatory comment about creating or overwriting files. Although it prints after saving, that is not a pre-action disclosure, and there is no visible README/SKILL.md context in this file to warn users about the write behavior.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.