Back to skill

Security audit

Smart Speak Multilingual TTS

Security checks for vulnerabilities and agentic risk

Overview

This skill is a purpose-aligned text-to-speech workflow, with some privacy and overwrite caveats users should understand before using it.

Install only if you are comfortable using edge-tts/ffmpeg for audio generation. Avoid sending confidential or regulated text unless you accept the TTS provider data flow, and choose a unique output path in the workspace to avoid overwriting an existing MP3.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (8)

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The documented behavior promises intelligent Pinyin-to-Hanzi conversion, language detection, and automatic voice selection, but the implementation reportedly does not perform those functions. This mismatch can mislead users and downstream agents into trusting transformations and language handling that never occur, causing unsafe automation decisions, incorrect content generation, or unexpected processing of user data.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
93% confidence
Finding

The skill instructs use of shell execution and file output but declares no explicit tool scope or permissions. This weakens least-privilege controls and can allow the agent to invoke more powerful capabilities than reviewers or policy expect, increasing the blast radius if the skill is misused or later modified.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The document specifies exact locale-bound voices (vi-VN, zh-CN, en-US) as the default behavior and says the skill 'must' assign them during synthesis. This can violate language/locale choice policy because it forces accent/locale variants without mentioning user preference, opt-in, or alternatives.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The script sends segment text to edge-tts, which typically relies on an external network-backed TTS service, without any notice, consent flow, or data-sensitivity check. If users provide private, regulated, or confidential text, the skill may exfiltrate that content to a third party unexpectedly.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/smart_speak.py (reported line 51)May include surrounding context.

python
"--write-media", temp_file
                ]
                
                result = subprocess.run(cmd, capture_output=True, text=True)
                if result.returncode != 0:
                    print(f"Error generating TTS for segment {i}: {result.stderr}", file=sys.stderr)
                    continue

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/smart_speak.py (reported line 72)May include surrounding context.

python
args.output
        ]
        
        result = subprocess.run(ffmpeg_cmd, capture_output=True, text=True)
        if result.returncode != 0:
            print(f"Error merging audio with ffmpeg: {result.stderr}", file=sys.stderr)
            sys.exit(1)

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
87% confidence
Finding

This markdown file describes generating output to an absolute filesystem path and delivering the resulting MP3, which affects user data on disk. The instructions do not include any warning or note about creating files, choosing a safe output path, or avoiding overwriting existing files.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

The script defaults to vi-VN-HoaiMyNeural, imposing a specific language/locale when a segment does not specify a voice. There is no documented opt-in or user choice mechanism for this locale default in the file.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.