Back to skill

Security audit

Gemini Tts

Security checks for vulnerabilities and agentic risk

Overview

This skill is a small Gemini text-to-speech wrapper whose network and file behavior matches its stated purpose, with some documentation and quality gaps.

Install only if you are comfortable sending the text you provide to Google's Gemini API using your GEMINI_API_KEY. Do not rely on the --voice option for custom persona output unless the script is fixed, because the current implementation always uses Puck and writes the result to output_voice.<type> in the working directory.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Taint TrackingDirect Taint Flow, Variable-Mediated Taint Flow, Credential Exfiltration Chain
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (7)

Tainted flow: 'req' from os.environ.get (line 26, credential/environment) → urllib.request.urlopen (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.

Content

Scanner excerpt · generate_voice.py (reported line 29)May include surrounding context.

python
req = urllib.request.Request(url, data=json.dumps(payload).encode('utf-8'), headers=headers, method='POST')
    
    try:
        with urllib.request.urlopen(req) as response:
            result = json.loads(response.read().decode('utf-8'))
            
            # 打印详细的返回结构,看看是不是有 inline_data 和 mime_type

Tp4

High
Category
MCP Tool Poisoning
Confidence
94% confidence
Finding

The primary purpose generally matches the description: this is a text-to-speech script using Gemini 2.5 Flash TTS and writing returned audio to disk. There are no evident unrelated triggers or undeclared sensitive capabilities beyond normal network access to the Gemini API and local file output as part of TTS generation. However, the description emphasizes custom/persona-driven voice output, while the implementation does not use the voice_id parameter or the --voice argument at all. Instead, it always requests the fixed prebuilt voice "Puck." That makes the declared customization/persona behavior materially overstated relative to the actual code.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
91% confidence
Finding

The skill advertises code capabilities that require environment-variable access and outbound network access, but the manifest does not declare any tool scope or permission boundaries. This creates an authorization and review gap: operators cannot clearly see or constrain what the skill needs, and an agent runtime may grant broader access than intended when handling API keys or network requests.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The manifest describes persona-driven voice output, and the CLI/function both accept a voice identifier, but the request payload ignores that input and hard-codes the Gemini voice_name to "Puck". This means the skill does not behave as described: it generates TTS, but not caller-selected persona-driven output.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The function signature accepts voice_id and the CLI exposes a --voice option, implying that callers can choose the generated voice. However, the implementation disregards that parameter and always sends "Puck" as the prebuilt voice, which contradicts the apparent documented/interface intent.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

This code sends the user-supplied --text content to an external Google API and writes returned audio bytes to a local file, but it provides no prior notice that input text leaves the local system or that a file will be created. The only visible output is after the request and write complete, which does not serve as advance disclosure.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
85% confidence
Finding

The inline comments are written in Chinese, which imposes a specific language in natural-language content embedded in the code. There is no indication that this locale choice is intentional, optional, or justified for a region-specific skill.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.