Back to skill

Security audit

Piper Tts Engine

Security checks for vulnerabilities and agentic risk

Overview

This TTS skill is mostly purpose-aligned, but its network API example and dependency guidance need review before installation.

Install only after reviewing the API deployment path. Keep the service bound to localhost unless remote access is required, add real authentication, TLS, rate limits, request size limits, output cleanup, and run it as an unprivileged user. Use isolated, pinned dependencies, and only train custom voices from recordings you are authorized to use.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T08 · Insecure Dependencies

Warning
Location
SKILL.md:287
Finding

Unverified and Unpinned Third-Party Package Installation

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
SKILL.md:201
Finding

Network-Exposed TTS API Lacks Implemented Authentication and Resource Controls

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (8)

Intent-Code Divergence

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The description explicitly says 'Use when 需要文本翻译、多语言转换、本地化处理时使用', which describes translation/localization use cases rather than speech synthesis. The rest of the file consistently documents TTS, voice training, SSML, and audio generation, so this is an active contradiction between intent text and actual documented behavior.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The description emphasizes a '本地离线文字转语音引擎' focused on local TTS capabilities, but later documentation adds serving the engine over HTTP with FastAPI and curl-accessible endpoints. Exposing a network service is a materially broader behavior than simple offline local synthesis and is not clearly scoped in the primary description text.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The description says "Use when 需要文本翻译、多语言转换、本地化处理时使用," which is a broad natural-language trigger covering translation and localization tasks rather than narrowly scoping the skill to text-to-speech. This ambiguity increases the chance the skill is invoked for common language tasks that do not actually fit the stated TTS purpose.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The custom voice training workflow uses user recordings and transcripts, which may contain biometric voice data and sensitive personal information. Without explicit consent, retention, and privacy warnings, users may process regulated or non-consensual voice data, creating privacy, compliance, and impersonation risks.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The API deployment section presents a network-accessible text submission endpoint without a clear warning that inputs leave local-only processing and become exposed over a service boundary. In practice, users may deploy it without authentication or transport protection, leading to interception of submitted text, abuse of the service, or leakage of generated content paths.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
93% confidence
Finding

The documented curl example sends user-provided text over plain HTTP to a service bound on 0.0.0.0, which is an external transmission of potentially sensitive content without transport protection. If used outside an isolated environment, attackers on the network could observe or tamper with requests, and unauthorized users could access the endpoint.

Content

Scanner excerpt · SKILL.md (reported line 230)May include surrounding context.

uvicorn tts_api_server:app --host 0.0.0.0 --port 8100

...

示例

curl -X POST http://localhost:8100/api/tts
-H "Content-Type: application/json"
-d '{"text":"您好,这是一条测试语音","voice":"zh_CN-huayan-medium"}'

text

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill claims API access should be restricted and HTTPS used, yet the example server binds to 0.0.0.0 and demonstrates unauthenticated plain-HTTP access. If copied into real use, this can expose submitted text, generated audio paths, and service control to any reachable network peer, enabling unauthorized use and data exposure.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
76% confidence
Finding

The description explicitly says "支持中文交互,无需复杂配置即开即用" while the skill also advertises multilingual support, but it does not indicate that users can choose another interaction language. This can be read as a language-default policy baked into the skill description without opt-in.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.