Back to skill

Security audit

ifly-hyper-tts

Security checks for vulnerabilities and agentic risk

Overview

This is a disclosed iFlytek text-to-speech skill, with some documentation and dependency hygiene issues but no evidence of hidden, destructive, persistent, or unrelated behavior.

Install only if you are comfortable sending the text you synthesize to iFlytek's service using your own API credentials. Avoid sensitive text unless you have reviewed iFlytek's terms and retention practices, use a virtual environment, pin `websocket-client`, and expect the current script to need a bug fix for the undefined `omni_params` variable before synthesis works reliably.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T08 · Insecure Dependencies

Warning
Location
README.md:8
Finding

Unpinned Third-Party Dependency Installation

Content
View full analysis
Remediation
View remediation
``` 2. Generate and record SHA-256 hashes for all required distributions, then install with hash enforcement: ```bash python3 -m pip install --require-hashes -r requirements.txt ``` 3. Commit the reviewed requirements or lockfile to the Skill package so dependency resolution is reproducible. 4. Recommend installation inside a dedicated virtual environment rather than the system Python environment: ```bash python3 -m venv .venv . .venv/bin/activate python3 -m pip install --require-hashes -r requirements.txt ``` 5. Use a trusted package index explicitly where appropriate, review dependency updates before changing the pin, and automate vulnerability and integrity checks in the release process. 6. Avoid recommending installation or execution with administrative or root privileges. ]]>
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (10)

Tp4

High
Category
MCP Tool Poisoning
Confidence
93% confidence
Finding

The documentation describes capabilities and parameters that are inconsistent with the implementation, including unsupported output formats and a role parameter that is not actually wired through. Security-relevant mismatches like this can cause callers to make unsafe assumptions about what data is sent, how outputs are produced, and whether guardrails are enforced, increasing the chance of unintended external transmission or operational failure.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
88% confidence
Finding

The skill appears to require environment access, file reads, and network access, but it does not declare an explicit tool scope or permission boundary. This creates a confused-deputy risk where the agent may invoke the skill without clear operator awareness that user text and local resources will be accessed and transmitted externally.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The trigger phrases are broad everyday language such as requests to 'read this text aloud,' making accidental or context-inappropriate activation more likely. In this skill's context, accidental invocation is more dangerous because activation causes user text to be transmitted to a third-party TTS provider and stored as a local file.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The description does not clearly warn that supplied text is sent to an external TTS service and that generated audio is written to a local file. This omission undermines informed consent and can lead users or orchestrating agents to disclose sensitive text to a third party without realizing it.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The invocation keywords are ambiguous and insufficiently constrained, which increases the risk that ordinary user phrasing triggers the skill unintentionally. Because this skill performs network transmission and local file output, false activation can expose sensitive text or create unexpected artifacts without the user's informed consent.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The FAQ advertises unsupported audio formats and an implicit_watermark parameter that contradict the declared interface and output behavior. This can mislead users or downstream agents into requesting unsupported options, with the result that data may be sent to the external service under false assumptions about format handling or provenance controls such as watermarking.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill sends user-supplied text to a third-party WebSocket TTS service, which can expose sensitive or private content if users are not clearly informed. In an agent skill context, users may ask it to read secrets, personal data, or internal documents aloud, making silent off-device transmission a real privacy risk.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The module and method documentation describe a text-to-speech client with controls like voice, speed, pitch, volume, and role, but the implementation logs omni_params, which is never defined anywhere in the file. This indicates the code was adapted from a different feature set and no longer matches the documented/current intent of this skill, causing the advertised synthesis flow to fail unexpectedly.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

One section says the skill '默认只开放以下 7 个精选发音人' and later repeats that it '默认只支持 7 个精选发音人'. But the usage examples and '官方默认免费发音人' section explicitly tell users to invoke additional voices such as x5_lingfeiyi_flow, indicating the effective scope is broader than the restriction stated in the documentation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

The natural-language description and defaults are framed around Chinese voices and Chinese-language usage, with no explicit opt-in or language-selection guidance in the user-facing instructions. Under the language/locale policy, forcing a specific language or locale without user choice can be a policy concern unless clearly justified.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.