Back to skill

Security audit

Free voice from Comfy UI + Qwen3 TTS

Security checks for vulnerabilities and agentic risk

Overview

This skill is a straightforward ComfyUI text-to-speech workflow with local file output and local HTTP use that are disclosed and aligned with its purpose.

Install only if you use this Windows ComfyUI setup and are comfortable with your text being sent to the local ComfyUI API and the resulting audio being saved in the listed output folder. Review or change the hardcoded paths and language if they do not match your environment.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
Findings (4)

Ae4

Medium
Category
analysis-evasion
Confidence
80% confidence
Finding

Suspicious Unicode normalization or mixed-script content

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill states that generated audio is saved to a fixed local directory but does not warn the user that outputs will persist on disk in a predictable location. This can expose sensitive spoken content to other local users, backups, indexing services, or later unintended disclosure if the generated audio contains private information.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The example request sets "language": "Russian", and the skill description does not offer a language choice or explain that the skill is intentionally limited to Russian. This creates a natural-language policy issue because the skill forces a specific language without documented user opt-in or clear regional justification.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The skill instructs the agent to send user-provided text to a local HTTP service at localhost:8000 without warning the user that their content will be transmitted outside the agent boundary. Even though the service is local, it is a separate process that may log prompts, expose them through history endpoints, or be bound insecurely, creating confidentiality and consent risks.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.