Back to skill

Security audit

music-with-comfyui

Security checks for vulnerabilities and agentic risk

Overview

The main music-generation skill is mostly disclosed and purpose-aligned, but the package includes an undeclared background watcher that can wait for a server and generate hard-coded songs without a current user request.

Install only if you are comfortable with a skill that can send your music tags and lyrics to the ComfyUI endpoint you configure and write generated audio locally. Review or remove artifact/_watch_and_generate.py before use, because it is not part of the disclosed on-demand workflow and can run a long wait-and-generate batch job with built-in content.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • System Prompt LeakageDirect Leakage, Indirect Extraction, Tool-Based Exfiltration
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
Findings (29)

Tainted flow: 'timeout' from os.environ.get (line 11, credential/environment) → urllib.request.urlopen (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.

Content

Scanner excerpt · _watch_and_generate.py (reported line 112)May include surrounding context.

python
def endpoint_up(url, timeout=3):
    try:
        with urllib.request.urlopen(url, timeout=timeout) as r:
            return r.status  # any HTTP status means it's listening
    except Exception:
        return False

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description presents an interactive skill for generating a song from user-supplied instructions via ComfyUI/AceStep. The supplied code instead waits for server availability and then batch-generates two specific built-in songs. This is a materially different primary behavior: automated fixed-content generation rather than user-directed composition. The ComfyUI/AceStep music-generation domain is related, but the implementation shown does not match the declared input/output behavior or trigger model.

Content

No source excerpt is available for this finding.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 163)May include surrounding context.

md
1. **Never re-read on repeat.** Don't re-read `SKILL.md`, `config.json`, or the script for a known config / repeat / same-job. Defaults are baked in.

Lp1

High
Category
MCP Least Privilege
Confidence
90% confidence
Finding

The script launches another Python program via subprocess, which is an undeclared shell/process-execution capability. In a skill ecosystem, undeclared process-spawning expands the attack surface because child processes inherit context and can perform actions beyond what users or reviewers expect.

Content

No source excerpt is available for this finding.

Lp1

High
Category
MCP Least Privilege
Confidence
90% confidence
Finding

The script launches another Python program via subprocess, which is an undeclared shell/process-execution capability. In a skill ecosystem, undeclared process-spawning expands the attack surface because child processes inherit context and can perform actions beyond what users or reviewers expect.

Content

No source excerpt is available for this finding.

Lp1

High
Category
MCP Least Privilege
Confidence
90% confidence
Finding

The script launches another Python program via subprocess, which is an undeclared shell/process-execution capability. In a skill ecosystem, undeclared process-spawning expands the attack surface because child processes inherit context and can perform actions beyond what users or reviewers expect.

Content

No source excerpt is available for this finding.

Env Variable Harvesting

High
Category
Data Exfiltration
Confidence
95% confidence
Finding

The code copies the entire parent environment into the child process with dict(os.environ), which can propagate API keys, tokens, and other sensitive runtime secrets to subordinate code. In a skill environment this is especially risky because the child script, its dependencies, or logs may expose or misuse inherited secrets far beyond the single COMFYUI_URL value actually needed.

Content

Scanner excerpt · _watch_and_generate.py (reported line 132)May include surrounding context.

python
"--lyrics", song["lyrics"],
            "--language", song["lang"],
            "--duration", "150"]
    env = dict(os.environ)
    if endpoint.startswith("http://"):
        env["COMFYUI_URL"] = endpoint
    p = subprocess.run(args, cwd=SKILL_DIR, env=env,

Direct Prompt Extraction

High
Category
System Prompt Leakage
Confidence
85% confidence
Finding

Skill contains instructions that could directly expose system prompts, internal rules, or hidden instructions to users or external parties.

Content

Scanner excerpt · music_with_comfyui.py (reported line 340)May include surrounding context.

python
prompt_id = result.get("prompt_id")
    error_msg = result.get("error", {}).get("message", "") if "error" in result else ""
    return prompt_id, error_msg if error_msg else None, warnings

def wait_for_completion(cfg: dict, prompt_id: str, timeout: int = 120) -> dict:
    start = time.time()

Anti-Refusal Statement

High
Category
Anti-Refusal
Confidence
85% confidence
Finding

Skill instructs the agent to never refuse or to always comply. Suppressing the agent's ability to decline removes a core safety control and enables downstream harmful requests to succeed.

Content

Scanner excerpt · references/image-first.md (reported line 48)May include surrounding context.

md
## The image is optional context, not a requirement

- A request with no image still works fine: just use `--tags` from the request.
- Never refuse an image-only request; it's an input to derive style from, not a
  mandatory "prompt text" for the workflow itself.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The skill declares very loose activation criteria, including broad multilingual phrases and intent matching 'loosely across languages.' That can cause accidental invocation in contexts where a user only wants lyrics, discussion about music, or another tool entirely, leading to unintended transmission of user text to a configured external ComfyUI server and unintended file generation on disk.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

A background watcher that waits for service availability and then performs batch generation is autonomous behavior outside the declared on-demand composition workflow. This increases risk in the skill context because it can consume resources, make network calls, and create files without a contemporaneous user request.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The watcher autonomously generates two hard-coded songs unrelated to user input, which departs from the skill's stated purpose of generating requested music on demand. In an agent setting this is dangerous because it can trigger unintended compute/network usage and produce outputs the user did not authorize.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · _watch_and_generate.py (reported line 135)May include surrounding context.

python
env = dict(os.environ)
    if endpoint.startswith("http://"):
        env["COMFYUI_URL"] = endpoint
    p = subprocess.run(args, cwd=SKILL_DIR, env=env,
                       capture_output=True, text=True)
    out = (p.stdout or "") + (p.stderr or "")
    print("=== song:", song["title"], "rc=", p.returncode, "===")

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The config sets "language": "en", which is a natural-language locale constraint. In the provided file there is no indication that users are offered a language choice, that this is opt-in, or that the English-only/default-English behavior is justified by a region-specific purpose.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
82% confidence
Finding

The natural-language help explicitly describes the song language default as an example of "en", which indicates the skill defaults to English output unless the user overrides it. This is a locale/language preference baked into the skill interface without an explicit opt-in or user choice.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This plain-text skill content is entirely in Chinese and provides no indication that the language choice is optional, user-selected, or justified by a region-specific requirement. Under the language/locale policy, forcing a specific language without opt-in is a natural-language policy violation.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The workflow explicitly stages generated audio in a persistent local outbound directory and only cleans it up after a successful send, but it provides no safeguards for failed sends, crashes, or concurrent sessions. This can leave user-generated media lingering on disk longer than expected, creating avoidable privacy and cross-session data exposure risk on shared or multi-tenant hosts.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The file states that Chinese/non-ASCII filenames reliably fail and instructs the operator to rename files to pure ASCII before sending. This is a language/locale-related constraint presented as a blanket rule, with no user choice or explicit policy justification.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The manifest explicitly says the skill is not for simple audio file conversion, but this delivery reference directs the agent to transcode generated audio to 44.1 kHz AAC for WhatsApp and to zip originals for document delivery. While delivery formatting is related to output handling, these are additional file-processing capabilities not justified by the stated generation-focused purpose.

Content

No source excerpt is available for this finding.

Scope Creep

Low
Category
Excessive Agency
Confidence
70% confidence
Finding

Skill's behavior or capabilities extend beyond its stated purpose. Scope creep allows an agent to perform actions unrelated to its documented functionality, increasing the attack surface.

Content

Scanner excerpt · SKILL.md (reported line 51)May include surrounding context.

md
## When to Use (Recognizing the Request)

This skill fires whenever the user asks to **write, compose, arrange, or make a song / melody / music track**, in any language — even if no genre or length is given. Common trigger phrases include, but are not limited to:

- **English:** "make me a song", "make up a song", "compose", "compose a song/track", "write a song", "arrange a song", "write some music", "make a melody", "produce a track", "write me a tune"
- **中文:** "写歌", "作曲", "编一首歌", "编一个歌", "编个曲", "编曲", "做首歌", "帮写首歌", "生成一首歌", "来首原创"

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
86% confidence
Finding

This file is named as an English text file, but line L36 includes Chinese characters ("盘" and "债") with no indication that multilingual output is optional or user-selected. That can violate a language/locale policy requiring skills or content to respect the user's chosen language unless explicitly documented.

Content

No source excerpt is available for this finding.

Vague Triggers

Low
Category
Not specified by scanner
Confidence
76% confidence
Finding

This is a manifest/config file, so vague natural-language configuration should be reviewed under SQP-1. The plain value "en" sets a language-related default but provides no surrounding documentation or constraints explaining when that default applies, whether users can override it, or whether non-English inputs are excluded.

Content

No source excerpt is available for this finding.

Dynamic attribute access via getattr()

Low
Category
Dangerous Code Execution
Confidence
50% confidence
Finding

Dynamic getattr() with a non-literal attribute name can access arbitrary object attributes, potentially bypassing access controls.

Content

Scanner excerpt · music_with_comfyui.py (reported line 523)May include surrounding context.

python
override = {}
    for flag_name in ("bpm", "keyscale", "language", "timesignature", "cfg_scale",
                      "temperature", "top_p", "top_k", "min_p", "generate_audio_codes"):
        override[flag_name] = getattr(args, flag_name, None)

    if args.prompt:
        tags = args.prompt

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
90% confidence
Finding

This plain-text file consists entirely of Chinese lyrics and stage directions, with no indication that language selection is optional or user-configurable. The policy category for all file types includes locale or language constraints that are imposed without user opt-in.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
83% confidence
Finding

The file's user-facing natural-language content is entirely in Chinese and provides no indication of optional language selection or locale choice. Under the stated policy, language-specific output without user opt-in can constitute a locale-policy issue unless clearly justified as region-specific.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.install_untrusted_source

Install source points to URL shortener or raw IP.

Warn
Code
suspicious.install_untrusted_source
Location
config.json:2