Back to skill

Security audit

Digital Singer

Security checks for vulnerabilities and agentic risk

Overview

This skill mostly matches a digital-singer app, but it exposes credentials and local control endpoints in ways users should review before installing.

Install only if you are comfortable reviewing and fixing the local server first: rotate/remove bundled API keys, protect or remove token/config endpoints, remove wildcard CORS, store NuwaAI credentials outside web-served paths, and replace shell-form FFmpeg execution with argument-based execution and song allowlists. Expect microphone audio to be sent to NuwaAI and chat text to be sent to the configured LLM provider.

Vulnerability Patterns
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (5)

T09 · Insecure Skill Coding Practices

Error
Location
server.mjs:313
Finding

Unauthenticated disclosure of NuwaAI credentials and access tokens

Content
View full analysis
Remediation
View remediation

T05 · Unauthorized Access and Privilege Escalation

Error
Location
server.mjs:339
Finding

Unauthenticated persistent overwrite of NuwaAI configuration

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
server.mjs:450
Finding

Shell command injection risk in the FFmpeg conversion endpoint

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
server.mjs:35
Finding

Hardcoded NuwaAI demo credential exposed in server source and token endpoint

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
chorus_agent.py:11
Finding

Hardcoded Dashscope API key in the Python agent

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
Findings (40)

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The documented behavior materially overstates what the skill actually does, including claims about NuwaAI avatar control, lip-sync, ASR-based singing analysis, and scoring. Description-behavior mismatch is dangerous because users may consent to run or trust the skill under false assumptions, which undermines informed consent, auditability, and safe deployment.

Content

No source excerpt is available for this finding.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 50)May include surrounding context.

md
The singing host (conversation agent) needs an OpenAI-compatible LLM API. Configure in `server.mjs` or via environment variables:

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 69)May include surrounding context.

md
The singing host (conversation agent) needs an OpenAI-compatible LLM API. Configure in `server.mjs` or via environment variables:

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 73)May include surrounding context.

md
The singing host (conversation agent) needs an OpenAI-compatible LLM API. Configure in `server.mjs` or via environment variables:

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 125)May include surrounding context.

md
The singing host (conversation agent) needs an OpenAI-compatible LLM API. Configure in `server.mjs` or via environment variables:

Missing User Warnings

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

A hardcoded live API key in source code is a direct secret-exposure vulnerability. Anyone with repository, package, log, or screenshot access can reuse the credential to incur cost, access associated services, or pivot into other connected resources.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The manifest promises digital-avatar singing with lip sync, but the implementation is only a CLI loop that plays local MP3 files. This mismatch can mislead reviewers into approving capabilities or permissions the code does not need, increasing the chance of overbroad trust and unnoticed behavior changes.

Content

No source excerpt is available for this finding.

Known Vulnerable Dependency: ws==8.20.0 — 2 advisory(ies): CVE-2026-45736 (ws: Uninitialized memory disclosure); CVE-2026-48779 (ws: Memory exhaustion DoS from tiny fragments and data chunks)

High
Category
Supply Chain
Confidence
96% confidence
Finding

The lockfile pins ws to version 8.20.0, and the finding cites known advisories for uninitialized memory disclosure and memory-exhaustion denial of service in that version. Because this skill appears to rely on WebSocket-based real-time audio/avatar interaction, an exposed or attacker-reachable WebSocket path would make these issues more relevant, especially for service disruption and possible data leakage.

Content

No source excerpt is available for this finding.

Known Vulnerable Dependency: ws==8.20.0 — 2 advisory(ies): CVE-2026-45736 (ws: Uninitialized memory disclosure); CVE-2026-48779 (ws: Memory exhaustion DoS from tiny fragments and data chunks)

High
Category
Supply Chain
Confidence
98% confidence
Finding

The package depends on ws 8.20.0, which is reported to have known vulnerabilities including potential uninitialized memory disclosure and memory exhaustion denial of service. Because ws is a WebSocket library typically exposed to untrusted remote input, exploitation could leak sensitive process memory or allow attackers to crash or degrade the singing/avatar service remotely.

Content

No source excerpt is available for this finding.

Hidden Instructions

High
Category
Prompt Injection
Confidence
70% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · public/index.html (reported line 111)May include surrounding context.

html
</head>
<body>

<!-- Welcome -->
<div id="welcome">
  <div class="welcome-box">
    <div class="hero">🎤</div>

Hidden Instructions

High
Category
Prompt Injection
Confidence
70% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · public/index.html (reported line 111)May include surrounding context.

html
</head>
<body>

<!-- Welcome -->
<div id="welcome">
  <div class="welcome-box">
    <div class="hero">🎤</div>

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

A live demo NuwaAI API key is hardcoded in the source and exposed through an endpoint that mints tokens on demand. Embedded service credentials are easily extracted and abused by anyone with code or API access, potentially leading to account takeover of the demo tenant, quota theft, and unauthorized use billed to the owner.

Content

No source excerpt is available for this finding.

Missing User Warnings

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The embedded demo API key is a real secret baked into code and later transmitted to NuwaAI without any user-facing warning. The main security issue is secret exposure: anyone with access to the repository or running service can recover and reuse the credential for unauthorized external API access.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
89% confidence
Finding

The skill documentation describes capabilities that require environment access, network access, and shell execution, but it does not declare any tool scope or permissions. This weakens transparency and reviewability, making it easier for a skill to overreach at runtime or for users to grant unsafe execution implicitly.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The skill states that ASR voice capture is used for user singing but does not clearly disclose that microphone audio may be captured and transmitted to NuwaAI or related services. Voice data is sensitive biometric-adjacent information, so insufficient notice can lead to privacy harm and non-compliant data handling.

Content

No source excerpt is available for this finding.

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
70% confidence
Finding

Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.

Content

Scanner excerpt · SKILL.md (reported line 93)May include surrounding context.

macOS

brew install ffmpeg

Ubuntu/Debian

sudo apt install ffmpeg

text

### 6. Node.js 18+ (Required)

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The file's natural-language description and all user-facing prompts are written to operate in Chinese, with no indication that the user may choose another language. This can violate language or locale policy where skills must not force a specific language without user opt-in.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The skill metadata says it is powered by a NuwaAI humanctrl API, but the code actually sends data to DashScope/Qwen and performs only local audio playback. This is a security-relevant integrity issue because users and integrators may make trust, privacy, and network-allowlisting decisions based on false implementation claims.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · chorus_agent.py (reported line 230)May include surrounding context.

python
part_label = "上半段" if part == "upper" else "下半段"

    # 同步播放,等播放完成再返回
    current_process = subprocess.Popen(
        ["afplay", file_path],
        stdout=subprocess.DEVNULL,
        stderr=subprocess.DEVNULL,

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The manifest says the skill includes scoring for duet/battle mode, implying some relationship to the user's actual singing. Here, battle_evaluate assigns both user and agent random scores and titles without analyzing any vocal input, making the implemented behavior a novelty generator rather than singing-performance scoring.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The code sends the full conversation history to an external LLM service even though the skill description focuses narrowly on digital singing. In this context, users may reasonably expect a local media feature, so undisclosed off-device transmission increases privacy and data-governance risk.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

User conversation data is transmitted to a third-party model API without any visible disclosure, consent flow, or data-minimization controls. In a karaoke-style skill, users may share personal text or voice-related content casually, making silent export to an external provider more sensitive than the manifest suggests.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
95% confidence
Finding

This request transmits conversation content and tool definitions to an external endpoint. External transmission is expected for cloud LLM use, but in this skill it is security-relevant because the feature description does not clearly justify or disclose remote processing, so user data can leave the host unexpectedly.

Content

Scanner excerpt · chorus_agent.py (reported line 317)May include surrounding context.

python
"tools": TOOLS,
    }

    resp = requests.post(DASHSCOPE_BASE_URL, headers=headers, json=payload, timeout=60)
    resp.raise_for_status()
    return resp.json()

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The code starts microphone capture and streams audio to a remote WebSocket service immediately after connection, but the UI does not provide an upfront privacy notice describing that speech is transmitted off-device for ASR/avatar control. In a karaoke/chat context users may expect mic use, but they may not expect continuous remote streaming as soon as the socket opens, which creates consent and privacy risk.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The page collects an API key in the browser and submits it to backend configuration storage without any visible warning about persistence, scope, or who can later use it. Because this skill is designed for end users and the key grants access to a third-party avatar service, silent storage increases the risk of credential misuse, accidental sharing across users, or unexpected retention.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.dangerous_exec, suspicious.env_credential_access, suspicious.exposed_secret_literal

Shell command execution detected (child_process).

Critical
Code
suspicious.dangerous_exec
Location
server.mjs:462

Environment variable access combined with network send.

Critical
Code
suspicious.env_credential_access
Location
server.mjs:21

File appears to expose a hardcoded API secret or token.

Critical
Code
suspicious.exposed_secret_literal
Location
chorus_agent.py:11

File appears to expose a hardcoded API secret or token.

Critical
Code
suspicious.exposed_secret_literal
Location
public/index.html:787

File appears to expose a hardcoded API secret or token.

Critical
Code
suspicious.exposed_secret_literal
Location
server.mjs:37