T09 · Insecure Skill Coding Practices
- Location
server.mjs:313- Finding
Unauthenticated disclosure of NuwaAI credentials and access tokens
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This skill mostly matches a digital-singer app, but it exposes credentials and local control endpoints in ways users should review before installing.
Install only if you are comfortable reviewing and fixing the local server first: rotate/remove bundled API keys, protect or remove token/config endpoints, remove wildcard CORS, store NuwaAI credentials outside web-served paths, and replace shell-form FFmpeg execution with argument-based execution and song allowlists. Expect microphone audio to be sent to NuwaAI and chat text to be sent to the configured LLM provider.
server.mjs:313Unauthenticated disclosure of NuwaAI credentials and access tokens
server.mjs:339Unauthenticated persistent overwrite of NuwaAI configuration
server.mjs:450Shell command injection risk in the FFmpeg conversion endpoint
server.mjs:35Hardcoded NuwaAI demo credential exposed in server source and token endpoint
chorus_agent.py:11Hardcoded Dashscope API key in the Python agent
The documented behavior materially overstates what the skill actually does, including claims about NuwaAI avatar control, lip-sync, ASR-based singing analysis, and scoring. Description-behavior mismatch is dangerous because users may consent to run or trust the skill under false assumptions, which undermines informed consent, auditability, and safe deployment.
Referenced artifact was not completely inspected
The singing host (conversation agent) needs an OpenAI-compatible LLM API. Configure in `server.mjs` or via environment variables:
Referenced artifact was not completely inspected
The singing host (conversation agent) needs an OpenAI-compatible LLM API. Configure in `server.mjs` or via environment variables:
Referenced artifact was not completely inspected
The singing host (conversation agent) needs an OpenAI-compatible LLM API. Configure in `server.mjs` or via environment variables:
Referenced artifact was not completely inspected
The singing host (conversation agent) needs an OpenAI-compatible LLM API. Configure in `server.mjs` or via environment variables:
A hardcoded live API key in source code is a direct secret-exposure vulnerability. Anyone with repository, package, log, or screenshot access can reuse the credential to incur cost, access associated services, or pivot into other connected resources.
The manifest promises digital-avatar singing with lip sync, but the implementation is only a CLI loop that plays local MP3 files. This mismatch can mislead reviewers into approving capabilities or permissions the code does not need, increasing the chance of overbroad trust and unnoticed behavior changes.
The lockfile pins ws to version 8.20.0, and the finding cites known advisories for uninitialized memory disclosure and memory-exhaustion denial of service in that version. Because this skill appears to rely on WebSocket-based real-time audio/avatar interaction, an exposed or attacker-reachable WebSocket path would make these issues more relevant, especially for service disruption and possible data leakage.
The package depends on ws 8.20.0, which is reported to have known vulnerabilities including potential uninitialized memory disclosure and memory exhaustion denial of service. Because ws is a WebSocket library typically exposed to untrusted remote input, exploitation could leak sensitive process memory or allow attackers to crash or degrade the singing/avatar service remotely.
Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.
</head>
<body>
<!-- Welcome -->
<div id="welcome">
<div class="welcome-box">
<div class="hero">🎤</div>
Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.
</head>
<body>
<!-- Welcome -->
<div id="welcome">
<div class="welcome-box">
<div class="hero">🎤</div>
A live demo NuwaAI API key is hardcoded in the source and exposed through an endpoint that mints tokens on demand. Embedded service credentials are easily extracted and abused by anyone with code or API access, potentially leading to account takeover of the demo tenant, quota theft, and unauthorized use billed to the owner.
The embedded demo API key is a real secret baked into code and later transmitted to NuwaAI without any user-facing warning. The main security issue is secret exposure: anyone with access to the repository or running service can recover and reuse the credential for unauthorized external API access.
The skill documentation describes capabilities that require environment access, network access, and shell execution, but it does not declare any tool scope or permissions. This weakens transparency and reviewability, making it easier for a skill to overreach at runtime or for users to grant unsafe execution implicitly.
The skill states that ASR voice capture is used for user singing but does not clearly disclose that microphone audio may be captured and transmitted to NuwaAI or related services. Voice data is sensitive biometric-adjacent information, so insufficient notice can lead to privacy harm and non-compliant data handling.
Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.
brew install ffmpeg
sudo apt install ffmpeg
### 6. Node.js 18+ (Required)
The file's natural-language description and all user-facing prompts are written to operate in Chinese, with no indication that the user may choose another language. This can violate language or locale policy where skills must not force a specific language without user opt-in.
The skill metadata says it is powered by a NuwaAI humanctrl API, but the code actually sends data to DashScope/Qwen and performs only local audio playback. This is a security-relevant integrity issue because users and integrators may make trust, privacy, and network-allowlisting decisions based on false implementation claims.
subprocess module calls execute external commands. Without careful input validation, this enables command injection.
part_label = "上半段" if part == "upper" else "下半段"
# 同步播放,等播放完成再返回
current_process = subprocess.Popen(
["afplay", file_path],
stdout=subprocess.DEVNULL,
stderr=subprocess.DEVNULL,
The manifest says the skill includes scoring for duet/battle mode, implying some relationship to the user's actual singing. Here, battle_evaluate assigns both user and agent random scores and titles without analyzing any vocal input, making the implemented behavior a novelty generator rather than singing-performance scoring.
The code sends the full conversation history to an external LLM service even though the skill description focuses narrowly on digital singing. In this context, users may reasonably expect a local media feature, so undisclosed off-device transmission increases privacy and data-governance risk.
User conversation data is transmitted to a third-party model API without any visible disclosure, consent flow, or data-minimization controls. In a karaoke-style skill, users may share personal text or voice-related content casually, making silent export to an external provider more sensitive than the manifest suggests.
This request transmits conversation content and tool definitions to an external endpoint. External transmission is expected for cloud LLM use, but in this skill it is security-relevant because the feature description does not clearly justify or disclose remote processing, so user data can leave the host unexpectedly.
"tools": TOOLS,
}
resp = requests.post(DASHSCOPE_BASE_URL, headers=headers, json=payload, timeout=60)
resp.raise_for_status()
return resp.json()
The code starts microphone capture and streams audio to a remote WebSocket service immediately after connection, but the UI does not provide an upfront privacy notice describing that speech is transmitted off-device for ASR/avatar control. In a karaoke/chat context users may expect mic use, but they may not expect continuous remote streaming as soon as the socket opens, which creates consent and privacy risk.
The page collects an API key in the browser and submits it to backend configuration storage without any visible warning about persistence, scope, or who can later use it. Because this skill is designed for end users and the key grants access to a third-party avatar service, silent storage increases the risk of credential misuse, accidental sharing across users, or unexpected retention.
Detected: suspicious.dangerous_exec, suspicious.env_credential_access, suspicious.exposed_secret_literal