T03 · Remote Payload Retrieval and Execution
- Location
SKILL.md:81- Finding
Execution of an Unrestricted Remote-Supplied Upgrade Command
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This video-analysis skill has review-worthy risk because it asks for cloud credentials in chat, exposes secrets in output/logs, and can run a server-supplied update command.
Install only if you are comfortable sending videos to the external Volcengine/Ark service and can provide tightly scoped, disposable credentials through a secure mechanism. Do not paste long-lived AK/SK values into chat, rotate any keys previously used with this workflow, and avoid approving remote update commands unless the publisher provides a verified package source and integrity checks.
SKILL.md:81Execution of an Unrestricted Remote-Supplied Upgrade Command
scripts/core/api/iccp/client.py:91Authentication Tokens and Request Data Persisted in Plaintext Logs
SKILL.md:41Long-Lived Cloud Credentials Exposed Through Chat and Shell Output
scripts/core/api/meida/chunks.py:234Account-Wide IAM Enumeration and Administrator Ownership for Media Uploads
SKILL.md:121Server-Side Request Forgery Risk in User-Controlled Video Download
scripts/requirements.txt:1Mandatory Installation of Unpinned Dependencies Without Integrity Verification
The workflow centers on uploading local media to an external service under the banner of video parsing, but the description does not make that trust boundary explicit. This is dangerous because local files may contain sensitive content, and users may not understand that data leaves the local environment and is processed remotely.
The workflow centers on uploading local media to an external service under the banner of video parsing, but the description does not make that trust boundary explicit. This is dangerous because local files may contain sensitive content, and users may not understand that data leaves the local environment and is processed remotely.
The workflow centers on uploading local media to an external service under the banner of video parsing, but the description does not make that trust boundary explicit. This is dangerous because local files may contain sensitive content, and users may not understand that data leaves the local environment and is processed remotely.
The workflow centers on uploading local media to an external service under the banner of video parsing, but the description does not make that trust boundary explicit. This is dangerous because local files may contain sensitive content, and users may not understand that data leaves the local environment and is processed remotely.
The workflow centers on uploading local media to an external service under the banner of video parsing, but the description does not make that trust boundary explicit. This is dangerous because local files may contain sensitive content, and users may not understand that data leaves the local environment and is processed remotely.
The workflow centers on uploading local media to an external service under the banner of video parsing, but the description does not make that trust boundary explicit. This is dangerous because local files may contain sensitive content, and users may not understand that data leaves the local environment and is processed remotely.
The workflow centers on uploading local media to an external service under the banner of video parsing, but the description does not make that trust boundary explicit. This is dangerous because local files may contain sensitive content, and users may not understand that data leaves the local environment and is processed remotely.
The workflow centers on uploading local media to an external service under the banner of video parsing, but the description does not make that trust boundary explicit. This is dangerous because local files may contain sensitive content, and users may not understand that data leaves the local environment and is processed remotely.
The workflow centers on uploading local media to an external service under the banner of video parsing, but the description does not make that trust boundary explicit. This is dangerous because local files may contain sensitive content, and users may not understand that data leaves the local environment and is processed remotely.
The workflow centers on uploading local media to an external service under the banner of video parsing, but the description does not make that trust boundary explicit. This is dangerous because local files may contain sensitive content, and users may not understand that data leaves the local environment and is processed remotely.
The workflow centers on uploading local media to an external service under the banner of video parsing, but the description does not make that trust boundary explicit. This is dangerous because local files may contain sensitive content, and users may not understand that data leaves the local environment and is processed remotely.
The workflow centers on uploading local media to an external service under the banner of video parsing, but the description does not make that trust boundary explicit. This is dangerous because local files may contain sensitive content, and users may not understand that data leaves the local environment and is processed remotely.
The workflow centers on uploading local media to an external service under the banner of video parsing, but the description does not make that trust boundary explicit. This is dangerous because local files may contain sensitive content, and users may not understand that data leaves the local environment and is processed remotely.
The authentication check command explicitly echoes API keys and secret access keys from environment variables. Printing secrets to stdout risks exposure in logs, transcripts, debugging output, and to any downstream observer, creating immediate compromise of cloud credentials.
The skill explicitly instructs the agent to ask users to paste cloud AK/SK credentials into chat and then export them for subsequent commands. This is a direct secret-handling anti-pattern: chat is not an appropriate credential channel, the values may be logged or exposed, and the request is not justified by the stated task of video analysis.
The skill tells the agent to collect AK/SK secrets directly in chat and reuse them within the session. This enables credential theft, accidental logging, downstream leakage to tools or transcripts, and unauthorized access to the user's cloud account; in the context of a video-analysis skill, it is severely disproportionate and highly suspicious.
The instructions solicit highly sensitive AK/SK credentials and provide no safety warning or secure handling guidance. In this context, the absence of warning is especially dangerous because the skill normalizes secret disclosure for a routine media task, increasing phishing-like risk and credential compromise.
Referenced artifact was not completely inspected
python3.12 -m pip install -r ./scripts/requirements.txt
The code logs the full request URL, headers, and body immediately before the outbound call. For the AK/SK flow, the Authorization signature and potentially sensitive video-analysis payloads are written to logs, which can expose secrets and user data to anyone with log access.
The V2 client logs request headers containing a bearer token in the Authorization header. A leaked bearer token can typically be replayed directly to access the remote API, making this a direct credential-exposure issue.
This code performs remote upload of local media to a Muse service, including hashing, state negotiation, material creation, and polling for media info, which goes beyond a skill described as local/video analysis and metadata extraction. In a media-analysis skill context, silently transmitting user-provided local files to an external service creates a significant data exfiltration and privacy risk, especially for sensitive videos.
The skill metadata claims video analysis functionality, but the actual code performs an unrelated network operation to register an 'Ark Claw' package. This deceptive mismatch is dangerous because a user invoking a media-processing skill would instead trigger an external service action they did not request, which could enroll accounts, consume quotas, or grant third-party benefits without informed consent.
The skill declares no explicit tool scope even though it clearly uses environment access, local file I/O, and network operations. Missing scope declarations weaken review and runtime containment, making it easier for a broadly triggered skill to access secrets and exfiltrate user-provided media without clear user-facing constraints.
The trigger includes vague 'equivalent intent' language, which can cause the skill to activate on loosely related requests. For a skill that accesses files, networks, and credentials, overbroad triggering materially increases the chance of unintended invocation and sensitive action in the wrong context.
The skill description, prompts, and output templates are entirely in Chinese and direct the agent to use fixed Chinese user-facing messages. There is no opt-in, language selection, or documented reason that the skill must operate only in Chinese, which creates a locale policy concern.
No suspicious patterns detected.