Back to skill

Security audit

Doubao Image Video Skill

Security checks for vulnerabilities and agentic risk

Overview

The skill is mostly a disclosed Doubao media-generation integration, but it warrants Review because it normalizes watermark/logo removal and includes unsafe setup/install guidance.

Review before installing. Use an isolated virtual environment, avoid the npm install path, do not source ~/.basic, set only ARK_API_KEY directly or through a trusted secret manager, and only send prompts or image URLs you are allowed to share with Volcengine/Doubao. Do not use the advertised watermark/logo removal examples on third-party content unless you have explicit authorization.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T08 · Insecure Dependencies

Warning
Location
doubao-skill.json:56
Finding

Unnecessary and Unpinned Third-Party Package Installation

Content
View full analysis

Vulnerability Details

File Location: doubao-skill.json:56-57, requirements.txt:1-4, SKILL.md:338-341
Vulnerability Type: Supply-chain exposure through unnecessary and insufficiently constrained dependencies
Risk Level: Medium

Vulnerable Code

doubao-skill.json:56-57:

json
"npm": "npm install doubao-skill",
"pip": "pip install requests aiohttp",

requirements.txt:1-4:

text
requests>=2.28.0
aiohttp>=3.8.0
pydantic>=1.9.0
pytimeparse>=1.1.5

SKILL.md:338-341:

bash
pip install -r requirements.txt

# Or install manually
pip install requests aiohttp pydantic pytimeparse

Technical Analysis

The Skill is implemented in Python and its executable code only imports requests from the listed third-party packages. No executable project file imports aiohttp, pydantic, or pytimeparse. The manifest nevertheless recommends installing these packages and additionally recommends npm install doubao-skill, even though the audited implementation does not require a Node.js package.

The Python requirements use open-ended lower bounds and provide no upper bounds, exact version locks, or package hashes. Consequently, installation may resolve to package releases that did not exist when the Skill was reviewed. The unnecessary npm installation creates an additional registry trust boundary unrelated to the declared Python implementation.

This does not establish that any currently named package is malicious. The vulnerability is the avoidable supply-chain exposure: more third-party code is downloaded and potentially executed than is required for the Skill's functionality.

Attack Path

  1. An attacker compromises a referenced package, its publisher account, or the package registry distribution path, or causes an unexpected package to be resolved.
  2. A user follows the manifest or documentation and runs the npm or pip installation command.
  3. The package manager ...[truncated 1020 chars]
Remediation
View remediation

Remediation Suggestions

  1. Remove the unused npm installation entry unless a specific audited Node.js component is genuinely required.
  2. Remove aiohttp, pydantic, and pytimeparse if they remain unused.
  3. Pin every required dependency to an exact reviewed version.
  4. Generate a lock file containing cryptographic hashes, for example with pip-tools, and install using hash verification.
  5. Use an isolated virtual environment and avoid privileged package installation.
  6. Add automated dependency review and vulnerability scanning to the release process.
  7. Keep the manifest, requirements.txt, and installation documentation synchronized so that they install only the minimum required dependencies.

T09 · Insecure Skill Coding Practices

Warning
Location
SKILL.md:48
Finding

Arbitrary Shell Command Execution Through Sourcing a Broad User Configuration File

Content
View full analysis

Vulnerability Details

File Location: SKILL.md:48-51, SKILL.md:298-301, SKILL.md:645-648; duplicated in references/SKILL.md:30-33, references/SKILL.md:273-276, and references/SKILL.md:614-617
Vulnerability Type: Unsafe shell configuration loading
Risk Level: Medium

Vulnerable Code

SKILL.md:48-51:

bash
export ARK_API_KEY="your_api_key_here"

# Or source a configuration file
source ~/.basic  # if this file exists

The same unsafe command is repeated in troubleshooting and security guidance:

bash
source ~/.basic

Technical Analysis

The shell source command does not parse a configuration file as passive key-value data. It executes every shell statement in the target file within the caller's current shell process.

The Skill only needs one environment variable, ARK_API_KEY, but recommends executing an entire nonstandard user file, ~/.basic. The project does not create, validate, permission-check, or constrain this file. If its content is attacker-controlled or has been modified by another compromised application, following the setup instructions executes arbitrary commands with the user's privileges.

This behavior is not required for API authentication and therefore exceeds the minimum privilege necessary for the declared functionality.

Attack Path

  1. An attacker or compromised local process obtains write access to the user's ~/.basic file.
  2. The attacker inserts shell commands into that file, such as commands that copy credentials, modify startup files, or download and run another payload.
  3. The user follows the Skill's installation or troubleshooting documentation and runs source ~/.basic.
  4. The current shell executes every injected command without displaying a separate executable or confirmation prompt.
  5. The malicious commands run with the user's privileges and inherit the shell's environment, potentially including ARK_API_KEY and other se ...[truncated 583 chars]
Remediation
View remediation

Remediation Suggestions

  1. Remove every instruction that recommends source ~/.basic.

  2. Prefer direct, scoped assignment:

    bash
    export ARK_API_KEY="your_api_key_here"
    
  3. If persistent configuration is necessary, use a dedicated file containing only the API key and load it through a parser that treats the content as data rather than executable shell code.

  4. Require restrictive permissions on credential files, such as owner read/write only.

  5. Use the host platform's secret manager or credential store where available.

  6. Never advise users to print the complete API key for verification; verify only that the variable is set or display a masked suffix.

  7. Apply the correction to both the primary and duplicated reference documentation.

Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • System Prompt LeakageDirect Leakage, Indirect Extraction, Tool-Based Exfiltration
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (50)

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

This second mismatch finding similarly indicates the skill orchestrates asynchronous workflows and task polling beyond the narrowly stated purpose, while image editing support appears inconsistent. Such ambiguity is dangerous in agent skills because hidden or under-declared behavior can bypass user expectations and security review.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
91% confidence
Finding

This second mismatch finding similarly indicates the skill orchestrates asynchronous workflows and task polling beyond the narrowly stated purpose, while image editing support appears inconsistent. Such ambiguity is dangerous in agent skills because hidden or under-declared behavior can bypass user expectations and security review.

Content

No source excerpt is available for this finding.

Direct Prompt Extraction

High
Category
System Prompt Leakage
Confidence
85% confidence
Finding

Skill contains instructions that could directly expose system prompts, internal rules, or hidden instructions to users or external parties.

Content

Scanner excerpt · scripts/cli.py (reported line 49)May include surrounding context.

python
if len(sys.argv) < 3:
                print("错误: 缺少 prompt 参数")
                print_help()
                return
            
            prompt = sys.argv[2]
            result = await handler({

Direct Prompt Extraction

High
Category
System Prompt Leakage
Confidence
85% confidence
Finding

Skill contains instructions that could directly expose system prompts, internal rules, or hidden instructions to users or external parties.

Content

Scanner excerpt · scripts/cli.py (reported line 76)May include surrounding context.

python
if len(sys.argv) < 3:
                print("错误: 缺少 prompt 参数")
                print_help()
                return
            
            prompt = sys.argv[2]
            result = await handler({

Env Variable Harvesting

High
Category
Data Exfiltration
Confidence
92% confidence
Finding

Using os.environ.copy() forwards the entire parent process environment into the child subprocess, potentially exposing unrelated secrets such as tokens, cloud credentials, or service configuration to doubao_demo.py and any libraries it loads. In a skill context, this is more dangerous because helper scripts are an additional trust boundary; if the companion script is modified, compromised, or logs its environment, multiple secrets can leak beyond the intended ARK_API_KEY.

Content

Scanner excerpt · scripts/doubao_skill.py (reported line 125)May include surrounding context.

python
cmd = ["python3", doubao_demo] + list(args)
            
            # 设置环境变量
            env = os.environ.copy()
            env["ARK_API_KEY"] = self.api_key
            
            # 异步执行

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

This markdown file presents all user-facing documentation and usage guidance only in Chinese, with no indication that language selection is optional or that the skill is region-specific. Under the policy rule, forcing a specific language without user opt-in is a natural-language policy concern.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The README explicitly advertises image editing for '去除水印' and even provides a watermark-removal example prompt, but includes no warning about authorization, copyright, ownership, or lawful use. This normalizes and facilitates a misuse-prone capability that can be used to strip attribution or rights-management markings from third-party content, increasing legal and abuse risk in the skill context.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
90% confidence
Finding

The skill documentation describes use of environment variables, network access, and shell commands, but the skill declares no explicit tool scope or permissions. This weakens review and sandboxing because operators cannot easily see that the skill needs sensitive capabilities such as outbound network access and access to secrets in the environment.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill documents watermark-removal functionality without any warning about legal, copyright, authorization, or terms-of-service risks. In context, this omission makes misuse more likely because the feature is presented as routine and frictionless rather than restricted to authorized use cases.

Content

No source excerpt is available for this finding.

Ssd 4

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

By repeatedly presenting watermark and logo removal as normal usage, the skill encourages misuse toward copyright-protection circumvention. The risk is elevated by the cumulative framing across examples, prompts, and troubleshooting guidance, which lowers user hesitation and makes abuse operationally easy.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

The manifest describes a skill for text-to-image, image editing, and text-to-video using the Doubao API. However, the documented usage directly invokes a local Python script via subprocess, introducing process-spawning capability that is unrelated to the core media-generation purpose and broader than what the skill description justifies.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The documentation normalizes watermark, logo, and text removal as a supported feature and expands it into broader content-altering regeneration. This materially increases abuse risk by facilitating removal of ownership markers and copyright-management information from third-party images.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The manifest advertises text-to-image, image editing, and text-to-video features using prompts and image URLs, and it requires an external ARK API key, but it does not clearly warn users that their prompts, uploaded/referenced image URLs, and related task data will be transmitted to a third-party service. This creates a privacy and data-handling risk because users may unknowingly send sensitive text or image content to an external provider.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The documentation presents watermark removal as normal usage without any warning about copyright, ownership, authorization, or legal restrictions. This omission increases the risk that users will treat potentially unlawful or abusive behavior as an endorsed feature of the skill.

Content

No source excerpt is available for this finding.

Ssd 4

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The workflow is framed around removing watermarks while keeping the underlying content intact, which operationally teaches users how to preserve value while stripping ownership cues. That makes the skill more dangerous in context because it is an image-editing tool and the documentation normalizes a misuse pattern directly tied to copyright evasion and content laundering.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The guide explicitly documents watermark removal as a supported capability and even suggests preserving the main content while removing attribution marks. That facilitates potentially infringing or deceptive use by helping users strip ownership or licensing indicators from images, which goes beyond neutral image editing guidance.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
78% confidence
Finding

The README is written entirely in Chinese and the examples and interface descriptions assume Chinese-language interaction, with no indication that users may choose another language or that the skill is intentionally limited to a Chinese-speaking or region-specific audience. This can violate the language/locale policy when a specific language is imposed without opt-in or justification.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The README explicitly advertises image editing for watermark removal, which expands the skill from generic editing into a misuse-prone capability associated with copyright and attribution bypass. In a generative media skill, documenting this use case normalizes abusive behavior and can encourage users to remove ownership markers from third-party content.

Content

No source excerpt is available for this finding.

Ssd 2

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The documentation includes direct instructions for removing watermarks from images, which is a misuse-enabling pattern with clear abuse potential. In the context of an image editing skill, this materially increases the likelihood of copyright infringement, provenance stripping, and circumvention of attribution or licensing controls.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The custom example prompt promotes removing logos while preserving the subject, which is a concrete misuse pattern for stripping brand or ownership indicators from images. Providing ready-to-use prompts lowers the barrier for policy-violating or infringing use and makes the skill more dangerous than a neutral editing interface.

Content

No source excerpt is available for this finding.

Ssd 2

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The example prompt to remove logos while preserving the subject is a paraphrased but unmistakable instruction for provenance or branding removal. Because it is presented as a copyable usage example, it serves as operational guidance for misuse rather than neutral documentation.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The documentation explicitly presents image editing for removing watermarks, which enables users to strip attribution or ownership markers from third-party content. In the context of an image-generation/editing skill, this meaningfully expands the tool toward misuse that can facilitate copyright infringement, fraud, or unauthorized redistribution.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The file documents watermark/logo removal without any caution about legal, copyright, trademark, or authorization risks. This omission is dangerous because it frames a potentially unlawful or policy-violating workflow as routine and acceptable, increasing the chance of misuse by untrained users.

Content

No source excerpt is available for this finding.

Ssd 4

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The guide does more than mention watermark removal; it operationalizes it with concrete commands and examples. That turns a potentially harmful manipulation technique into an easy, repeatable workflow and increases misuse potential in a way that is more dangerous than abstract discussion alone.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The API reference formalizes AI-based watermark removal as a supported capability, not merely an incidental example. That makes misuse easier to operationalize and signals endorsement of a high-risk content manipulation use case that can bypass attribution and rights-management signals.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.