Back to skill

Security audit

04 Text To Image

Security checks for vulnerabilities and agentic risk

Overview

This skill does what it says: it sends a text prompt to a configured image-generation API and returns an image URL.

Install only if you trust the configured image-generation provider. Use an HTTPS API_BASE that you control or trust, use a limited-scope API key with quotas, and avoid submitting confidential prompts unless that provider is approved for such data.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Warning
Location
skill.js:4
Finding

Unrestricted Network Destination Receives API Credentials and User Prompts

Content
View full analysis

Vulnerability Details

File Location: skill.js, lines 4–18
Vulnerability Type: Unvalidated credential-bearing outbound request
Risk Level: Medium

Vulnerable Code

js
const API_KEY = env.API_KEY;
const API_BASE = env.API_BASE;
const MODEL_NAME = env.MODEL_NAME;

const res = await fetch(`${API_BASE}/txt2img`, {
  method: "POST",
  headers: {
    "Authorization": "Bearer " + API_KEY,
    "Content-Type": "application/json"
  },
  body: JSON.stringify({
    model: MODEL_NAME,
    prompt: prompt,
    negative_prompt: negative_prompt,
    ratio: "9:16"
  })
});

Technical Analysis

The destination of the outbound request is constructed directly from the environment-controlled API_BASE value without validating its scheme, hostname, port, resolved address, or trust relationship. The request sends the API_KEY as a bearer credential and transmits the user-supplied prompt and negative_prompt.

External network access and transmission of prompts are necessary for the declared remote text-to-image functionality. However, allowing an unrestricted destination is broader than the minimum privilege required. If an attacker can influence runtime configuration, the request can be redirected to an attacker-controlled endpoint or potentially an internal network service. The implementation also does not explicitly require HTTPS or define a restrictive redirect policy, so credential confidentiality depends entirely on external configuration and runtime behavior.

No evidence was found that the Skill intentionally harvests credentials or sends them to a hidden, hard-coded destination.

Attack Path

  1. An attacker gains the ability to modify or influence the Skill's API_BASE environment value.
  2. The attacker sets API_BASE to a server under their control or to a reachable internal endpoint.
  3. A user invokes the Skill with a text-to-image prompt.
  4. The Skill sends a POST request ...[truncated 1099 chars]
Remediation
View remediation

Remediation Suggestions

  1. Replace unrestricted API_BASE configuration with a fixed trusted provider endpoint where operationally possible.
  2. If configurability is required, parse the URL with a standards-compliant URL parser and enforce an explicit allowlist of trusted hostnames and ports.
  3. Require the https: scheme and reject plaintext HTTP, embedded URL credentials, unexpected ports, fragments, and malformed URLs.
  4. Prevent SSRF by rejecting loopback, private, link-local, multicast, and other non-public resolved addresses unless a specifically approved private provider is required. Account for DNS rebinding by validating resolved addresses at connection time.
  5. Disable redirects for credential-bearing requests, or validate every redirect destination against the same scheme, hostname, port, and address restrictions before forwarding the Authorization header.
  6. Use a provider-scoped, least-privilege API key with quota and billing limits, and establish regular key rotation and revocation procedures.
  7. Validate that API_KEY, API_BASE, and MODEL_NAME are present and valid before issuing a request.
  8. Clearly document that prompts and negative prompts are transmitted to an external image-generation provider so users can avoid submitting confidential content.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
Findings (4)

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill explicitly states that use requires configuring API_KEY, API_BASE, and MODEL_NAME, which indicates prompts and likely generated content are sent to an external model service, but it provides no user warning about that data flow. This can cause unintended disclosure of sensitive prompts, proprietary story material, or personal data to third-party infrastructure, especially in creative workflows where users may paste confidential content.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The code makes an HTTP POST request to an external /txt2img endpoint and includes user-provided prompt/negative_prompt data along with an authorization bearer token. There is no confirmation prompt, logging, comment, or docstring warning the user that their input will be transmitted to a remote service.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
81% confidence
Finding

All user-facing instructions and descriptions in the file are in Chinese, and the skill does not indicate that this locale is optional or limited to a justified region-specific use case. This can be a natural-language policy concern when a skill effectively forces a specific language without offering user choice or documenting the constraint.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
95% confidence
Finding

The manifest description is written only in Chinese ("根据提示词生成短剧画面"), which indicates a fixed language/locale presentation without offering a user choice or documenting a region-specific need. Under the policy, language constraints should be optional or explicitly justified.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.