Back to skill

Security audit

IMA AI Text To Speech — seed-tts, DouBao

Security checks for vulnerabilities and agentic risk

Overview

This appears to be a real remote text-to-speech skill, but it needs review because its script can be pointed at an arbitrary API server and then send the user's API key and text there despite documentation saying the key only goes to IMA.

Review before installing. Use a scoped or test IMA_API_KEY, do not use --base-url unless you fully trust the endpoint, and avoid sending confidential scripts, personal data, or regulated content. The safest fix would be to remove or strictly allowlist the base URL override before normal use.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/ima_tts_create.py:841
Finding

Unrestricted API Base URL Allows API Key and TTS Prompt Exfiltration

Content
View full analysis

Vulnerability Details

File Location: scripts/ima_tts_create.py:66-73, 550-554, 841-864, 922
Vulnerability Type: User-controlled network destination for authenticated requests
Risk Level: High

The script permits callers to override the API base URL without validating its scheme, hostname, port, or trust relationship. The IMA API key is then attached to requests sent to that destination. This violates the documented guarantee that the credential is sent only to api.imastudio.com.

Vulnerable Code

Credential-bearing headers are constructed for all API requests:

python
def make_headers(api_key: str, language: str = "en") -> dict:
    return {
        "Authorization":  f"Bearer {api_key}",
        "Content-Type":   "application/json",
        "User-Agent":     "IMA-OpenAPI-Client/Skill-1.0.0",
        "x-app-source":   "ima_skills",
        "x_app_language": language,
    }

Task creation sends the authenticated request, including the user-provided prompt in the generated payload, to the supplied base URL:

python
url = f"{base_url}/open/v1/tasks/create"
headers = make_headers(api_key)
logger.info(f"Attempt {attempt_num}: attribute_id={attribute_id}, credit={credit}")

try:
    resp = requests.post(url, json=payload, headers=headers, timeout=30)

The command-line interface accepts an unrestricted destination:

python
p.add_argument("--base-url", default=DEFAULT_BASE_URL,
               help="API base URL")

That destination is used with the API key retrieved from either a command-line argument or the environment:

python
def main():
    args   = build_parser().parse_args()
    base   = args.base_url
    apikey = args.api_key or os.getenv("IMA_API_KEY")
    if not apikey:
        logger.error("API key is required. Use --api-key or set IMA_API_KEY environment variable")
        sys.exit(1)

    start_time = time.time()
    masked_key = f"{ap
...[truncated 3290 chars]
Remediation
View remediation

Remediation Suggestions

  1. Remove the production endpoint override. Use the fixed DEFAULT_BASE_URL for all normal Skill execution:

    python
    base = "https://api.imastudio.com"
    
  2. If endpoint overrides are required for development, require explicit opt-in and exact allowlisting. Parse the URL and validate every relevant component:

    python
    from urllib.parse import urlparse
    
    APPROVED_HOSTS = {"api.imastudio.com"}
    
    def validate_base_url(value: str) -> str:
        parsed = urlparse(value)
        if parsed.scheme != "https":
            raise ValueError("The API endpoint must use HTTPS")
        if parsed.hostname not in APPROVED_HOSTS:
            raise ValueError("Unapproved API hostname")
        if parsed.username or parsed.password:
            raise ValueError("Embedded URL credentials are prohibited")
        if parsed.port not in (None, 443):
            raise ValueError("Unexpected API port")
        return value.rstrip("/")
    
  3. Constrain redirects. Disable automatic redirects for credential-bearing requests or verify that every redirect target uses HTTPS and remains on the exact approved hostname before resending authorization headers.

  4. Separate development credentials. If testing against non-production endpoints is genuinely necessary, use a distinct option that never accepts the production IMA_API_KEY and requires a scoped test credential.

  5. Fully redact credentials from logs. The current code records the first ten characters of the API key. Replace this with a constant marker such as api_key=[REDACTED].

  6. Add security regression tests confirming rejection of:

    • Plain HTTP URLs.
    • Unapproved hostnames.
    • Lookalike and subdomain hosts.
    • URLs with embedded credentials.
    • Unexpected ports.
    • Cross-host redirects.
  7. Align documentation with enforcement. Retain the claim that credentials are sent only to `api.ima ...[truncated 64 chars]

Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Rogue AgentSelf-Modification, Session Persistence
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
Findings (27)

Tainted flow: 'task_id' from os.getenv (line 921, credential/environment) → requests.post (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.

Content

Scanner excerpt · scripts/ima_tts_create.py (reported line 357)May include surrounding context.

python
"Check the IMA dashboard for status."
            )

        resp = requests.post(url, json={"task_id": task_id},
                             headers=headers, timeout=30)
        resp.raise_for_status()
        data = resp.json()

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared purpose is a text-to-speech/voiceover capability, but the code shown does not perform any media processing, speech synthesis, caption handling, or script conversion. Instead, it is an infrastructure module for logging, including filesystem access to create directories and files in the user's home directory and deleting old log files. While logging can be a supporting detail in a larger skill, this specific code chunk's actual behavior is materially different from the declared end-user functionality and includes undeclared local file management capabilities.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The keyword list includes very broad, common phrases such as generic read-aloud, voice, and article-listening terms that can match many unrelated user requests. This can cause the skill to trigger unintentionally, increasing the chance of routing users into the wrong capability, exposing text to external TTS processing, or creating confusing behavior without clear user intent.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The prescribed acknowledgment, pre-generation, and progress messages are written as fixed Chinese responses, and the section title frames them as the required UX flow. This can violate language/locale policy if users have not chosen Chinese and no language choice mechanism is documented.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The trigger guidance for when the skill should engage is broad enough that the skill may activate on loosely related requests such as narration, dubbing, or voiceover without clear user intent. In agent systems, over-broad activation can cause accidental external transmission of user text to the third-party API when the user did not specifically consent to using this service.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · SKILL-DETAIL.md (reported line 239)May include surrounding context.

md
tree = r.json()["data"]
# ... find type=3 node, get attribute_id, model_id, model_version, credit ...

# 2. Create task
body = {
    "task_type": "text_to_speech",
    "enable_multi_model": False,

External Transmission

Medium
Category
Data Exfiltration
Confidence
70% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · SKILL-DETAIL.md (reported line 262)May include surrounding context.

md
},
    }],
}
r = requests.post(f"{BASE}/open/v1/tasks/create", headers=HEADERS, json=body)
task_id = r.json()["data"]["id"]

# 3. Poll

External Transmission

Medium
Category
Data Exfiltration
Confidence
70% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · SKILL-DETAIL.md (reported line 267)May include surrounding context.

md
# 3. Poll
while True:
    r = requests.post(f"{BASE}/open/v1/tasks/detail", headers=HEADERS, json={"task_id": task_id})
    task = r.json()["data"]
    medias = task.get("medias") or []
    if not medias:

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The document claims to be API documentation for TTS generation, but it also specifies persistent per-user storage behavior in a local preferences file. That hidden expansion of scope matters because it introduces data retention and profiling behavior that users and integrators may not expect from a simple text-to-speech skill.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

Local per-user memory storage is not clearly necessary for the stated purpose of one-shot TTS generation, so it creates unnecessary retention of user preference data. Even if the fields are not highly sensitive, storing user-linked preferences without strong justification increases privacy risk and broadens the skill's effective capability.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The documented storage of voice preferences lacks any user-facing disclosure that data is being retained locally across sessions. This creates a privacy and transparency problem because users may not realize their speaker choices are being associated with a user identifier and stored on disk.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
93% confidence
Finding

The skill declares capabilities that involve environment access, network access, and local file writes, but it does not explicitly scope or constrain those powers with a permissions or allowed-tools declaration. In an agent setting, missing tool-scope boundaries increases the chance of overbroad execution, misuse of credentials, or unintended file/network operations beyond the user's expectations.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The trigger phrase for requests like '帮我制作旁白/配音' is broad and can match loosely scoped user prompts, increasing the risk that sensitive or unintended content is sent to the remote TTS workflow without sufficiently explicit confirmation. In a conversational agent, ambiguous activation can cause privacy-impacting actions or unwanted external API usage.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill sends user-provided text to a third-party remote TTS service, but the user-facing description does not clearly warn that their content leaves the local environment. This creates a meaningful privacy and consent issue, especially if users provide confidential scripts, captions, or personal information expecting only local processing.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill specifies the header x_app_language: en, which imposes a fixed language/locale choice. The document does not offer user opt-in or explain why English is required, even though much of the skill content and trigger guidance is in Chinese.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · scripts/ima_tts_create.py (reported line 7)May include surrounding context.

python
Version: 1.0.0

Specialized script for text-to-speech generation via IMA Open API.
Handles: product list query → virtual param resolution → task create → poll status

Usage:
  python3 ima_tts_create.py \\

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The header builder defaults the x_app_language value to "en", causing requests to use English unless another language is explicitly supplied. This is a natural-language policy concern because it forces a language preference by default rather than requiring user choice or neutral behavior.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
91% confidence
Finding

The script transmits user-supplied prompt content and generated parameters to an external third-party API. In the context of an agent skill, this is a real data egress point: sensitive scripts, captions, or proprietary text entered by the user may be sent off-host without explicit per-run disclosure beyond the code itself.

Content

Scanner excerpt · scripts/ima_tts_create.py (reported line 299)May include surrounding context.

python
f"credit={model_params['credit']}, attribute_id={model_params['attribute_id']}")

    try:
        resp = requests.post(url, json=payload, headers=headers, timeout=30)
        resp.raise_for_status()
        data = resp.json()

External Transmission

Medium
Category
Data Exfiltration
Confidence
80% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · scripts/ima_tts_create.py (reported line 357)May include surrounding context.

python
"Check the IMA dashboard for status."
            )

        resp = requests.post(url, json={"task_id": task_id},
                             headers=headers, timeout=30)
        resp.raise_for_status()
        data = resp.json()

External Transmission

Medium
Category
Data Exfiltration
Confidence
92% confidence
Finding

This alternate task-creation path with reflection/retry also sends prompt-related parameters and user-selected options to the external IMA API. Because retries may include different parameter combinations and log reflection state, it increases the number of outbound transmissions of user data within one execution.

Content

Scanner excerpt · scripts/ima_tts_create.py (reported line 554)May include surrounding context.

python
logger.info(f"Attempt {attempt_num}: attribute_id={attribute_id}, credit={credit}")

        try:
            resp = requests.post(url, json=payload, headers=headers, timeout=30)
            resp.raise_for_status()
            data = resp.json()
            code = data.get("code")

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
81% confidence
Finding

The phrase "Multilingual and crosslingual reading for mixed Chinese-English scripts" describes a language behavior preference, and elsewhere the keyword list reinforces this with "中英文混读" and "explicit_language." There is no accompanying natural-language indication that language or locale choice is user-selectable or opt-in.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

The skill persists user voice preferences and logs to local files, but the description does not clearly warn users about that retention behavior. Even though the data appears low sensitivity, silent persistence can expose usage patterns, content metadata, or preference history to other local processes or users.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
96% confidence
Finding

The dependency is specified as requests>=2.25.0, which allows installation of any newer release and makes builds non-reproducible. This can unintentionally pull in a vulnerable or behavior-changing version over time, especially because the package has a history of security advisories.

Content

Scanner excerpt · requirements.txt (reported line 4)May include surrounding context.

text
# Python dependencies for ima-tts-ai skill
# Install with: pip install -r requirements.txt

requests>=2.25.0

Unverifiable Dependency: requests has 16 known advisory(ies) (CVE-2014-1830 (Exposure of Sensitive Information to an Unauthorized Actor in Requests); CVE-2024-47081 (Requests vulnerable to .netrc credentials leak via malicious URLs); CVE-2024-35195 (Requests `Session` object does not verify requests after making first request wi) +13 more), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
90% confidence
Finding

The manifest references requests without pinning an exact version, while the package has multiple known advisories across releases. Because the resolved version is unverifiable from this file alone, deployments may install an affected release, leaving the skill exposed to known dependency flaws such as credential leakage or other request-handling issues.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
92% confidence
Finding

The module docstring and inline documentation are written entirely in Chinese, which imposes a specific language/locale without any stated opt-in or explanation that the skill is region-specific. This matches the language-policy concern for natural-language content in code files.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.