Back to skill

Security audit

Flowyaipc Herdsman Skill En

Security checks for vulnerabilities and agentic risk

Overview

The skill is a mostly disclosed Herdsman integration, but its media download helper can leak API credentials to untrusted URLs and writes files with little containment.

Review before installing or exposing this to agents with API keys. Use it only with trusted Herdsman endpoints, avoid auto-downloading media from untrusted services, keep outputs in a dedicated directory, and do not submit sensitive images, audio, voices, or business documents unless you intend them to be processed by the configured service.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/herdsman_client.py:238
Finding

Bearer Credential Disclosure Through Unrestricted Cross-Origin Media Downloads

Content
View full analysis
str: request = Request(url, headers=self._headers()) actual_timeout = timeout if timeout is not None else self.timeout try: with urlopen(request, timeout=actual_timeout) as response: data = response.read() except HTTPError as exc: raw = exc.read().decode("utf-8", errors="replace") raise HerdsmanAPIError(raw or f"http error {exc.code}", status_code=exc.code, body=raw) from exc except URLError as exc: raise HerdsmanAPIError(f"download failed: {exc}") from exc return write_bytes(output_path, data) ``` The authorization header used by this function is constructed as follows: ```python def _headers(self, extra_headers: Optional[Dict[str, str]] = None) -> Dict[str, str]: headers: Dict[str, str] = {} if self.api_key: headers["Authorization"] = "Bearer " + self.api_key if extra_headers: headers.update(extra_headers) return headers ``` A representative caller trusts a URL contained in the model-service response: ```python url = item.get("url", "") if not url: print(f"Image {index + 1} returned empty") continue if target_path and (args.download or args.auto_save): try: saved = client.download_to_file(url, target_path, timeout=120) print(f"Image {index + 1} downloaded: {saved}") except HerdsmanAPIError as exc: print(json.dumps(exc.to_dict(), indent=2, ensure_ascii=False), file=sys.stderr) ...[truncated 2853 chars]
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (26)

Tp4

High
Category
MCP Tool Poisoning
Confidence
94% confidence
Finding

The declared description presents the package as a general integration layer/protocol connector for model services, including OpenAI, Anthropic, or AGUI-compatible connections. The supplied code chunk instead implements a specific executable TTS utility for Herdsman audio synthesis. Its primary purpose is generating speech from text, optionally using voice cloning inputs and downloading resulting audio files to disk. Those are concrete end-user media-generation capabilities not reflected in the high-level integration description. While being under a Herdsman integration package could be related context, this script's actual behavior is materially more specific and operational than the declared purpose.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description emphasizes integration/protocol support for connecting agent platforms to model service providers (OpenAI, Anthropic, AGUI-compatible services). The supplied code does not implement provider connectivity, protocol handling, or integration logic. Instead, it performs local media preprocessing: validating file paths, creating output directories, invoking ffmpeg as a subprocess, and writing converted WAV files. This is a materially different primary purpose and includes undeclared capabilities related to audio transformation, filesystem operations, and subprocess execution.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The declared description presents the package as a general integration/protocol package for connecting agent platforms to model services. The supplied code chunk instead implements a specific operational script for image editing via the Herdsman API. That is a materially more specific and active capability than a passive integration/protocol description. It also handles local file input, local file output, and optional downloading of remote image results, none of which are reflected in the declared purpose or permissions. While it does relate to Herdsman, the actual behavior is not accurately represented by the high-level description alone.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The declared description presents the package as an integration layer/protocol support component for connecting agent platforms to model services. The supplied code chunk instead implements an end-user/CLI image generation script with concrete content-generation behavior, local filesystem writes, and optional network downloading of generated images. Those are materially more specific and operational capabilities than the declared purpose, and they are not clearly represented by the description.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
94% confidence
Finding

The declared description presents the package as a general integration layer and protocol-spec bundle for connecting agent platforms to OpenAI-, Anthropic-, or AGUI-compatible services. The supplied code chunk instead implements a concrete OCR utility: it accepts a local image path, verifies file existence, reads the image, converts it to a data URL, calls a Herdsman OCR API, and prints or saves recognition results. That is a materially different primary purpose and introduces undeclared capabilities involving local file input/output and OCR-specific processing. While it does use Herdsman infrastructure, the behavior shown is not accurately represented by the declared description.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The declared description frames the package as a general integration layer and protocol-spec bundle for connecting agent platforms to OpenAI, Anthropic, or AGUI-compatible services via Herdsman. The supplied code chunk instead implements a specific executable utility for transcribing audio through a Herdsman endpoint. It processes user-provided media inputs, including local files and remote URLs, and performs a concrete transcription task. That is a materially more specific and different behavior than the declared high-level integration/connectivity purpose, and it introduces media-processing capability not reflected in the description.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
93% confidence
Finding

The declared description presents this as a general integration package and protocol support layer for connecting agent platforms to model services. The supplied code instead implements a concrete standalone transcription script with a narrow purpose: ingesting local audio, calling a specific local OpenAI-compatible ASR endpoint, and optionally saving outputs to disk. While this still relates to Herdsman, the primary behavior is materially more specific and operational than the declared package-level integration description. The local file read/write behavior and the dedicated transcription capability are not represented in the declared purpose.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
92% confidence
Finding

The declared description presents the skill as a general integration package for Herdsman and protocol specs for connecting to multiple provider types. The supplied code chunk instead implements a specific end-user utility: voice-clone TTS synthesis against a local Herdsman service using the qwen3-tts-voiceclone model. It reads a reference audio file from disk, sends content to a local HTTP endpoint, and writes synthesized audio to an output file. Those concrete capabilities and the script’s primary purpose are materially narrower and different from the declared generic integration/package description, so this is a description-behavior mismatch.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
90% confidence
Finding

The skill documents scripts and workflows that clearly rely on shell, filesystem, network, and likely environment access, but the manifest does not declare any tool scope such as permissions or allowed-tools. That omission makes the package harder for agent platforms to sandbox correctly and increases the chance an agent grants broader execution than intended.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The skill encourages OCR, speech transcription, and voice cloning workflows over user-supplied images and audio without warning that these inputs may contain sensitive biometric, personal, or confidential information. In agent environments, missing privacy guidance increases the chance users process or persist sensitive media without consent, minimization, or retention controls.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The examples instruct users to submit local images and audio for OCR, image editing, transcription, and voice cloning, which can expose sensitive personal or business content to the service. The document lacks any explicit warning that local media contents are sent to an API endpoint and may contain sensitive data.

Content

No source excerpt is available for this finding.

Unbounded Resource Access

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill allows unbounded resource consumption (API calls, storage, compute). Without rate limits or quotas, a compromised or misbehaving agent can cause denial-of-service or cost overruns.

Content

Scanner excerpt · references/error-codes.md (reported line 174)May include surrounding context.

}

text

Handling: Check VRAM, memory, and runtime status first; do not retry indefinitely.

### Language not supported

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The script sets --language to default to Chinese, which imposes a specific language choice even when the user does not explicitly request it. This is a natural-language policy concern because the tool does not offer a neutral default or require opt-in before applying a locale-specific setting.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
86% confidence
Finding

When --download is used in streaming mode, the script automatically fetches a stream_url supplied by the API/server. Because base_url is user-configurable and the server controls stream_url, this can be abused to trigger unintended outbound requests or retrieval of attacker-influenced content, which is risky in an integration component expected to mediate remote model services.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/audio_speech.py (reported line 108)May include surrounding context.

python
print(f"Stream URL: {full_url}")
            if args.download:
                import subprocess
                subprocess.run(["curl", "-o", auto_output_path(), full_url], check=False)
        else:
            print(json.dumps(result, indent=2, ensure_ascii=False))
        return

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/convert_audio.py (reported line 40)May include surrounding context.

python
# Check if ffmpeg is available
    try:
        subprocess.run(["ffmpeg", "-version"], capture_output=True, check=True)
    except (subprocess.CalledProcessError, FileNotFoundError):
        print("Error: ffmpeg not found, please install it and add to PATH", file=sys.stderr)
        sys.exit(1)

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/convert_audio.py (reported line 54)May include surrounding context.

python
output_path
    ]

    result = subprocess.run(cmd, capture_output=True, text=True)

    if result.returncode != 0:
        print(f"ffmpeg conversion failed:\n{result.stderr}", file=sys.stderr)

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The helper reads arbitrary local files and converts them to a data URL, which is then used by methods like OCR and media submission to send file contents to remote API endpoints. Although the behavior is functional, this file does not include any confirmation prompt, user-visible logging, or explicit warning that local file data may be uploaded.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The download_to_file method fetches attacker-controlled or caller-supplied URLs and writes the response to an arbitrary caller-specified path with no path restrictions, overwrite controls, size checks, or scheme validation. In a larger agent context, untrusted inputs could cause arbitrary file overwrite, persistence in sensitive locations, or SSRF-style retrieval of internal resources followed by local storage.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The payload hard-codes "language": "Chinese", which imposes a specific language/locale behavior on every invocation. This matches the policy-violation category because the user is not given any language choice or opt-in, and the file does not document a justified region-specific constraint.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/tts_voice_clone.py (reported line 158)May include surrounding context.

python
# Try to get duration info via ffprobe (if available)
    try:
        import subprocess
        probe = subprocess.run(
            ["ffprobe", "-v", "quiet", "-show_entries",
             "format=duration,bit_rate", "-of", "csv=p=0", saved],
            capture_output=True, text=True, timeout=10

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
83% confidence
Finding

The skill explicitly instructs users to save generated images and reused outputs to disk but does not mention that saved files may persist sensitive or proprietary data locally. This can lead to unintended retention, later disclosure, or cross-session leakage on shared systems.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
78% confidence
Finding

The natural-language recommendations favor a Chinese-language model and note differing Chinese script outputs, but the surrounding guidance does not frame this as user choice or a region-specific requirement. This can be interpreted as a locale/language preference embedded in the skill guidance without explicit opt-in.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
81% confidence
Finding

This markdown file documents commands that auto-save generated images, edit local images, and write OCR results to disk, but it does not include any warning that these operations create or overwrite local files. Under the markdown-specific warning rule, behaviors affecting user data or system state should be disclosed so users understand the impact before running the examples.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

The save_base64_file method decodes arbitrary base64 content and writes it directly to any caller-provided path, enabling arbitrary file creation or overwrite when used with untrusted input. In an agent integration package, this becomes more dangerous because upstream model/tool outputs may be attacker-influenced and could be used to plant files in sensitive or executable locations.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.prompt_injection_instructions

Prompt-injection style instruction pattern detected.

Warn
Code
suspicious.prompt_injection_instructions
Location
references/api-examples.md:44