Back to skill

Security audit

Image-crawler

Security checks for vulnerabilities and agentic risk

Overview

The skill does what it says, but its bulk image scraping and downloading lacks enough safety limits and compliance guardrails for automatic use.

Review before installing. Use only for images you have rights to collect, keep counts small, write into a dedicated quota-limited folder, and treat downloaded files as untrusted until validated. Do not rely on the anti-scraping tips to bypass site rules or terms of service.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/image_crawler.py:190
Finding

Unbounded and Insufficiently Validated Remote Image Downloads

Content
View full analysis
Remediation
View remediation
MAX_IMAGE_BYTES: return False except ValueError: return False written = 0 with open(temp_path, "wb") as f: for chunk in resp.iter_content(8192): if not chunk: continue written += len(chunk) if written > MAX_IMAGE_BYTES: raise ValueError("Image exceeds maximum permitted size") f.write(chunk) ``` 2. **Require explicit image MIME types** - Apply the same validation to Baidu and Bing. - Allow only necessary values such as `image/jpeg`, `image/png`, and `image/webp`. - Do not treat `application/octet-stream` as sufficient proof that the response is an image. 3. **Validate decoded image content** - Use a maintained image-processing library to parse and verify the completed file. - Check the actual decoded format, dimensions, and pixel count. - Configure a maximum pixel count to prevent decompression-bomb attacks. - Do not rely on URL extensions or server-controlled HTTP headers. 4. **Use temporary files and atomic publication** - Write each response to a securely created temporary file inside the output directory. - Delete it on every rejection or exception. - Atomically rename it to the final generated filename only after size, MIME, and image-structure checks succeed. 5. **Apply task-level resource controls** - Limit total downloaded bytes, candidate URLs, redirects, and execution time for each crawler run. - Consider filesystem quotas or process ...[truncated 315 chars]
Vulnerability Patterns
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (9)

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The code’s primary purpose broadly matches image crawling/downloading, but the declared description materially overstates implemented capabilities. The provided chunk only covers a simplified Baidu crawler/downloader. There is no evidence of Bing integration, keyword expansion, duplicate detection, persistent dedup state, or stall monitoring logic. Simple loop output during downloads does not satisfy the declared advanced progress/stall features. Therefore the description does not accurately represent this code chunk.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

该代码与“图片采集/爬虫工具”的大方向一致,确实执行图片搜索和下载,但与声明相比缺少多项核心能力。最明显的是搜索引擎范围不符:声明写明支持百度和 Bing,而此代码块只实现 Bing。其次,声明中的高级功能——关键词拓展、内容 hash 去重、跨次运行持久化、停滞检测——在代码中均未体现。代码只有简单的 URL 级去重,且仅在内存中对当前 search 结果生效;命令行输出进度也不足以等同于“进度监控和停滞检测”。因此描述对该代码块的能力有明显夸大,属于描述与实际行为不匹配。

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
86% confidence
Finding

The skill instructs the agent to invoke a Python crawler, write downloaded files to disk, and access external search engines, but it declares no explicit tool scope or permissions boundaries. That creates an over-privileged and ambiguous execution model where the agent may use file and network capabilities without clear restriction, increasing the chance of misuse or unintended access.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The skill explicitly facilitates bulk scraping and downloading of third-party images to a user-chosen directory without any guidance on copyright, privacy, or site terms. In this context, the omission materially increases legal/compliance and privacy risk, especially because the workflow encourages automated large-scale collection from search engines.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The documentation explicitly advises how to bypass or mitigate anti-scraping defenses by slowing requests, changing the User-Agent, and spacing batches, but provides no warning about legal, contractual, or account-abuse risks. In an image-crawling skill, this guidance materially enables policy evasion and misuse against third-party services rather than merely explaining benign configuration.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This code forces Baidu requests to use the zh-CN locale via the Accept-Language header, which is a natural-language locale choice embedded in code rather than selected by the user. The policy allows locale constraints only when user choice or clear justification is provided, neither of which appears here.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The Bing client sets Accept-Language to prefer zh-CN and zh ahead of English, imposing a language/locale preference in the skill's behavior. Because the user is not offered a locale choice and no justification is documented, this is a natural-language policy violation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
91% confidence
Finding

The request headers hard-code Accept-Language to zh-CN,zh;q=0.9, which imposes a specific language/locale behavior. The file does not offer a user opt-in or configuration for locale, and there is no documented justification that this skill must be China-specific for compliance or other policy reasons.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
95% confidence
Finding

The request header sets Accept-Language to "zh-CN,zh;q=0.9,en;q=0.8", which imposes a specific language/locale preference on all requests. The file does not offer a user opt-in or explain why a Chinese locale is required, so this is a natural-language locale policy concern.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.