Back to skill

Security audit

vaDocparse

Security checks for vulnerabilities and agentic risk

Overview

This is a real remote document parser, but it needs Review because it can upload readable local documents over HTTP and stores service credentials in local configuration.

Install only if you trust the remote parsing service and its network path. Treat every parsed file as uploaded to that service, avoid sensitive documents unless HTTPS and retention policies are clear, restrict the service to approved document directories, and avoid putting real API keys in command lines, dry-run logs, .env files, or broadly readable OpenClaw config.

Vulnerability Patterns
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
Findings (4)

T09 · Insecure Skill Coding Practices

Error
Location
mcp/docparse.py:305
Finding

Document Contents and API Credentials May Be Transmitted over Plaintext HTTP

Content
View full analysis
Remediation
View remediation

T05 · Unauthorized Access and Privilege Escalation

Error
Location
mcp/docparse.py:174
Finding

Unrestricted Readable File Paths Can Be Uploaded to the Remote Parsing Service

Content
View full analysis
tuple[bool, dict | str]: """Validate input file without exiting. Returns (ok, resolved_path_or_error_dict). """ path = Path(file_path) if not path.exists(): return False, { "code": "F001", "message": "解析失败:未找到待解析文件。建议:请确认文件是否已上传,或检查文件路径是否正确。", } if not path.is_file() or not os.access(path, os.R_OK): return False, { "code": "F002", "message": "解析失败:当前文件无法读取。建议:请检查文件权限后重试。", } ext = path.suffix.lower() if ext not in SUPPORTED_EXT: mime, _ = mimetypes.guess_type(str(path)) if mime not in SUPPORTED_MIME: return False, { "code": "F003", "message": "解析失败:当前文件格式不支持文档解析。支持格式:PDF、PNG、JPG、JPEG、BMP、WEBP、TIFF。建议:请先将文件转换为 PDF 后重试。", } return True, str(path.resolve()) ``` ```python # ---- input validation ---- ok, check_result = _check_file(input_file) if not ok: return {"success": False, "error": check_result} input_path = check_result # ---- read file ---- try: with open(input_path, "rb") as f: file_b64 = base64.b64encode(f.read()).decode("utf-8") except OSError: return { "success": False, "error": { "code": "F002", "message": "解析失败:当前文件无法读取。建议:请检查文件权限后重试。", }, } # ---- call MCP ---- try: client_kwargs = {"timeout": timeout} if api_key: client_kwargs["auth"] = api_key client = Client(mcp_url, **client_kwargs) async with client: result = await client.call_tool( "docparse", { "file_content_base64": file_b64, "file_name": Path(input_path).name, "output_format" ...[truncated 2146 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
setup.py:91
Finding

Installer Persists API Credentials in Plaintext and Exposes Them in Dry-Run Output

Content
View full analysis
Remediation
View remediation

T08 · Insecure Dependencies

Warning
Location
requirements.txt:1
Finding

Unpinned Dependencies Are Installed without Artifact Integrity Verification

Content
View full analysis
=3.0.0 mcp>=1.0.0 ``` ```python def install_dependencies(skill_dir: Path, dry_run: bool): """Install Python dependencies.""" req_file = skill_dir / "requirements.txt" try: import fastmcp import mcp return except ImportError: pass if not req_file.is_file(): packages = ["fastmcp>=3.0.0", "mcp>=1.0.0"] else: packages = None if packages: cmd = [sys.executable, "-m", "pip", "install"] + packages elif req_file.is_file(): cmd = [ sys.executable, "-m", "pip", "install", "-r", str(req_file), ] result = subprocess.run( cmd, capture_output=True, text=True, timeout=300, ) ``` The installation documentation also recommends: ```bash pip install -r requirements.txt -i https://pypi.tuna.tsinghua.edu.cn/simple ``` ### Technical Analysis Both dependencies use open-ended lower-bound constraints rather than reviewed exact versions. A future release satisfying `>=3.0.0` or `>=1.0.0` can therefore be selected automatically. No hashes are supplied to authenticate the expected distributions, and transitive dependencies are not locked. The documentation additionally recommends a third-party package mirror. A mirror is not inherently malicious, but using an additional package-distribution trust boundary without hashes means a compromised mirror, compromised upstream project, or malicious future release could supply altered code. Python package installation can execute build-system code, and installed package code is later imported and executed by this Skill. ### Attack Path 1. An attacker compromises an allowed package re ...[truncated 1126 chars]
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (41)

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The skill is presented as simple document parsing, but its instructions extend into package installation, config discovery, environment inspection, and modification of local/OpenClaw configuration. That mismatch can cause users or orchestrators to grant broader trust than warranted, enabling stealthy persistence, secret collection, or system reconfiguration beyond OCR functionality.

Content

No source excerpt is available for this finding.

Missing User Warnings

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The skill sends user document contents to a remote parsing service but does not present a clear user-facing warning at the trigger/description level. This is dangerous because documents often contain confidential data, and users may assume local processing unless remote transfer is prominently disclosed before use.

Content

No source excerpt is available for this finding.

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · mcp/docparse.py (reported line 87)May include surrounding context.

python
Returns a dict of key=value pairs, ignoring comments and blank lines.
    """
    skill_dir = _SCRIPT_DIR.parent  # mcp/ to docparse/
    env_path = skill_dir / ".env"
    env_vars = {}

    if env_path.is_file():

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · setup.py (reported line 98)May include surrounding context.

python
Returns a dict of key=value pairs, ignoring comments and blank lines.
    """
    skill_dir = _SCRIPT_DIR.parent  # mcp/ to docparse/
    env_path = skill_dir / ".env"
    env_vars = {}

    if env_path.is_file():

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · mcp/docparse.py (reported line 16)May include surrounding context.

python
功能:
  1. 自动识别 OpenClaw 安装路径与 workspace 位置
  2. 安装 Python 依赖(fastmcp, mcp)
  3. 创建 .env 配置文件(如不存在)
  4. 将 MCP Server 配置注入 openclaw.json
  5. 检测并修复 docparse.py 中硬编码的路径,使其自适应

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · mcp/docparse.py (reported line 44)May include surrounding context.

python
功能:
  1. 自动识别 OpenClaw 安装路径与 workspace 位置
  2. 安装 Python 依赖(fastmcp, mcp)
  3. 创建 .env 配置文件(如不存在)
  4. 将 MCP Server 配置注入 openclaw.json
  5. 检测并修复 docparse.py 中硬编码的路径,使其自适应

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · mcp/docparse.py (reported line 79)May include surrounding context.

python
功能:
  1. 自动识别 OpenClaw 安装路径与 workspace 位置
  2. 安装 Python 依赖(fastmcp, mcp)
  3. 创建 .env 配置文件(如不存在)
  4. 将 MCP Server 配置注入 openclaw.json
  5. 检测并修复 docparse.py 中硬编码的路径,使其自适应

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · mcp/docparse.py (reported line 81)May include surrounding context.

python
功能:
  1. 自动识别 OpenClaw 安装路径与 workspace 位置
  2. 安装 Python 依赖(fastmcp, mcp)
  3. 创建 .env 配置文件(如不存在)
  4. 将 MCP Server 配置注入 openclaw.json
  5. 检测并修复 docparse.py 中硬编码的路径,使其自适应

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · mcp/docparse.py (reported line 82)May include surrounding context.

python
功能:
  1. 自动识别 OpenClaw 安装路径与 workspace 位置
  2. 安装 Python 依赖(fastmcp, mcp)
  3. 创建 .env 配置文件(如不存在)
  4. 将 MCP Server 配置注入 openclaw.json
  5. 检测并修复 docparse.py 中硬编码的路径,使其自适应

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · mcp/docparse.py (reported line 108)May include surrounding context.

python
功能:
  1. 自动识别 OpenClaw 安装路径与 workspace 位置
  2. 安装 Python 依赖(fastmcp, mcp)
  3. 创建 .env 配置文件(如不存在)
  4. 将 MCP Server 配置注入 openclaw.json
  5. 检测并修复 docparse.py 中硬编码的路径,使其自适应

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · mcp/docparse.py (reported line 118)May include surrounding context.

python
功能:
  1. 自动识别 OpenClaw 安装路径与 workspace 位置
  2. 安装 Python 依赖(fastmcp, mcp)
  3. 创建 .env 配置文件(如不存在)
  4. 将 MCP Server 配置注入 openclaw.json
  5. 检测并修复 docparse.py 中硬编码的路径,使其自适应

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · mcp/docparse.py (reported line 127)May include surrounding context.

python
功能:
  1. 自动识别 OpenClaw 安装路径与 workspace 位置
  2. 安装 Python 依赖(fastmcp, mcp)
  3. 创建 .env 配置文件(如不存在)
  4. 将 MCP Server 配置注入 openclaw.json
  5. 检测并修复 docparse.py 中硬编码的路径,使其自适应

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · setup.py (reported line 8)May include surrounding context.

python
功能:
  1. 自动识别 OpenClaw 安装路径与 workspace 位置
  2. 安装 Python 依赖(fastmcp, mcp)
  3. 创建 .env 配置文件(如不存在)
  4. 将 MCP Server 配置注入 openclaw.json
  5. 检测并修复 docparse.py 中硬编码的路径,使其自适应

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · setup.py (reported line 97)May include surrounding context.

python
功能:
  1. 自动识别 OpenClaw 安装路径与 workspace 位置
  2. 安装 Python 依赖(fastmcp, mcp)
  3. 创建 .env 配置文件(如不存在)
  4. 将 MCP Server 配置注入 openclaw.json
  5. 检测并修复 docparse.py 中硬编码的路径,使其自适应

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · setup.py (reported line 109)May include surrounding context.

python
功能:
  1. 自动识别 OpenClaw 安装路径与 workspace 位置
  2. 安装 Python 依赖(fastmcp, mcp)
  3. 创建 .env 配置文件(如不存在)
  4. 将 MCP Server 配置注入 openclaw.json
  5. 检测并修复 docparse.py 中硬编码的路径,使其自适应

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · setup.py (reported line 114)May include surrounding context.

python
功能:
  1. 自动识别 OpenClaw 安装路径与 workspace 位置
  2. 安装 Python 依赖(fastmcp, mcp)
  3. 创建 .env 配置文件(如不存在)
  4. 将 MCP Server 配置注入 openclaw.json
  5. 检测并修复 docparse.py 中硬编码的路径,使其自适应

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · setup.py (reported line 116)May include surrounding context.

python
功能:
  1. 自动识别 OpenClaw 安装路径与 workspace 位置
  2. 安装 Python 依赖(fastmcp, mcp)
  3. 创建 .env 配置文件(如不存在)
  4. 将 MCP Server 配置注入 openclaw.json
  5. 检测并修复 docparse.py 中硬编码的路径,使其自适应

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · setup.py (reported line 122)May include surrounding context.

python
功能:
  1. 自动识别 OpenClaw 安装路径与 workspace 位置
  2. 安装 Python 依赖(fastmcp, mcp)
  3. 创建 .env 配置文件(如不存在)
  4. 将 MCP Server 配置注入 openclaw.json
  5. 检测并修复 docparse.py 中硬编码的路径,使其自适应

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · setup.py (reported line 124)May include surrounding context.

python
功能:
  1. 自动识别 OpenClaw 安装路径与 workspace 位置
  2. 安装 Python 依赖(fastmcp, mcp)
  3. 创建 .env 配置文件(如不存在)
  4. 将 MCP Server 配置注入 openclaw.json
  5. 检测并修复 docparse.py 中硬编码的路径,使其自适应

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · setup.py (reported line 360)May include surrounding context.

python
功能:
  1. 自动识别 OpenClaw 安装路径与 workspace 位置
  2. 安装 Python 依赖(fastmcp, mcp)
  3. 创建 .env 配置文件(如不存在)
  4. 将 MCP Server 配置注入 openclaw.json
  5. 检测并修复 docparse.py 中硬编码的路径,使其自适应

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · setup.py (reported line 361)May include surrounding context.

python
功能:
  1. 自动识别 OpenClaw 安装路径与 workspace 位置
  2. 安装 Python 依赖(fastmcp, mcp)
  3. 创建 .env 配置文件(如不存在)
  4. 将 MCP Server 配置注入 openclaw.json
  5. 检测并修复 docparse.py 中硬编码的路径,使其自适应

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
93% confidence
Finding

The skill declares no explicit tool scope or permissions despite instructing use of environment variables, filesystem access, network connectivity, shell commands, and file writes. In an agent setting this over-privileges the skill, making unintended config changes, secret access, or remote exfiltration of document data more likely if the skill is invoked broadly or misused.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The trigger terms include common verbs like 'read' and 'extract text,' which can cause accidental activation in unrelated conversations. Because this skill sends documents to a remote service and may perform setup actions, over-broad invocation increases the chance of unintended data handling or execution of privileged actions.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
86% confidence
Finding

The natural-language instructions, examples, and user-facing outputs are entirely in Chinese, including mandated response text such as error and success messages. There is no indication that users may choose another language or that the locale restriction is required for a region-specific compliance purpose.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The activation guidance lacks crisp boundaries for when not to run, beyond a limited unsupported-format note. In an agent environment, ambiguous routing can cause the skill to process sensitive files or initiate remote parsing when the user only asked for summarization, viewing, or local-only handling.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.