Back to skill

Security audit

文档整理技能 (convert-markdown)

Security checks for vulnerabilities and agentic risk

Overview

This is a coherent document-to-Markdown conversion skill, with ordinary file read/write behavior but supply-chain and privacy precautions users should handle before use.

Install only in a virtual environment, prefer exact pinned package versions or a lockfile, and run conversions on a limited input folder rather than broad personal or business directories. Treat OCR, audio, spreadsheets, and PDFs as potentially sensitive, review generated Markdown/index files before sharing them, and use overwrite options only after confirming output paths.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T08 · Insecure Dependencies

Warning
Location
SKILL.md:39
Finding

Unpinned and Overly Broad Python Dependencies

Content
View full analysis
=3.10", "pip_packages": [ "markitdown[docx,xlsx,pdf]>=0.1.5" ] } ``` ### Technical Analysis The documented installation commands do not pin MarkItDown or its transitive dependencies to reviewed versions, and they do not require package hashes. The manifest similarly permits any MarkItDown release equal to or newer than version `0.1.5`. As a result, identical installation commands can resolve to different package artifacts over time. The recommended `[all]` extra also expands the dependency graph to optional document, media, OCR, archive, and network-related components that may not be necessary for a particular deployment. A larger dependency graph increases the supply-chain attack surface. This issue does not establish that MarkItDown or any current dependency is malicious. The risk is that a compromised upstream account, malicious future release, dependency-confusion event, or compromised transitive package could be incorporated without a project change or an additional review. ### Attack Path 1. An attacker compromises an upstream dependency, its publishing account, or a transitive package used by one of the selected extras. 2. The attacker publishes a malicious version that satisfies `>=0.1.5` or is selected by an unconstrained `pip install` command. 3. A user follows the project documentation or an automated environment processes `manifest.json`. 4. Pip resolves and installs the malicious o ...[truncated 879 chars]
Remediation
View remediation
``` 2. Generate and commit a lock file that records all transitive versions. 3. Require cryptographic hashes during installation, for example with a hash-locked requirements file and: ```bash pip install --require-hashes -r requirements.lock ``` 4. Replace the default recommendation for `[all]` with the minimum extras required for each use case. 5. Install packages from a trusted, explicitly configured index and disable unintended fallback indexes where practical. 6. Run dependency vulnerability and provenance checks in CI before publishing the Skill. 7. Test dependency updates in an isolated environment and require security review before updating the lock file. 8. Run conversion in a sandboxed, non-privileged environment with limited filesystem and network access, particularly when processing untrusted documents. ]]>

T08 · Insecure Dependencies

Warning
Location
SKILL.md:70
Finding

Unversioned NPX Commands Can Download and Execute Mutable Registry Content

Content
View full analysis
Remediation
View remediation
convert --input document.pdf --output document.md ``` 2. Prefer installing the package as a locked project dependency and invoking the local binary through an npm script. 3. Commit `package-lock.json` or another supported lock file and use `npm ci` for reproducible installation. 4. Verify registry provenance and package integrity before execution. 5. Configure the expected npm registry explicitly and protect against unintended registry substitution. 6. Avoid running NPX or document conversion with administrator or root privileges. 7. Execute the converter in a restricted environment with access only to required input and output paths. 8. Require security review and automated dependency scanning before changing the locked package version. ]]>
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (59)

Tp4

High
Category
MCP Tool Poisoning
Confidence
89% confidence
Finding

整体上,代码的主目的与声明基本一致,都是使用 MarkItDown 进行文件到 Markdown 的转换,并支持目录批量处理,因此不存在完全不同的功能或未声明的敏感能力。问题在于声明明确写到支持“图片、音频等多种格式文件批量转换”,而代码在批量目录处理时只筛选固定扩展名集合,不包括常见图片和音频格式。虽然单文件转换路径可能借助底层库处理更多格式,但从该代码块本身无法验证图片、音频的批量支持,因此声明与实际可确认行为之间存在一定不一致,属于描述范围大于代码可证实能力的轻度到中度偏差。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

声明描述覆盖了“批量将多种文件转换为 Markdown”这一部分,代码也确实使用 MarkItDown 执行该功能。但该代码块的主要目的明显更具体:它不仅转换文档,还构建知识库索引,按目录组织条目,并输出机器可读与人类可读的目录文件。这属于超出声明的实质性能力,而不是单纯的实现细节。另一个差异是,声明提到支持图片、音频等多种格式,但此代码实际只显式筛选 pdf/docx/pptx/xlsx/txt/html/csv/json/xml/zip/epub/png/jpg/jpeg,并未体现音频支持。因此描述与实际行为存在明显但非恶意的能力/范围不一致。

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The natural-language documentation is presented in Chinese throughout, which effectively forces a specific language for users without any opt-in or explanation of locale scope. The policy requires flagging language or locale constraints unless the skill offers a language choice or clearly documents a justified regional restriction.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The README promotes OCR and extraction for PDFs, images, and other documents without warning that these files may contain confidential, personal, or regulated information. In this skill's context, the omission is meaningful because the tool's entire purpose is to ingest and transform potentially sensitive files at scale.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The README advertises audio transcription capability without warning that uploaded or processed audio may contain sensitive personal, legal, or regulated content. For a document-processing skill, users may reasonably feed confidential recordings into the tool without understanding privacy, retention, or third-party model implications.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
70% confidence
Finding

npx commands without a version suffix (e.g. @1.0.0) create a rug-pull risk if the upstream server is compromised and publishes a malicious update.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
70% confidence
Finding

npx commands without a version suffix (e.g. @1.0.0) create a rug-pull risk if the upstream server is compromised and publishes a malicious update.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
70% confidence
Finding

npx commands without a version suffix (e.g. @1.0.0) create a rug-pull risk if the upstream server is compromised and publishes a malicious update.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
70% confidence
Finding

npx commands without a version suffix (e.g. @1.0.0) create a rug-pull risk if the upstream server is compromised and publishes a malicious update.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
93% confidence
Finding

The README instructs users to run npx convert-markdown without pinning a package version. npx may fetch and execute the latest published package, so if the package is compromised, typosquatted, or unexpectedly changed, users could execute attacker-controlled code.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
94% confidence
Finding

This command tells users to execute npx convert-markdown convert ... without a pinned version. Because npx resolves packages dynamically, a compromised upstream release or namespace confusion can lead to arbitrary code execution on the user's machine.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
94% confidence
Finding

The unpinned npx batch conversion example can cause users to download and run whatever package is current at execution time. In documentation for a file-conversion tool that processes local documents, this increases risk because the executed package may gain access to sensitive files supplied as inputs.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
94% confidence
Finding

The README provides another unpinned npx command for batch processing. Dynamic retrieval at runtime exposes users to supply-chain attacks, and the batch-processing context could amplify impact by granting malicious code access to entire directories of documents.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The README documents an --overwrite option but gives no warning that existing files may be replaced. In a batch document-conversion context, users may unintentionally destroy or replace important Markdown outputs or other files if paths are mistaken, leading to data loss.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
94% confidence
Finding

This overwrite example still relies on unpinned npx, so it can execute an untrusted package version. Because the command also writes output files, a malicious package could both exfiltrate inputs and tamper with local content under the guise of conversion.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
94% confidence
Finding

This quick-start batch conversion command uses npx convert-markdown without version pinning, exposing users to runtime package substitution or malicious updates. Since the tool is intended to process knowledge-base documents, compromise could expose large sets of internal files.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
94% confidence
Finding

The OCR example uses unpinned npx, which can fetch and execute attacker-controlled code if the package or dependency chain is compromised. OCR workflows commonly involve sensitive scanned documents, so the consequences can include data theft beyond mere command execution.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
94% confidence
Finding

This spreadsheet conversion example again uses unpinned npx, creating a supply-chain execution risk. Because spreadsheets often contain business or financial data, a compromised package could directly access and exfiltrate high-value information during processing.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
92% confidence
Finding

The skill documents shell commands and file-writing behavior but does not declare any tool scope such as permissions or allowed-tools. This creates an implicit-trust situation where an agent or operator may execute filesystem and shell actions without an explicit least-privilege boundary, increasing the chance of unintended file modification or command execution.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
78% confidence
Finding

The description and instructions are presented in Chinese, but the file does not indicate that the skill is region-specific or offer an alternative language option. Under the language/locale policy, forcing a specific language without user opt-in or justification is a policy concern.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

The skill promotes OCR, audio transcription, and YouTube-related processing without warning about privacy, copyright, or sensitive-data handling. Users may feed confidential documents, images, or recordings into conversion workflows without understanding retention, licensing, or exposure risks.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
87% confidence
Finding

The skill recommends invoking an NPX-based CLI without pinning an exact package version. Unpinned package execution can pull updated or maliciously substituted code at runtime, introducing supply-chain risk and reducing reproducibility.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
87% confidence
Finding

The NPX CLI reference is versionless, so users may execute whatever package version resolves at the time of use. That exposes users to supply-chain compromise, breaking changes, or malicious package takeover.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
94% confidence
Finding

Running npx convert-markdown without a pinned version may fetch and execute remote code dynamically. In a skill that performs file conversion and writes outputs, that increases the blast radius of a compromised dependency.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
94% confidence
Finding

This example executes a versionless NPX package for document conversion, creating a supply-chain risk path. Because the command handles user-supplied files and output locations, a malicious package could read, alter, or exfiltrate local content.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.dangerous_exec

Shell command execution detected (child_process).

Critical
Code
suspicious.dangerous_exec
Location
bin/convert-markdown.js:33