Back to skill

Security audit

everything to markdown 中文版

Security checks for vulnerabilities and agentic risk

Overview

This is a coherent Chinese-language document conversion skill, with ordinary but broad dependency and media-processing risks users should understand before installing.

Install this only if you are comfortable adding MarkItDown and its optional dependencies. Prefer a virtual environment, pin reviewed versions where possible, and avoid processing confidential documents, recordings, or URLs unless you understand how the installed converters handle that content.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T08 · Insecure Dependencies

Warning
Location
SKILL.md:15
Finding

Unpinned Third-Party Dependencies Permit Supply-Chain Compromise

Content
View full analysis

Vulnerability Details

File Location: SKILL.md:15-18, SKILL.md:410, SKILL.md:413, SKILL.md:440, and README.md:14-17
Vulnerability Type: Unpinned and unverified third-party package installation
Risk Level: Medium

Vulnerable Code

SKILL.md:15-18:

yaml
install:
  - id: install-markitdown
    kind: exec
    command: pip3 install 'markitdown[all]'

SKILL.md:410:

bash
pip install 'markitdown[all]'

SKILL.md:413:

bash
conda install -c conda-forge markitdown

SKILL.md:440:

bash
pip install markitdown-mcp

README.md:14-17:

bash
pip3 install 'markitdown[all]'

Technical Analysis

The Skill installs markitdown, all packages selected by its all extra, and the optional markitdown-mcp package without fixed versions or integrity hashes. Consequently, the installed code is determined by mutable package-repository state at installation time rather than by the version reviewed during this audit.

The markitdown[all] specification also expands the supply-chain and parser attack surface by installing optional transitive dependencies. None of the installation instructions use a lock file, hash verification, an explicitly trusted package index, or an isolated execution policy.

The metadata installation command is particularly relevant because a compatible Skill framework may execute it as part of automated Skill setup. Package installation can execute package build hooks, and installed packages later execute within the MarkItDown command, Python API, or MCP server process.

This finding does not establish that the current upstream packages are malicious. The vulnerability is the absence of controls that bind installation to reviewed artifacts.

Attack Path

  1. An attacker compromises the publishing account, build pipeline, distribution artifact, or transitive dependency of one of the referenced packages.
  2. The attacker publishes a malicious version that satisfies the unrestricted package ...[truncated 1341 chars]
Remediation
View remediation

Remediation Suggestions

  1. Pin every direct dependency to a reviewed version:
bash
python3 -m pip install 'markitdown[all]==REVIEWED_VERSION'
python3 -m pip install 'markitdown-mcp==REVIEWED_VERSION'
  1. Generate a lock file containing exact versions for all transitive dependencies. Review and update it through a controlled dependency-update process.

  2. Require cryptographic hashes for downloaded artifacts, for example with a hash-locked requirements file:

bash
python3 -m pip install --require-hashes -r requirements.lock
  1. Avoid the broad all extra unless every optional converter is required. Install only the format-specific extras needed by the Skill to reduce the dependency and parser attack surface.

  2. Configure an explicit trusted package index or an internally controlled package mirror. Do not allow fallback to untrusted or unintended indexes.

  3. Install packages inside a dedicated virtual environment or container rather than into the system Python environment:

bash
python3 -m venv .venv
. .venv/bin/activate
python3 -m pip install --require-hashes -r requirements.lock
  1. Run document conversion and the MCP server as a non-privileged account with restricted filesystem and network access. Do not install or invoke these components as root.

  2. Add automated dependency scanning, provenance verification, and periodic review of pinned updates. Recreate and test the lock file whenever an approved dependency version changes.

  3. Make the installation policy consistent across SKILL.md and README.md so users cannot bypass the hardened process by following alternate instructions.

Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
Findings (6)

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The README advertises audio transcription and YouTube subtitle extraction but does not warn users that these features may involve external network access, third-party services, or processing of potentially sensitive media. In an agent/LLM workflow, this omission is more dangerous because users may automate document handling at scale and unknowingly expose private content, credentials spoken in audio, or copyrighted material to remote services.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
82% confidence
Finding

Natural-language policy requires avoiding forced language or locale constraints unless users are given a choice or the restriction is clearly justified. This skill's name, tags, and full documentation indicate a Chinese-only experience, but there is no explicit opt-in, alternative language option, or justification for the locale constraint.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
85% confidence
Finding

The skill advertises OCR, audio transcription, and YouTube processing but does not warn that these capabilities may invoke optional components, external models, network access, or privacy-sensitive handling of uploaded media and extracted content. In an agent/LLM workflow, users may process confidential files or recordings without realizing sensitive data could be exposed beyond the local environment.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
93% confidence
Finding

Natural-language policy violations apply to all file types. This README presents the skill description only in Chinese, which can amount to forcing a specific language without user opt-in or an explicit justification for the locale constraint.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
80% confidence
Finding

Recommending markitdown[all] installs the broadest optional dependency set, which may enable network-facing, media-processing, or other higher-risk functionality without clearly informing users. This increases attack surface and may surprise operators who expected a minimal local document-conversion tool.

Content

No source excerpt is available for this finding.

Rp1

Low
Category
MCP Rug Pull
Confidence
60% confidence
Finding

pip install without ==version installs the latest release, which could include malicious changes.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.