Back to skill

Security audit

Microsoft MarkItDown

Security checks for vulnerabilities and agentic risk

Overview

This is a straightforward MarkItDown document-conversion skill with normal dependency and privacy considerations, and I found no hidden exfiltration, destructive behavior, or deceptive instructions.

Install this in a dedicated virtual environment or container, prefer pinned reviewed versions instead of the unpinned examples, and avoid converting confidential documents or private media unless you trust the environment and understand whether the underlying converter may fetch remote URL content.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T08 · Insecure Dependencies

Warning
Location
SKILL.md:25
Finding

Unpinned Third-Party Dependencies Permit Supply-Chain Compromise

Content
View full analysis

Vulnerability Details

File Location: SKILL.md:25, SKILL.md:30, SKILL.md:35, and SKILL.md:104
Vulnerability Type: Unpinned and unverified third-party dependencies
Risk Level: Medium

Affected code snippets:

SKILL.md:25

bash
pipx install 'markitdown[all]'

SKILL.md:30

bash
pip install 'markitdown[all]'

SKILL.md:35

bash
pip install 'markitdown[pdf,docx,pptx]'

SKILL.md:104

bash
pip install markitdown-mcp

Technical Analysis

The documented installation commands resolve mutable, unpinned versions of markitdown, markitdown-mcp, and their transitive dependencies from the configured package index. The [all] extra substantially expands the dependency tree and therefore the supply-chain attack surface.

The project does not provide a lockfile, constraints file, package hashes, exact version requirements, or index restrictions. Consequently, the code installed by these commands can change after this Skill has been reviewed. A malicious or compromised future package release, compromised transitive dependency, dependency-confusion package from an untrusted configured index, or unexpected upstream change could introduce arbitrary code into the user's environment.

The referenced package names and upstream links appear legitimate, and the audited project contains no evidence that the current packages are malicious. The risk arises from executing future dependency versions without reproducibility or integrity verification.

Attack Path

  1. An attacker compromises a relevant package or transitive dependency, publishes a malicious future release, or introduces a higher-priority package through an untrusted package index.
  2. A user follows one of the installation commands in SKILL.md.
  3. pip or pipx resolves the latest compatible package graph because no exact versions, hashes, or trusted-index restrictions are specified.
  4. The package ma ...[truncated 1067 chars]
Remediation
View remediation

Remediation Suggestions

  1. Pin every direct dependency to a reviewed exact version, including markitdown and markitdown-mcp.
  2. Generate and commit a reproducible lockfile or constraints file that also constrains transitive dependencies.
  3. Require cryptographic hashes for downloaded distributions, such as through a requirements file used with pip install --require-hashes.
  4. Install only the format-specific extras required by the deployment rather than [all].
  5. Restrict installation to an explicitly trusted package index and prevent unintended fallback to user-controlled or internal indexes.
  6. Review package provenance, release signatures where available, maintainers, and dependency changes before updating pinned versions.
  7. Perform installation and conversion in a dedicated, non-privileged virtual environment or container with minimal filesystem and network access.
  8. Add automated dependency auditing and update pins only through a reviewed change process.
  9. Pin and verify markitdown-mcp independently because it exposes an integration service and may have a dependency graph distinct from the core converter.

A hardened installation flow should use reviewed exact versions and a hash-locked requirements file rather than resolving mutable latest releases directly from the documentation.

Vulnerability Patterns
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (9)

Tp2

High
Category
MCP Tool Poisoning
Confidence
85% confidence
Finding

Mixing characters from multiple Unicode scripts in a single identifier is a common technique to create visually ambiguous tool names.

Content

No source excerpt is available for this finding.

Tp2

High
Category
MCP Tool Poisoning
Confidence
85% confidence
Finding

Mixing characters from multiple Unicode scripts in a single identifier is a common technique to create visually ambiguous tool names.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
94% confidence
Finding

The code’s core purpose does align partially with the description: it is a wrapper that invokes markitdown to convert documents, including optional page selection and alternate output formats. However, the declared description specifically claims support for OCR, audio transcription, YouTube extraction, and broad format coverage such as images and audio. This code chunk does not implement or explicitly handle those capabilities; it only passes an input path/URL to the external markitdown binary and, in batch mode, only scans for a narrow set of extensions (.pdf, .docx, .pptx, .xlsx, .xls, .html, .htm). Because the claimed specialized capabilities and related triggers are not substantiated by the supplied code, the description overstates what this chunk actually does.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
87% confidence
Finding

The skill clearly instructs use of shell commands such as pip/pipx installs and command-line execution, but it does not declare any explicit tool scope or permissions. In an agent ecosystem, that omission weakens policy enforcement and can allow broader-than-expected command execution pathways when the skill is invoked.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The description promotes OCR, audio transcription, and YouTube extraction without warning that these operations may handle sensitive personal or confidential content and may involve remote retrieval/processing. Users and orchestrators may invoke the skill without understanding the privacy implications, leading to unintended disclosure of document contents, media, or accessed URLs.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

A broad trigger such as convert document can cause accidental invocation on unrelated or sensitive user requests, especially in agents that auto-select skills by fuzzy matching. Because this skill can lead to shell usage, file processing, and possibly network-backed content extraction, unintended activation increases the chance of privacy or safety mistakes.

Content

No source excerpt is available for this finding.

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
70% confidence
Finding

Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.

Content

Scanner excerpt · SKILL.md (reported line 138)May include surrounding context.

brew install openjdk@17

Ubuntu

sudo apt install openjdk-17-jdk

text

### Permission denied

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/convert.py (reported line 50)May include surrounding context.

python
cmd.append("--quiet")
    
    try:
        result = subprocess.run(
            cmd,
            capture_output=True,
            text=True,

Rp1

Low
Category
MCP Rug Pull
Confidence
78% confidence
Finding

Installing an MCP server package without pinning a version makes builds non-reproducible and exposes users to unexpected upstream changes or a compromised future release. Because MCP components may interface directly with LLM tooling, a malicious or breaking update could alter behavior or expand attack surface unexpectedly.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.