Back to skill

Security audit

pdf-ppt-docx-xlsx-tools

Security checks for vulnerabilities and agentic risk

Overview

This appears to be a normal local document conversion skill, with some hardening issues users should handle carefully.

Install only if you need local document conversion. Run it on documents and output directories you choose, review before allowing system package installs, and avoid opening or hosting HTML generated from untrusted DOCX files unless it is sanitized first.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/docx_to_html.py:29
Finding

Stored HTML Injection in DOCX-to-HTML Conversion

Content
View full analysis
{text}" elif "heading 2" in style: return f"

{text}

" elif "heading 3" in style: return f"

{text}

" elif "list" in style: return f"
  • {text}
  • " else: return f"

    {text}

    " ``` ### Technical Analysis Paragraph text extracted from a potentially untrusted DOCX file is interpolated directly into HTML elements without HTML encoding or sanitization. Consequently, characters such as `<`, `>`, `&`, single quotes, and double quotes retain their markup semantics. An attacker can place active HTML in a DOCX paragraph, for example: ```html ``` The converter writes this content unchanged inside a generated paragraph: ```html

    ``` When the generated file is opened in a browser or published by a web service, the event handler executes as JavaScript. Other payloads may create deceptive forms, initiate browser requests, alter displayed content, or interact with data available to the generated document's origin. The same unsafe conversion pattern is also documented in the inline example at `SKILL.md:232-251`, which may cause users to reproduce the vulnerability even if the standalone script is corrected. ### Attack Path 1. An attacker creates a DOCX document containing HTML or JavaScript-bearing markup in a paragraph. 2. The victim or an automated service runs: ```bash python scripts/docx_to_html.py attacker.docx output.html ``` 3. `para_to_html()` reads the attacker-controlled paragraph text and inserts it directly into an ...[truncated 1170 chars]
    Remediation
    View remediation
    {text}" elif "heading 2" in style: return f"

    {text}

    " elif "heading 3" in style: return f"

    {text}

    " elif "list" in style: return f"
  • {text}
  • " return f"

    {text}

    " ``` Additional hardening measures: 1. Apply contextual encoding to every untrusted value written into HTML, including any future attributes, links, image names, table cells, comments, or metadata. 2. If preserving limited document formatting requires accepting HTML, process it through a mature allowlist-based sanitizer rather than interpolating it directly. 3. Disallow scripts, inline event handlers, unsafe URL schemes, active SVG content, embedded objects, and dangerous CSS. 4. Serve generated HTML from an isolated origin with no sensitive cookies or application data. 5. Apply a restrictive Content Security Policy where generated content is hosted, such as disabling scripts unless they are explicitly required. 6. Add regression tests covering script tags, event-handler attributes, malformed tags, SVG payloads, entity-encoded payloads, and quotation-mark breakout attempts. 7. Correct the corresponding unsafe inline implementation in `SKILL.md:232-251`. ]]>
    Vulnerability Patterns
    • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
    • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
    • Tool MisuseTool Parameter Abuse, Chaining Abuse, Unsafe Defaults
    • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
    • Behavioral ASTexec() Call, eval() Call, Dynamic Import
    Findings (37)

    Tp4

    High
    Category
    MCP Tool Poisoning
    Confidence
    95% confidence
    Finding

    声明描述的是一个覆盖多种文档格式的综合转换与处理工具集,重点在“PDF、PPTX、DOCX、XLSX 四种格式之间的互转及衍生操作”。但给出的代码块功能范围明显更窄,只处理 DOCX 文件,并且仅支持合并、提取文本/表格、查看文档信息。其主要目的不是“多格式转换”,而是“DOCX 专项处理”。此外,代码还提供了文档元数据/info 功能,这属于声明中未明确列出的能力。虽然其中的合并、提取文本与声明中的部分子能力一致,但整体描述与实际行为在核心能力范围上存在实质性不匹配。

    Content

    No source excerpt is available for this finding.

    Tp4

    High
    Category
    MCP Tool Poisoning
    Confidence
    98% confidence
    Finding

    声明描述的是一个广义的多格式文档转换与处理工具集,核心能力应包括多种 Office/PDF 格式之间转换以及若干衍生操作。但该代码块的实际范围明显更窄,仅针对 PDF 文件执行合并、拆分、加水印、提取页面和查看信息,属于 PDF 专用处理脚本,而非多格式转换工具。虽然其中的合并、拆分、加水印与声明中的部分 PDF 操作一致,但大量关键宣称能力缺失,导致描述对代码能力有明显夸大。此外,代码还提供了 PDF 元数据/页面信息查看,这一能力未在描述中明确提及。综合来看,声明与实际行为存在实质性不匹配。

    Content

    No source excerpt is available for this finding.

    Tool Parameter Abuse

    High
    Category
    Tool Misuse
    Confidence
    93% confidence
    Finding

    The use of shell=True for executable lookup is a tool-parameter abuse pattern because it delegates parsing and execution to the shell unnecessarily. In the context of a document-conversion skill that may run in varied environments, this can make PATH hijacking or shell-environment manipulation more relevant, even though the immediate command strings are fixed.

    Content

    Scanner excerpt · scripts/docx_to_pdf.py (reported line 35)May include surrounding context.

    python
    # 尝试 which/where
        for cmd in ["where soffice", "which soffice", "which libreoffice"]:
            try:
                result = subprocess.run(cmd, shell=True, capture_output=True, text=True)
                if result.returncode == 0 and result.stdout.strip():
                    return result.stdout.strip().split("\n")[0]
            except Exception:
    

    Tool Parameter Abuse

    High
    Category
    Tool Misuse
    Confidence
    90% confidence
    Finding

    This is a real tool-parameter abuse pattern because the code launches a shell for executable discovery instead of using safer APIs. In the context of a document-conversion skill, the danger is not from direct user command injection into the string, but from executing attacker-influenced binaries or shell behavior in compromised or multi-tenant environments.

    Content

    Scanner excerpt · scripts/pptx_to_images.py (reported line 35)May include surrounding context.

    python
    return path
        for cmd in ["where soffice", "which soffice", "which libreoffice"]:
            try:
                result = subprocess.run(cmd, shell=True, capture_output=True, text=True)
                if result.returncode == 0 and result.stdout.strip():
                    return result.stdout.strip().split("\n")[0]
            except Exception:
    

    Tool Parameter Abuse

    High
    Category
    Tool Misuse
    Confidence
    96% confidence
    Finding

    Using shell=True for tool discovery is an instance of unsafe tool invocation. In the context of a document-conversion skill, which may be triggered automatically and process untrusted inputs in diverse environments, relying on shell command resolution can allow a malicious PATH entry or environment manipulation to redirect execution to an attacker-controlled program.

    Content

    Scanner excerpt · scripts/pptx_to_pdf.py (reported line 29)May include surrounding context.

    python
    return path
        for cmd in ["where soffice", "which soffice", "which libreoffice"]:
            try:
                result = subprocess.run(cmd, shell=True, capture_output=True, text=True)
                if result.returncode == 0 and result.stdout.strip():
                    return result.stdout.strip().split("\n")[0]
            except Exception:
    

    Tool Parameter Abuse

    High
    Category
    Tool Misuse
    Confidence
    95% confidence
    Finding

    The tool performs executable discovery through shell=True, which creates an unnecessary command-execution surface during parameter handling. In the context of a document-conversion skill that may run on varied developer or automation hosts, this is more dangerous because environment manipulation or PATH hijacking could redirect the conversion flow to attacker-controlled executables.

    Content

    Scanner excerpt · scripts/xlsx_to_pdf.py (reported line 29)May include surrounding context.

    python
    return path
        for cmd in ["where soffice", "which soffice", "which libreoffice"]:
            try:
                result = subprocess.run(cmd, shell=True, capture_output=True, text=True)
                if result.returncode == 0 and result.stdout.strip():
                    return result.stdout.strip().split("\n")[0]
            except Exception:
    

    Undeclared Tool Scope

    Medium
    Category
    MCP Least Privilege
    Confidence
    89% confidence
    Finding

    The skill clearly instructs shell execution and file-writing operations but does not declare any tool scope, permissions, or allowed-tools boundaries. That makes it easier for an agent runtime to invoke broader capabilities than users or platform policy expect, increasing the chance of unsafe command execution or uncontrolled file modification.

    Content

    No source excerpt is available for this finding.

    Vague Triggers

    Medium
    Category
    Not specified by scanner
    Confidence
    95% confidence
    Finding

    描述中写明“当用户需要转换文档格式、处理 PDF、操作 Office 文件时使用此技能”,其中“处理 PDF”“操作 Office 文件”范围非常广,覆盖大量普通办公表达,缺少明确边界或排除条件。这种宽泛触发语容易导致技能在非目标场景下被误调用。

    Content

    No source excerpt is available for this finding.

    Natural-Language Policy Violations

    Medium
    Category
    Not specified by scanner
    Confidence
    93% confidence
    Finding

    技能名称与正文并非问题,但描述文本整体仅以中文表述技能用途,没有说明是否仅面向中文用户,也未给出语言可选项。若组织要求不得在无用户选择的情况下强制特定语言或区域设置,这种单一语言约束属于自然语言策略风险。

    Content

    No source excerpt is available for this finding.

    Missing User Warnings

    Medium
    Category
    Not specified by scanner
    Confidence
    91% confidence
    Finding

    The skill repeatedly writes, converts, merges, splits, and exports files, often in batch form, without warning about overwrites, destination control, or accidental disclosure through generated outputs. In an agent setting, missing guardrails around filesystem changes can cause data loss, unintended persistence of sensitive content, or writing files to unsafe locations.

    Content

    No source excerpt is available for this finding.

    Context-Inappropriate Capability

    Medium
    Category
    Not specified by scanner
    Confidence
    93% confidence
    Finding

    The manifest describes a document conversion and processing toolkit, which justifies local file parsing and writing. However, the documentation explicitly depends on shelling out to system tools (soffice) and even shows OS package-manager commands (winget, brew, apt, later which/where), introducing command-execution capability that is not itself part of the stated user-facing purpose.

    Content

    No source excerpt is available for this finding.

    Sudo/Root Execution

    Medium
    Category
    Privilege Escalation
    Confidence
    87% confidence
    Finding

    The skill documentation includes a privileged installation command using sudo, which can lead operators or automated agents to perform system-wide changes with elevated privileges. In this context the command is for dependency setup rather than exploitation, but encouraging root-level execution without strong warnings increases the blast radius of mistakes and supply-chain risk from package installation.

    Content

    Scanner excerpt · SKILL.md (reported line 27)May include surrounding context.

    brew install --cask libreoffice

    Ubuntu/Debian

    sudo apt install libreoffice

    text
    
    PDF 转图片如需高质量渲染,可选装 poppler(`PyMuPDF` 内置渲染已足够,poppler 仅作为备选)。
    

    Natural-Language Policy Violations

    Medium
    Category
    Not specified by scanner
    Confidence
    95% confidence
    Finding

    The module docstring and subsequent console messages are written entirely in Chinese, which imposes a specific language on users. Under the policy, language constraints should either be optional/opt-in or clearly justified as region-specific; neither is present here.

    Content

    No source excerpt is available for this finding.

    Natural-Language Policy Violations

    Medium
    Category
    Not specified by scanner
    Confidence
    97% confidence
    Finding

    The file's docstring and runtime messages are written in Chinese only, which imposes a specific language on users without any opt-in or documented justification. This matches the policy concern for language or locale restrictions in natural-language content.

    Content

    No source excerpt is available for this finding.

    Natural-Language Policy Violations

    Medium
    Category
    Not specified by scanner
    Confidence
    91% confidence
    Finding

    This file presents its docstring usage instructions in Chinese, and later user-facing print messages are also Chinese-only. Under the stated policy, forcing a specific language without offering a user choice or documenting a justified locale constraint is a natural-language policy violation.

    Content

    No source excerpt is available for this finding.

    subprocess module call

    Medium
    Category
    Dangerous Code Execution
    Confidence
    92% confidence
    Finding

    This code uses subprocess.run with shell=True to execute command strings for tool discovery. Although the strings are currently constant, shell invocation is unnecessary and increases attack surface by relying on shell parsing and potentially attacker-influenced PATH or shell behavior, especially in hostile or multi-user environments.

    Content

    Scanner excerpt · scripts/docx_to_pdf.py (reported line 35)May include surrounding context.

    python
    # 尝试 which/where
        for cmd in ["where soffice", "which soffice", "which libreoffice"]:
            try:
                result = subprocess.run(cmd, shell=True, capture_output=True, text=True)
                if result.returncode == 0 and result.stdout.strip():
                    return result.stdout.strip().split("\n")[0]
            except Exception:
    

    Natural-Language Policy Violations

    Medium
    Category
    Not specified by scanner
    Confidence
    95% confidence
    Finding

    This code file contains natural-language content in the module docstring that fixes the interaction language to Chinese. Under the policy rule, forcing a specific language without user opt-in is a natural-language policy violation unless the locale restriction is explicitly justified.

    Content

    No source excerpt is available for this finding.

    Natural-Language Policy Violations

    Medium
    Category
    Not specified by scanner
    Confidence
    97% confidence
    Finding

    This code file contains natural-language strings exclusively in Chinese in the module docstring and throughout user-facing output, which imposes a specific language on users. The policy allows locale constraints only when opt-in is offered or the constraint is clearly documented and justified, neither of which is present here.

    Content

    No source excerpt is available for this finding.

    Natural-Language Policy Violations

    Medium
    Category
    Not specified by scanner
    Confidence
    94% confidence
    Finding

    The module docstring presents the tool's interface and user-facing descriptions entirely in Chinese, which indicates a fixed language choice for user interaction. Under the stated policy, forcing a specific language without offering a user choice or documenting a justified locale constraint is a natural-language policy violation.

    Content

    No source excerpt is available for this finding.

    subprocess module call

    Medium
    Category
    Dangerous Code Execution
    Confidence
    91% confidence
    Finding

    The code uses subprocess.run with shell=True to execute command strings when locating LibreOffice. Even though the current strings are hardcoded, shell invocation unnecessarily exposes the process to PATH/environment manipulation and shell-resolution risks, which is more concerning in an agent skill that may run in varied host environments.

    Content

    Scanner excerpt · scripts/pptx_to_images.py (reported line 35)May include surrounding context.

    python
    return path
        for cmd in ["where soffice", "which soffice", "which libreoffice"]:
            try:
                result = subprocess.run(cmd, shell=True, capture_output=True, text=True)
                if result.returncode == 0 and result.stdout.strip():
                    return result.stdout.strip().split("\n")[0]
            except Exception:
    

    subprocess module call

    Medium
    Category
    Dangerous Code Execution
    Confidence
    70% confidence
    Finding

    subprocess module calls execute external commands. Without careful input validation, this enables command injection.

    Content

    Scanner excerpt · scripts/pptx_to_images.py (reported line 76)May include surrounding context.

    python
    # LibreOffice 转 PDF
        tmp_dir = tempfile.mkdtemp()
        subprocess.run(
            [soffice_path, "--headless", "--convert-to", "pdf", "--outdir", tmp_dir, pptx_path],
            check=True, capture_output=True,
        )
    

    subprocess module call

    Medium
    Category
    Dangerous Code Execution
    Confidence
    94% confidence
    Finding

    This call uses shell=True to execute command strings while resolving the soffice path. Even though the current strings are constant, invoking a shell unnecessarily increases attack surface and can enable execution of an attacker-controlled binary via PATH manipulation or shell/environment quirks, which is more concerning in an agent skill that may run in varied environments.

    Content

    Scanner excerpt · scripts/pptx_to_pdf.py (reported line 29)May include surrounding context.

    python
    return path
        for cmd in ["where soffice", "which soffice", "which libreoffice"]:
            try:
                result = subprocess.run(cmd, shell=True, capture_output=True, text=True)
                if result.returncode == 0 and result.stdout.strip():
                    return result.stdout.strip().split("\n")[0]
            except Exception:
    

    Natural-Language Policy Violations

    Medium
    Category
    Not specified by scanner
    Confidence
    95% confidence
    Finding

    This code file contains natural-language strings entirely in Chinese in the module docstring and later user-facing output, which imposes a specific language on users. The policy allows locale constraints only when they are justified or when users are given an explicit choice, neither of which is present here.

    Content

    No source excerpt is available for this finding.

    subprocess module call

    Medium
    Category
    Dangerous Code Execution
    Confidence
    92% confidence
    Finding

    This call uses shell=True to execute lookup commands, which is unsafe because shell resolution depends on the runtime environment and can be influenced by PATH or shell behavior. In hostile or multi-user environments, an attacker may be able to cause execution of an unintended binary or alter lookup results, leading to command execution or use of a malicious soffice path.

    Content

    Scanner excerpt · scripts/xlsx_to_pdf.py (reported line 29)May include surrounding context.

    python
    return path
        for cmd in ["where soffice", "which soffice", "which libreoffice"]:
            try:
                result = subprocess.run(cmd, shell=True, capture_output=True, text=True)
                if result.returncode == 0 and result.stdout.strip():
                    return result.stdout.strip().split("\n")[0]
            except Exception:
    

    subprocess module call

    Medium
    Category
    Dangerous Code Execution
    Confidence
    70% confidence
    Finding

    subprocess module calls execute external commands. Without careful input validation, this enables command injection.

    Content

    Scanner excerpt · scripts/docx_to_pdf.py (reported line 76)May include surrounding context.

    python
    cmd = [soffice_path, "--headless", "--convert-to", "pdf", "--outdir", out_dir, xlsx_path]
        print(f"执行: {' '.join(cmd)}")
        result = subprocess.run(cmd, capture_output=True, text=True)
    
        if result.returncode != 0:
            print(f"错误: {result.stderr}")
    

    Static analysis

    No suspicious patterns detected.