Back to skill

Security audit

Book to Skill Converter 书本即技能

Security checks for vulnerabilities and agentic risk

Overview

This skill coherently helps turn a user-provided book into a draft skill, with some implementation hygiene risks but no evidence of deception, exfiltration, or hidden persistence.

Install only if you are comfortable letting the agent read the specific book files you provide and create draft skill files where you choose. Review generated skills before installing or reusing them, avoid processing untrusted MOBI files in sensitive directories, and prefer pinned, trusted parser dependencies instead of ad hoc pip installs.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T08 · Insecure Dependencies

Warning
Location
scripts/book_extractor.py:14
Finding

Unpinned Third-Party Dependency Installation Guidance

Content
View full analysis

Vulnerability Details

File Location: scripts/book_extractor.py:14-20, scripts/book_extractor.py:33-42, scripts/book_extractor.py:51-62, and scripts/book_extractor.py:64-72
Vulnerability Type: Supply-chain exposure through unpinned dependencies
Risk Level: Medium

Vulnerable Code

python
def extract_text_from_pdf(file_path):
    """从PDF提取文本"""
    try:
        import PyPDF2
        with open(file_path, 'rb') as f:
            reader = PyPDF2.PdfReader(f)
            text = ""
            for page in reader.pages:
                text += page.extract_text() + "\n"
            return text
    except ImportError:
        return "请安装 PyPDF2: pip install PyPDF2"
python
def extract_text_from_epub(file_path):
    """从EPUB提取文本"""
    try:
        import ebooklib
        from ebooklib import epub
        book = epub.read_epub(file_path)
        text = ""
        for item in book.get_items():
            if item.get_type() == 9:  # DOCUMENT
                text += item.get_content().decode('utf-8') + "\n"
        return text
    except ImportError:
        return "请安装 ebooklib: pip install ebooklib"
python
def extract_text_from_mobi(file_path):
    """从MOBI提取文本"""
    try:
        import mobi
        from pathlib import Path
        output_path = Path(file_path).with_suffix('.html')
        mobi.extract(file_path, output_path)
        with open(output_path, 'r', encoding='utf-8') as f:
            text = f.read()
        # 清理临时文件
        output_path.unlink(missing_ok=True)
        return text
    except ImportError:
        return "请安装 mobi: pip install mobi"
python
def extract_text_from_docx(file_path):
    """从DOCX提取文本"""
    try:
        import docx
        doc = docx.Document(file_path)
        text = ""
        for para in doc.paragraphs:
            text += para.text + "\n"
        return text
  
...[truncated 1984 chars]
Remediation
View remediation

Remediation Suggestions

  • Add a reviewed dependency manifest containing exact versions for all runtime and transitive dependencies.
  • Generate and verify cryptographic hashes, then install with a command such as pip install --require-hashes -r requirements.txt.
  • Use a lock-file workflow appropriate to the selected package manager.
  • Install packages only from explicitly trusted repositories and disable unintended extra indexes.
  • Perform dependency vulnerability and provenance scanning in CI.
  • Replace the unconstrained installation messages with instructions referencing the project's locked installation process.
  • Run document parsers in an isolated, least-privileged environment because they process attacker-controlled file formats.

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/book_extractor.py:41
Finding

Predictable and Unsafe MOBI Extraction Path

Content
View full analysis

Vulnerability Details

File Location: scripts/book_extractor.py:41-62
Vulnerability Type: Unsafe temporary-file and extraction-path handling
Risk Level: Medium

Vulnerable Code

python
def extract_text_from_mobi(file_path):
    """从MOBI提取文本"""
    try:
        import mobi
        from pathlib import Path
        output_path = Path(file_path).with_suffix('.html')
        mobi.extract(file_path, output_path)
        with open(output_path, 'r', encoding='utf-8') as f:
            text = f.read()
        # 清理临时文件
        output_path.unlink(missing_ok=True)
        return text
    except ImportError:
        return "请安装 mobi: pip install mobi"

Technical Analysis

The output path is derived predictably from the attacker-influenced input path by replacing its extension with .html. Extraction occurs beside the input rather than inside a securely created, isolated temporary directory.

The code does not reject an existing output path, check for symbolic links, verify the resolved destination, or confirm that every generated extraction path remains within a controlled directory. It also assumes that the second argument to mobi.extract identifies the exact HTML file later opened, even though extraction libraries may treat that argument as an output directory and may return the actual generated path separately.

Finally, cleanup calls unlink() on the predictable path without proving that it is a regular file created by this invocation. Depending on the extraction library's handling of existing paths and symbolic links, this creates opportunities for path collisions, local race conditions, unintended writes, deletion of an existing path entry, or denial of service.

Attack Path

  1. An attacker supplies or influences a MOBI input path in a directory where the attacker can also create adjacent entries.
  2. The attacker prepares the corresponding path produced by `Path(file_path).with ...[truncated 1234 chars]
Remediation
View remediation

Remediation Suggestions

  • Create a unique extraction directory with tempfile.TemporaryDirectory() rather than deriving a destination beside the input.
  • Pass that isolated directory to the MOBI extraction library and use the filepath returned by the library instead of assuming a fixed .html output filename.
  • Resolve and validate every extracted path, ensuring it remains beneath the temporary directory before opening it.
  • Reject symbolic links and non-regular files when selecting extracted content.
  • Avoid deleting individual paths derived from user input; allow the temporary-directory context manager to clean up only the directory created for the current invocation.
  • Verify the extraction API's return contract and handle directories, multiple generated files, malformed archives, and extraction errors explicitly.
  • Apply limits for extracted size, file count, nesting depth, and processing time to reduce archive-based denial-of-service risks.
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (8)

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

声明描述的是一个更高层的能力链路:用户上传书籍后,系统应分析内容、提炼技能/思维方式,并生成新的可复用 skill。但提供的代码块只实现了文件读取与文本提取,按格式调用不同解析器获取书籍原文,没有任何 NLP/信息抽取、结构化总结、技能识别、思维方式归纳或 skill 创建逻辑。虽然文本提取可能是该目标的前置支持步骤,但就当前代码块本身而言,其实际行为与声明的主要用途存在实质差异,因此应判定为不匹配。

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
89% confidence
Finding

The skill instructs reading user-supplied local files in multiple formats but declares no explicit tool scope or permissions boundary. That omission increases the chance the agent can access files more broadly than intended, especially in a meta-skill that may chain into other skills for PDF/DOCX parsing.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The description and the entire skill guidance are written as if the skill's output and interaction are expected to be in Chinese, and no opt-in or language-selection mechanism is provided. This can violate language/locale policy when a skill implicitly enforces one language without giving the user a choice.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

This markdown file contains user-facing instructional content exclusively in Chinese, and there is no indication that users can opt into another language or that the skill is intended only for a Chinese-language context. Under the policy rule for language or locale constraints, forcing a specific language without user opt-in is a natural-language policy violation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

This code file contains natural-language strings and docstrings exclusively in Chinese, including the title, usage text, and error messages. Under the policy, forcing a specific language without user opt-in is a natural-language policy violation unless the locale restriction is justified, which is not indicated here.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

This markdown template forces Chinese-language headings such as '核心要点', '具体步骤', and '常见问题' for all generated skills. That creates a language/locale policy issue because users are not given an opt-in or alternative language option, and the file does not document a justified region-specific constraint.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
77% confidence
Finding

The naming rule specifies a language preference for generated skill names, which is a locale-related instruction. Because the file does not indicate that this preference is optional or user-selected, it may impose a language convention without opt-in.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

The skill says generated output will be saved locally but gives no safety guidance about file creation, path selection, or overwrite behavior. In practice, this can lead to accidental overwriting of existing files or creation of artifacts in sensitive locations when handling user-specified paths.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.