Back to skill

Security audit

EPUB Bilingual Converter Skill

Security checks for vulnerabilities and agentic risk

Overview

This skill is a coherent local EPUB-to-bilingual-EPUB workflow, with ordinary file-processing risks but no evidence of hidden, deceptive, credential-seeking, persistent, or exfiltrating behavior.

Install dependencies in an isolated virtual environment, consider pinning beautifulsoup4 and lxml, process only EPUBs you trust, and choose a fresh dedicated output directory because the skill refreshes the summary subdirectory during extraction.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (2)

T08 · Insecure Dependencies

Warning
Location
SKILL.md:28
Finding

Unpinned Third-Party Dependency Installation

Content
View full analysis

Vulnerability Details

File Location: SKILL.md:28-33
Vulnerability Type: Unpinned dependencies and unsafe supply-chain trust
Risk Level: Medium

Vulnerable Code

bash
python3 -m pip install beautifulsoup4 lxml

The same installation pattern also appears in README.md:23-35, including:

bash
python3 -m pip install beautifulsoup4 lxml
python3 -m pip install pytest

Technical Analysis

The skill directs users or agents to install packages without version constraints or cryptographic hash verification. Consequently, installation resolves whichever versions the configured Python package index serves at execution time. This prevents reproducible dependency resolution and implicitly trusts future upstream releases, mirrors, and local package-index configuration.

The risk is elevated for dependencies that may include native components, such as lxml, because package installation and later import can execute package-controlled code with the privileges of the invoking process.

No evidence indicates that the named packages are currently malicious. The vulnerability is the unsafe, mutable dependency-resolution process.

Attack Path

  1. A user invokes the skill in an environment where one or more dependencies are absent.
  2. The agent follows the documented instruction and runs the unpinned pip install command.
  3. An attacker has compromised a future upstream release, a configured package mirror, or the package-index resolution path.
  4. pip downloads and installs the attacker-controlled distribution.
  5. Installation hooks or subsequent imports execute attacker-controlled code under the invoking user's account.

Impact Assessment

Successful exploitation could provide arbitrary code execution with the privileges of the user running pip or the conversion scripts. Accessible scope could include project files, EPUB inputs and outputs, environment variables, user-readable files, ...[truncated 146 chars]

Remediation
View remediation

Remediation Suggestions

  • Define reviewed, exact dependency versions in a requirements or lock file.
  • Record and verify distribution hashes using pip install --require-hashes.
  • Install dependencies in a dedicated virtual environment with least privilege.
  • Document and enforce the expected trusted package index rather than inheriting arbitrary index configuration.
  • Use automated dependency review and vulnerability scanning before updating pinned versions.
  • Separate runtime dependencies from development-only packages such as pytest.

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/extract.py:396
Finding

Unbounded Processing of Attacker-Controlled EPUB Archive Members

Content
View full analysis

Vulnerability Details

File Location: scripts/extract.py:396-423
Vulnerability Type: Resource-exhaustion vulnerability in archive processing
Risk Level: Medium

Vulnerable Code

python
for href in html_spine:
    soup = BeautifulSoup(zf.read(href).decode("utf-8", errors="ignore"), "lxml")
    page_type = classify_page(soup, href)
    if page_type == "section_index":
        heading = soup.find(re.compile("^h[1-3]$"))
        if heading:
            current_section = norm_text(heading.get_text(" "))
        continue
    if page_type != "article":
        continue
    paragraphs = get_translatable_paragraphs(soup)
    if not paragraphs:
        continue
    title = extract_title(soup)
    num = len(articles) + 1
    section = infer_section(soup, href, current_section)
    plain_text = "\n\n".join(paragraphs)[:8000]
    image_filename = copy_first_image(zf, soup, href, names, summary_dir, num, title)

Related unbounded image-member processing occurs in scripts/extract.py:368-379:

python
image_path = resolve_zip_path(href, src, names)
if not image_path or image_path in seen:
    continue
seen.add(image_path)
ext = Path(image_path).suffix.lower()
if ext not in IMAGE_EXTS:
    ext = ".jpg"
out_name = f"{num:02d} {safe_filename(title)}{ext}"
out_path = summary_dir / out_name
out_path.write_bytes(zf.read(image_path))

Assembly also reads every member fully in scripts/assemble.py:600-612:

python
for name in names:
    if name == "mimetype":
        continue
    data = zin.read(name)
    if is_ad_page(name, data):
        continue
    if name in article_by_href:
        data = process_article_html(data, article_by_href[name])
    elif name.lower().endswith(HTML_EXTS):
        data = process_other_html(data, articles)
    info = zin.getinfo(name)
    zout.writestr(info, data)

Technical Analysis

EPUB files are ZIP a ...[truncated 2074 chars]

Remediation
View remediation

Remediation Suggestions

  • Inspect every ZipInfo entry before reading archive data.
  • Reject archives exceeding configured limits for member count, individual uncompressed size, and cumulative uncompressed size.
  • Reject suspicious compression ratios and encrypted members.
  • Set explicit maximum sizes for HTML/XHTML, XML, CSS, fonts, and image resources.
  • Read and copy large resources incrementally where practical instead of using ZipFile.read().
  • Apply parsing time and memory limits by running untrusted EPUB processing in a constrained subprocess or sandbox.
  • Enforce disk quotas and verify available storage before writing summaries or the output EPUB.
  • Add tests covering ZIP bombs, oversized members, excessive member counts, and parser-complexity inputs.
  • Apply equivalent validation in both extraction and assembly so a crafted source cannot bypass limits by entering at a later workflow stage.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (18)

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description covers a multi-stage EPUB bilingual-production pipeline: extraction, translation/summary alignment support, and final EPUB assembly with preserved structure and a report. The supplied code chunk implements only the extraction stage. It reads an EPUB, finds the OPF/spine, filters HTML spine items, classifies pages, extracts translatable paragraphs, infers titles/sections, copies one image per article, and writes extraction.json. It also includes math text normalization for inline equation titles. There is no code for translation orchestration, summary filling, rebuilding EPUBs, preserving CSS/TOC/index pages in an output EPUB, or writing report.txt. Because the actual code is a narrower subset of the declared end-to-end behavior, the description does not accurately represent what this chunk itself actually does.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The declared description presents a full EPUB transformation pipeline: extracting content from EPUBs, coordinating translation/summaries, and rebuilding bilingual EPUBs. The actual code chunk does none of that. It is only a test file validating behavior of a token-estimation utility operating on already-structured JSON payloads. While token estimation could be a supporting component in a larger EPUB translation workflow, this code’s primary purpose is materially narrower and different from the declared end-to-end EPUB processing/assembly capability.

Content

No source excerpt is available for this finding.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 15)May include surrounding context.

md
4. Assemble the bilingual EPUB, summary text files, and `report.txt` with `scripts/assemble.py`.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 52)May include surrounding context.

md
4. Assemble the bilingual EPUB, summary text files, and `report.txt` with `scripts/assemble.py`.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 237)May include surrounding context.

md
4. Assemble the bilingual EPUB, summary text files, and `report.txt` with `scripts/assemble.py`.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 240)May include surrounding context.

md
4. Assemble the bilingual EPUB, summary text files, and `report.txt` with `scripts/assemble.py`.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 254)May include surrounding context.

md
4. Assemble the bilingual EPUB, summary text files, and `report.txt` with `scripts/assemble.py`.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 293)May include surrounding context.

md
4. Assemble the bilingual EPUB, summary text files, and `report.txt` with `scripts/assemble.py`.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
70% confidence
Finding

Without declared permissions the skill's intent is opaque and cannot be validated.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.md (reported line 93)May include surrounding context.

md
The user's initial request to translate or convert an EPUB is not enough to bypass this gate. The confirmation must happen after the concrete estimate report is shown.

If the user only asks for token estimation and does not ask to translate or convert, run extraction if needed, show the estimate, and stop without asking for translation confirmation.

The estimator is a lightweight character heuristic, not a tokenizer-exact counter. It estimates only translation-stage model usage:

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

This markdown file includes example configuration that specifies target_language as Simplified Chinese, and no nearby text states that language selection is user-driven. Because SQP-3 applies to all file types and forbids forcing a specific language without user opt-in, this is a natural-language locale policy concern.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The manifest example again sets target_language to Simplified Chinese, which can be read as a default enforced locale rather than a user-selected parameter. The document does not explicitly justify this locale constraint as region-specific, so it risks violating the language-choice policy.

Content

No source excerpt is available for this finding.

Ae4

Medium
Category
analysis-evasion
Confidence
80% confidence
Finding

Suspicious Unicode normalization or mixed-script content

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

This Python file contains a hard-coded Chinese summary template string, which makes generated summaries use a specific language regardless of the user's requested locale. That is a natural-language locale policy issue because the script does not offer a language choice or document that the tool is intentionally China-specific.

Content

No source excerpt is available for this finding.

Ae4

Medium
Category
analysis-evasion
Confidence
80% confidence
Finding

Suspicious Unicode normalization or mixed-script content

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The script unconditionally deletes the output_dir/summary directory with shutil.rmtree() before recreating it. In an agent setting where output_dir may be supplied dynamically, a mistaken or attacker-influenced path can cause unintended data loss by wiping an existing directory tree without warning, confirmation, or path safety checks.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The CLI sets --target-language to Chinese by default, which imposes a specific language choice even when the user does not explicitly request it. This is a natural-language locale policy concern because the skill does not offer neutral default behavior or require user opt-in for the language selection.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

The script embeds special token multipliers for specific target languages such as Chinese, Japanese, Korean, Spanish, and others, while all unspecified languages fall back to a default heuristic. This creates language-dependent behavior in the skill without any explicit user choice about locale handling or justification for why certain languages receive different treatment.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.dynamic_code_execution

Dynamic code execution detected.

Critical
Code
suspicious.dynamic_code_execution
Location
scripts/assemble.py:391

Dynamic code execution detected.

Critical
Code
suspicious.dynamic_code_execution
Location
scripts/patch_bilingual_math.py:11

Dynamic code execution detected.

Critical
Code
suspicious.dynamic_code_execution
Location
tests/test_assemble_order.py:12

Dynamic code execution detected.

Critical
Code
suspicious.dynamic_code_execution
Location
tests/test_estimate_tokens.py:13

Dynamic code execution detected.

Critical
Code
suspicious.dynamic_code_execution
Location
tests/test_extract_feynman_style.py:11