Back to skill

Security audit

pdf-translate-skill

Security checks for vulnerabilities and agentic risk

Overview

This PDF translation skill is useful and mostly transparent, but its arXiv workflow handles remote archives and LaTeX source in ways that can affect local files without strong containment.

Install only if you are comfortable running document-processing tools on files you trust, preferably inside a disposable container or isolated workspace. Be especially careful with arXiv mode: remote source archives should be treated as untrusted, safely extracted, and compiled without access to your home directory, secrets, or important project files.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (3)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/download_arxiv_source.py:64
Finding

Unsafe Extraction of Remotely Supplied Tar Archives

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
scripts/compile_latex.py:131
Finding

Untrusted Downloaded LaTeX Is Compiled Without Isolation

Content
View full analysis
Remediation
View remediation

T08 · Insecure Dependencies

Warning
Location
SKILL.md:37
Finding

Unpinned Third-Party Dependency Installation

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (15)

Natural-Language Policy Violations

High
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill repeatedly mandates Chinese-specific output behavior, such as requiring the ctex package for Chinese support and later stating to always use ctex with XeLaTeX. This imposes a specific language/locale expectation rather than offering translation to a user-selected target language, which is a natural-language policy violation under the language-choice rule.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

High
Category
Not specified by scanner
Confidence
94% confidence
Finding

The translation approach and example convert English section titles and body text into Chinese, without framing Chinese as merely one possible target language. This creates a default locale/language bias instead of preserving user choice, which violates the language/locale policy for natural-language instructions.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The Local PDF Mode explicitly says 'Always use ctex package with XeLaTeX' and the template includes Chinese support unconditionally. That requirement forces a specific language/locale setup for every translation job, even when the user may want another language, which conflicts with the policy requiring language choice or justified constraint.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The skill instructs the agent to download remote arXiv source archives and extract/write files locally, but the description does not warn users that untrusted remote content will be fetched and stored on disk. In this context, the omission matters because arXiv source bundles can contain arbitrary TeX trees and auxiliary files, increasing the chance of users unknowingly authorizing risky file operations.

Content

No source excerpt is available for this finding.

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
70% confidence
Finding

Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.

Content

Scanner excerpt · SKILL.md (reported line 49)May include surrounding context.

md
# LaTeX (choose one)
# Ubuntu/Debian:
sudo apt install texlive-xetex texlive-lang-chinese

# macOS (Homebrew):
brew install mactex

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
70% confidence
Finding

Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.

Content

Scanner excerpt · references/troubleshooting.md (reported line 109)May include surrounding context.

  • Install TeX Live (Linux/Mac):
    bash
    # Ubuntu/Debian
    sudo apt install texlive-xetex texlive-lang-chinese
    
    # macOS (Homebrew)
    brew install mactex
    

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
93% confidence
Finding

This code invokes a full LaTeX engine on a user-supplied .tex file. Even though subprocess.run is used safely without shell=True, the dangerous part is that LaTeX itself can execute file I/O and, depending on engine configuration or packages, may permit command execution or arbitrary file reads from untrusted documents. In an agent skill context, compiling attacker-controlled TeX is risky because the skill is effectively handing untrusted input to a complex interpreter.

Content

Scanner excerpt · scripts/compile_latex.py (reported line 156)May include surrounding context.

python
print(f"\nPass {pass_num}/{passes}...")

        try:
            process = subprocess.run(
                cmd_base,
                capture_output=True,
                text=True,

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The script performs filesystem cleanup by deleting generated auxiliary files via unlink(), but there is no explicit user warning, confirmation prompt, or visible disclosure that these files will be removed. While cleanup may be convenient, it is still a destructive operation affecting user files in the output directory.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/download_arxiv_source.py (reported line 86)May include surrounding context.

python
url,
        ]

        result = subprocess.run(curl_cmd, capture_output=True, text=True, timeout=300)

        if result.returncode != 0:
            # Try wget as fallback

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/download_arxiv_source.py (reported line 92)May include surrounding context.

python
# Try wget as fallback
            print("curl failed, trying wget...")
            wget_cmd = ["wget", "-O", temp_path, "-e", "tries=3", url]
            result = subprocess.run(
                wget_cmd, capture_output=True, text=True, timeout=300
            )

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
99% confidence
Finding

The script extracts a remote tar.gz from arXiv directly with tar.extractall(output_path) without validating member paths. A crafted archive could contain absolute paths or '../' traversal entries and overwrite files outside the intended directory, which is especially dangerous because the archive comes from untrusted remote content.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
87% confidence
Finding

The README explicitly advertises downloading arXiv source files and processing document content, but it does not disclose that the skill may initiate network access or handle externally sourced data. In an agent setting, missing user-facing disclosure can lead to unexpected outbound requests and ingestion of untrusted content, increasing the chance of unsafe execution paths or privacy surprises even if the behavior is part of the intended functionality.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
89% confidence
Finding

As a markdown file, this guide should warn about behaviors that could affect user data or system integrity. The manual cleanup command uses a broad rm pattern that deletes files in the working directory, but the surrounding text does not caution users to verify the target directory or review matched files first.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
78% confidence
Finding

The module docstring frames the tool around XeLaTeX and Chinese/CJK support as the default behavior, which can be interpreted as a locale-specific bias without explicit user opt-in. The file does provide a pdfLaTeX alternative, but the natural-language description still privileges a specific language/locale context rather than presenting it as optional.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

The module docstring says the tool 'Downloads source TeX files from arXiv', which implies a narrower action focused on TeX sources. In practice, the code downloads the source archive and calls tar.extractall(output_path), writing all archive contents to the output directory, including non-TeX files.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.