T09 · Insecure Skill Coding Practices
- Location
scripts/extract_hwp.py:165- Finding
Path Traversal Enables Arbitrary Relative File Overwrite
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This document-extraction skill appears purpose-related but needs review because it overstates PDF/OCR support and can execute local programs and write extracted contents to paths influenced by user input.
Install only if you trust the callers and can run it in a constrained workspace. Treat extracted document text as retained on disk, validate --id to a simple identifier, do not pass untrusted --venv paths, and do not rely on the advertised PDF/OCR behavior until it is implemented or the documentation is corrected.
scripts/extract_hwp.py:165Path Traversal Enables Arbitrary Relative File Overwrite
scripts/extract_hwp.py:64Untrusted Interpreter Path Allows Arbitrary Local Executable Launch
The skill description claims reliable PDF/scan extraction and OCR fallback, but the static finding indicates the implementation does not actually provide those capabilities and instead falls back differently. This mismatch is security-relevant because downstream agents may trust the advertised extraction guarantees, make decisions on incomplete or incorrect text, or skip safer handling paths under the false assumption that OCR/PDF parsing occurred.
The skill advertises behavior that involves shell execution and writing files, but it does not declare any tool scope or permission boundaries in the manifest. This creates an authorization and transparency gap: an agent or user may invoke the skill without understanding that it can execute commands and persist data to disk, increasing the risk of unsafe file modification or command execution in broader runtime contexts.
The operative usage instructions and behavior descriptions are provided in Korean only, with no indication that users may request another language. This can violate a language/locale policy when a skill forces a specific language without opt-in or documented justification.
subprocess module calls execute external commands. Without careful input validation, this enables command injection.
def run_cmd(cmd, timeout=30, env=None):
try:
p = subprocess.run(cmd, stdout=subprocess.PIPE, stderr=subprocess.PIPE, timeout=timeout, env=env)
return p.returncode, p.stdout.decode('utf-8', errors='replace'), p.stderr.decode('utf-8', errors='replace')
except subprocess.TimeoutExpired:
return -1, '', 'timeout'
The script writes extracted document text to a predictable local JSON file named from the provided record ID, even though the skill is presented as an extraction pipeline rather than a persistence/export tool. This can silently retain sensitive document contents on disk, increasing the risk of unintended disclosure to other users, processes, logs, backups, or later workflow steps.
Extracted document text is written to a local JSON file without any user-facing disclosure or consent mechanism. In the context of HWP/HWPX/PDF extraction, the content is likely to include sensitive personal, legal, or business data, so silent local retention materially increases confidentiality risk.
The markdown states that extracted output is saved to disk, but it does not warn users about this filesystem side effect or clarify where data will persist. In document-processing workflows, silent persistence can expose sensitive contents, create unintended data retention, or overwrite files in shared workspaces if callers assume the operation is read-only.
The top-level documentation states the priority includes 'ocr(not implemented)', while the manifest says the pipeline attempts OCR with safe fallbacks. The actual code never performs OCR and instead falls back from HWP/HWPX handling directly to the strings command.
The code writes a temporary Python script to disk, marks it executable, and runs it as a subprocess to invoke pyhwp. This subprocess execution is only described internally in a code comment, not disclosed to the user via runtime messaging or the top-level usage documentation.
No suspicious patterns detected.