T09 · Insecure Skill Coding Practices
- Location
scripts/import_doc.py:136- Finding
Automatic transmission of extracted document images over plaintext HTTP
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This skill appears to do its advertised document-import job, but it needs review because it can automatically upload extracted document images over configurable/plain HTTP and leave extracted images in temporary folders.
Install only if you are comfortable with converted documents being saved into the configured Obsidian path and embedded images being uploaded to the configured image host. Use a trusted HTTPS image host, prefer an isolated virtual environment, avoid confidential documents until uploads are made opt-in or local-only, and clean temporary asset folders after use.
scripts/import_doc.py:136Automatic transmission of extracted document images over plaintext HTTP
scripts/import_doc.py:283Sensitive extracted images are retained in predictable shared temporary directories
README.md:19Third-party Python dependencies are installed without version or integrity pinning
The upload target is derived from configuration that falls back to environment variables, and the script blindly performs HTTP PUT requests to that URL with extracted document images. If an attacker can influence the environment or config, this becomes an SSRF/data-exfiltration path that transmits document contents to an arbitrary host, which is especially sensitive because the skill processes knowledge-base documents that may contain proprietary or personal data.
headers={'Content-Type': 'application/octet-stream'},
)
with urlopen(req, timeout=DUFs_CONFIG["timeout"]) as response:
if response.status in (200, 201):
return url
The README prominently advertises automatic upload of extracted document images to an external image host, but does not clearly warn that potentially sensitive document content may be transmitted off-device. In the context of a knowledge-base importer handling Word, Excel, PPT, and PDF files, this can expose confidential screenshots, diagrams, embedded scans, or other sensitive content to third-party infrastructure over the network.
Docker image references without a specific tag (:latest is implicit) or digest (@sha256:...) can be silently replaced by a malicious image.
Docker image references without a specific tag (:latest is implicit) or digest (@sha256:...) can be silently replaced by a malicious image.
The skill documents capabilities that read environment variables, read/write local files, and perform network uploads, but it does not declare any tool scope or permissions boundaries. That makes the skill easier to invoke with broader-than-expected access and reduces transparency for users about sensitive operations such as writing into a local Obsidian vault and uploading extracted images to an external host.
The skill description explains document conversion but does not warn that imported content will be written into the local knowledge base and that images may be uploaded to an external image host. This creates a meaningful data exposure risk because users may provide sensitive documents without realizing content persistence and third-party transfer are part of the workflow.
The module docstring, warnings, help text, and console output are all written in Chinese, which effectively forces a specific language for users. There is no opt-in, alternate locale, or documentation that this tool is intentionally limited to a Chinese-speaking environment.
The skill uploads extracted images to an external host automatically and does not present an explicit warning or consent step about transmitting document-derived data off-device. In the context of a knowledge-import skill, this materially increases the risk of accidental leakage of confidential document content, including screenshots, scans, and embedded diagrams.
Data from a source is assigned to a variable that is later passed to a sink, creating a variable-mediated taint flow.
img_name = f"word_{timestamp}_{len(image_paths) + 1}.{ext}"
img_path = os.path.join(output_dir, img_name)
with open(img_path, 'wb') as f:
f.write(image)
image_paths.append(img_path)
Data from a source is assigned to a variable that is later passed to a sink, creating a variable-mediated taint flow.
img_name = f"word_{timestamp}_{len(image_paths) + 1}.{ext}"
img_path = os.path.join(output_dir, img_name)
with open(img_path, 'wb') as f:
f.write(image)
image_paths.append(img_path)
Data from a source is assigned to a variable that is later passed to a sink, creating a variable-mediated taint flow.
img_name = f"pdf_img_p{page_num+1}_{img_index+1}.{image_ext}"
img_path = os.path.join(output_dir, img_name)
with open(img_path, "wb") as f:
f.write(image_bytes)
image_paths.append(img_path)
The operational documentation and user-facing instructions are presented entirely in Chinese, which can constitute a language-policy issue when no opt-in, alternative language, or locale justification is provided. This may restrict accessibility for users expecting language choice.
Natural-language instructions in this file consistently require Chinese comprehension, but the skill does not state that Chinese is optional or that the skill is intentionally limited to a Chinese-speaking context. Under the policy, forcing a specific language without user opt-in is a language/locale policy concern.
The script creates directories and writes output files into the configured knowledge-base path, which can modify user-managed content. While this is part of the tool's purpose, the file lacks a clear warning that running it will create and save Markdown files under the target knowledge-base location by default.
No suspicious patterns detected.