Back to skill

Security audit

Knowledge Importer

Security checks for vulnerabilities and agentic risk

Overview

This skill appears to do its advertised document-import job, but it needs review because it can automatically upload extracted document images over configurable/plain HTTP and leave extracted images in temporary folders.

Install only if you are comfortable with converted documents being saved into the configured Obsidian path and embedded images being uploaded to the configured image host. Use a trusted HTTPS image host, prefer an isolated virtual environment, avoid confidential documents until uploads are made opt-in or local-only, and clean temporary asset folders after use.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (3)

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/import_doc.py:136
Finding

Automatic transmission of extracted document images over plaintext HTTP

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/import_doc.py:283
Finding

Sensitive extracted images are retained in predictable shared temporary directories

Content
View full analysis
Remediation
View remediation

T08 · Insecure Dependencies

Note
Location
README.md:19
Finding

Third-party Python dependencies are installed without version or integrity pinning

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Taint TrackingDirect Taint Flow, Variable-Mediated Taint Flow, Credential Exfiltration Chain
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (14)

Tainted flow: 'DUFs_CONFIG' from os.getenv (line 31, credential/environment) → urllib.request.urlopen (network output)

Critical
Category
Data Flow
Confidence
94% confidence
Finding

The upload target is derived from configuration that falls back to environment variables, and the script blindly performs HTTP PUT requests to that URL with extracted document images. If an attacker can influence the environment or config, this becomes an SSRF/data-exfiltration path that transmits document contents to an arbitrary host, which is especially sensitive because the skill processes knowledge-base documents that may contain proprietary or personal data.

Content

Scanner excerpt · scripts/import_doc.py (reported line 151)May include surrounding context.

python
headers={'Content-Type': 'application/octet-stream'},
            )

            with urlopen(req, timeout=DUFs_CONFIG["timeout"]) as response:
                if response.status in (200, 201):
                    return url

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The README prominently advertises automatic upload of extracted document images to an external image host, but does not clearly warn that potentially sensitive document content may be transmitted off-device. In the context of a knowledge-base importer handling Word, Excel, PPT, and PDF files, this can expose confidential screenshots, diagrams, embedded scans, or other sensitive content to third-party infrastructure over the network.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
75% confidence
Finding

Docker image references without a specific tag (:latest is implicit) or digest (@sha256:...) can be silently replaced by a malicious image.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
75% confidence
Finding

Docker image references without a specific tag (:latest is implicit) or digest (@sha256:...) can be silently replaced by a malicious image.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
92% confidence
Finding

The skill documents capabilities that read environment variables, read/write local files, and perform network uploads, but it does not declare any tool scope or permissions boundaries. That makes the skill easier to invoke with broader-than-expected access and reduces transparency for users about sensitive operations such as writing into a local Obsidian vault and uploading extracted images to an external host.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill description explains document conversion but does not warn that imported content will be written into the local knowledge base and that images may be uploaded to an external image host. This creates a meaningful data exposure risk because users may provide sensitive documents without realizing content persistence and third-party transfer are part of the workflow.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The module docstring, warnings, help text, and console output are all written in Chinese, which effectively forces a specific language for users. There is no opt-in, alternate locale, or documentation that this tool is intentionally limited to a Chinese-speaking environment.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The skill uploads extracted images to an external host automatically and does not present an explicit warning or consent step about transmitting document-derived data off-device. In the context of a knowledge-import skill, this materially increases the risk of accidental leakage of confidential document content, including screenshots, scans, and embedded diagrams.

Content

No source excerpt is available for this finding.

Tainted flow: 'img_path' from os.getenv (line 245, credential/environment) → open (file write)

Medium
Category
Data Flow
Confidence
65% confidence
Finding

Data from a source is assigned to a variable that is later passed to a sink, creating a variable-mediated taint flow.

Content

Scanner excerpt · scripts/import_doc.py (reported line 217)May include surrounding context.

python
img_name = f"word_{timestamp}_{len(image_paths) + 1}.{ext}"
                img_path = os.path.join(output_dir, img_name)

                with open(img_path, 'wb') as f:
                    f.write(image)

                image_paths.append(img_path)

Tainted flow: 'img_path' from os.getenv (line 245, credential/environment) → open (file write)

Medium
Category
Data Flow
Confidence
65% confidence
Finding

Data from a source is assigned to a variable that is later passed to a sink, creating a variable-mediated taint flow.

Content

Scanner excerpt · scripts/import_doc.py (reported line 247)May include surrounding context.

python
img_name = f"word_{timestamp}_{len(image_paths) + 1}.{ext}"
                img_path = os.path.join(output_dir, img_name)

                with open(img_path, 'wb') as f:
                    f.write(image)

                image_paths.append(img_path)

Tainted flow: 'img_path' from os.getenv (line 245, credential/environment) → open (file write)

Medium
Category
Data Flow
Confidence
65% confidence
Finding

Data from a source is assigned to a variable that is later passed to a sink, creating a variable-mediated taint flow.

Content

Scanner excerpt · scripts/import_doc.py (reported line 424)May include surrounding context.

python
img_name = f"pdf_img_p{page_num+1}_{img_index+1}.{image_ext}"
                img_path = os.path.join(output_dir, img_name)
                
                with open(img_path, "wb") as f:
                    f.write(image_bytes)
                
                image_paths.append(img_path)

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
92% confidence
Finding

The operational documentation and user-facing instructions are presented entirely in Chinese, which can constitute a language-policy issue when no opt-in, alternative language, or locale justification is provided. This may restrict accessibility for users expecting language choice.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
91% confidence
Finding

Natural-language instructions in this file consistently require Chinese comprehension, but the skill does not state that Chinese is optional or that the skill is intentionally limited to a Chinese-speaking context. Under the policy, forcing a specific language without user opt-in is a language/locale policy concern.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
85% confidence
Finding

The script creates directories and writes output files into the configured knowledge-base path, which can modify user-managed content. While this is part of the tool's purpose, the file lacks a clear warning that running it will create and save Markdown files under the target knowledge-base location by default.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.