Back to skill

Security audit

Ingest Paper

Security checks for vulnerabilities and agentic risk

Overview

The skill does what it says by storing research materials in a Gitea knowledge base, but it uses over-privileged Gitea admin credentials over plaintext HTTP and has broad file-read and persistence behaviors that need review.

Install only in a controlled environment after changing GITEA_URL to HTTPS, replacing the site-admin token with the narrowest possible service credential, pinning dependencies in a virtual environment, and deciding whether remote logs and uploaded source files are acceptable for your data.

Vulnerability Patterns
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
Findings (4)

T09 · Insecure Skill Coding Practices

Error
Location
env-example.txt:4
Finding

Site-Administrator Token and Private Repository Content Transmitted over Plaintext HTTP

Content
View full analysis
Remediation
View remediation

T05 · Unauthorized Access and Privilege Escalation

Error
Location
scripts/gitea_api.py:78
Finding

Knowledge-Base Ingestion Requires Excessive Site-Administrator Privileges

Content
View full analysis
bool: code, _ = _api("GET", "/admin/users", params={"limit": 1}) return code == 200 ``` ```python def create_repo_for_user(username: str, name: str, description: str) -> dict: if repo_exists(username, name): return { "created": False, "html_url": f"{GITEA_URL}/{username}/{name}", } code, data = _api("POST", f"/admin/users/{username}/repos", json_body={ "name": name, "private": True, "description": description, "auto_init": True, "default_branch": "main", }) ``` ### Technical Analysis The declared ingestion functionality needs to read and update a registered user's `paper-kb` repository. That operation does not inherently require the ability to administer all users or create repositories through `/admin/users/{username}/repos`. The shared API module nevertheless expects a site-administrator token and includes administrative user-repository provisioning operations. This combines routine document ingestion with privileged account administration in the same credential and code boundary. The capability appears related to initialization rather than normal ingestion. Keeping it in the ingestion Skill exceeds the minimum permissions necessary for saving documents, updating indexes, and reading the registered user's catalog. ### Attack Path 1. An attacker obtains the Skill's token through plaintext network interception, local `.env` access, logs, backups, or another process running under the same account. 2. The attacker authenticates directly to the Gitea API with the token. 3. The attacker invokes site-administration ...[truncated 972 chars]
Remediation
View remediation

T08 · Insecure Dependencies

Warning
Location
requirements.txt:1
Finding

Unpinned Dependencies Are Installed Directly into the System Python Environment

Content
View full analysis
=2.28 python-dotenv>=1.0 pymupdf>=1.24 python-docx>=1.1 openpyxl>=3.1 ``` ```bash echo "[1/3] Installing Python dependencies..." pip3 install -r requirements.txt --break-system-packages 2>/dev/null || pip3 install -r requirements.txt ``` ### Technical Analysis All dependencies use open-ended lower bounds rather than exact versions. A setup run can therefore install any future version satisfying the constraint. The dependency graph is not protected by hashes, so package artifacts are not verified against an audited lock file. The setup script also passes `--break-system-packages`, explicitly allowing modification of the host's externally managed Python installation. If the command runs with elevated privileges, package installation code may execute with those privileges and alter packages used by unrelated applications. No evidence of an intentionally malicious package or typosquatted name was found. The risk arises from unsafe supply-chain controls and system-wide installation behavior. ### Attack Path 1. A dependency or transitive dependency publishes a compromised future release that still satisfies the lower-bound constraint. 2. An operator runs `setup.sh`. 3. `pip` resolves and downloads the compromised release because no exact version or artifact hash is required. 4. Package build or installation logic executes in the installer context. 5. The compromised package gains the permissions of the account running setup and remains available in the system Python environment. 6. Other Python applications may subsequently import the modified dependency, extending the impact beyond this Skill. ### Impact Assessment Potential impact includes: - Arbitrary code execution during installation or import. - Compromise of the ...[truncated 404 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/save_paper.py:66
Finding

Unrestricted Local File Paths Can Be Read and Uploaded to Gitea

Content
View full analysis
Remediation
View remediation
Path: candidate = Path(value) if candidate.is_symlink(): raise ValueError("Symbolic links are not allowed") resolved = candidate.resolve(strict=True) if WORKSPACE not in resolved.parents: raise ValueError("Input file is outside the approved workspace") if not resolved.is_file(): raise ValueError("Input must be a regular file") return resolved ``` 3. Use a separate workspace for each user or request to prevent cross-user temporary-file access. 4. Create temporary directories with restrictive permissions such as mode `0700`. 5. Reject symlinks and revalidate immediately before opening the file to reduce time-of-check/time-of-use races. 6. Enforce maximum file sizes before reading content into memory or uploading it. 7. Validate actual file signatures rather than trusting filename extensions. 8. Limit binary upload arguments to known files produced by the current ingestion operation. 9. Pass opaque file handles or generated workspace identifiers between processing stages instead of arbitrary filesystem paths. 10. Run the Skill under a dedicated operating-system account with no access to unrelated credentials or user files. ]]>
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (47)

Credential Access

High
Category
Privilege Escalation
Confidence
91% confidence
Finding

The example configuration explicitly instructs users to copy the file to .env and populate a site administrator token, creating a workflow centered on high-privilege credentials. Combined with the hardcoded external Gitea endpoint and the requirement that the token be a site admin token, this increases the risk of credential exposure, over-privileged deployment, and compromise of the Gitea instance if the .env file is mishandled, committed, or accessed by the skill runtime.

Content

Scanner excerpt · env-example.txt (reported line 2)May include surrounding context.

text
# paper-kb / ingest_paper 环境配置
# 复制本文件为 .env 并填入真实值

# Gitea 服务器地址(末尾不要带斜杠)
GITEA_URL=http://43.134.182.170:3000

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/gitea_api.py (reported line 6)May include surrounding context.

python
gitea_api.py — paper-kb 系统的 Gitea API 封装(共用模块)

职责:
  - 读取 .env 配置(GITEA_URL / GITEA_ADMIN_TOKEN)
  - 封装所有 Gitea REST API 调用
  - 提供 users.json(用户映射表)的读写,带并发冲突重试

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/gitea_api.py (reported line 27)May include surrounding context.

python
gitea_api.py — paper-kb 系统的 Gitea API 封装(共用模块)

职责:
  - 读取 .env 配置(GITEA_URL / GITEA_ADMIN_TOKEN)
  - 封装所有 Gitea REST API 调用
  - 提供 users.json(用户映射表)的读写,带并发冲突重试

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · setup.sh (reported line 9)May include surrounding context.

sh
gitea_api.py — paper-kb 系统的 Gitea API 封装(共用模块)

职责:
  - 读取 .env 配置(GITEA_URL / GITEA_ADMIN_TOKEN)
  - 封装所有 Gitea REST API 调用
  - 提供 users.json(用户映射表)的读写,带并发冲突重试

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · setup.sh (reported line 10)May include surrounding context.

sh
gitea_api.py — paper-kb 系统的 Gitea API 封装(共用模块)

职责:
  - 读取 .env 配置(GITEA_URL / GITEA_ADMIN_TOKEN)
  - 封装所有 Gitea REST API 调用
  - 提供 users.json(用户映射表)的读写,带并发冲突重试

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · setup.sh (reported line 11)May include surrounding context.

sh
gitea_api.py — paper-kb 系统的 Gitea API 封装(共用模块)

职责:
  - 读取 .env 配置(GITEA_URL / GITEA_ADMIN_TOKEN)
  - 封装所有 Gitea REST API 调用
  - 提供 users.json(用户映射表)的读写,带并发冲突重试

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · setup.sh (reported line 14)May include surrounding context.

sh
gitea_api.py — paper-kb 系统的 Gitea API 封装(共用模块)

职责:
  - 读取 .env 配置(GITEA_URL / GITEA_ADMIN_TOKEN)
  - 封装所有 Gitea REST API 调用
  - 提供 users.json(用户映射表)的读写,带并发冲突重试

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/gitea_api.py (reported line 33)May include surrounding context.

python
except ImportError:
        return
    here = Path(__file__).resolve().parent
    for candidate in (here / ".env", here.parent / ".env"):
        if candidate.exists():
            load_dotenv(candidate)
            return

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill is designed to send user-supplied documents and derived metadata into external systems, specifically a Gitea knowledge base and optionally Feishu, but the description does not present a clear upfront disclosure or consent boundary before storage/synchronization. This creates a real privacy and data-governance risk because users may submit sensitive research notes, meeting minutes, or internal files without realizing they will be persisted and shared to third-party services.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The instruction '全部中文' forces all generated output into Chinese regardless of user preference. This is a language policy constraint and no opt-in, language choice, or region-specific justification is provided in the file.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
86% confidence
Finding

The manifest contains a Chinese-only natural-language description, which may imply a fixed language/locale without user opt-in. The file does not document that the skill is intended only for Chinese-speaking users or that language is configurable.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This code file contains natural-language documentation entirely in Chinese, including the module description and usage/output explanation. The file does not offer a language choice or document that the skill is intentionally limited to a Chinese-speaking context, which creates a locale policy violation under the rule for forced language without opt-in.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
82% confidence
Finding

The module reads GITEA_ADMIN_TOKEN from the environment and places it into the Authorization header for all API calls. Access to sensitive credentials and authenticated transmission to a remote service are safety-relevant operations, but the code provides no user-facing notice that an admin token is required and will be used.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
86% confidence
Finding

This code automatically creates the system-config repository and initializes users.json, which are file-writing operations that affect persistent user/system data. While the docstrings describe the behavior for developers, there is no user-facing disclosure, confirmation prompt, or visible logging in this code to warn that remote repository state will be created or modified.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The write_users function performs persistent remote file updates with retry logic, changing the shared users.json mapping in a repository. Although the function is documented for developers, the code contains no confirmation, print/log disclosure, or other user-facing warning that shared configuration data will be modified.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The module-level documentation and generated content strings are written entirely in Chinese, including headings such as knowledge-base index and operation log, with no indication that users may choose another language. This can violate a language/locale policy when the skill implicitly forces one locale without opt-in or clear justification.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

This function writes catalog data to the remote repository via g.put_file, but the code provides no confirmation prompt, print/log disclosure to the user, or comment/docstring warning about modifying repository contents. Because this is a file write operation that changes persisted user data, the file lacks an explicit user disclosure in the code shown.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The function rebuilds and writes index.md to the repository, replacing existing content, yet there is no confirmation prompt or user-facing warning in the code. This is a persistent file modification affecting user-managed knowledge-base structure, so some explicit disclosure is expected.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This code automatically persists operation descriptions and user queries to log.md without any visible consent, redaction, or sensitivity filtering. Because queries can contain proprietary research topics, credentials pasted by mistake, or personal data, the logging creates a durable disclosure surface inside the repository and can expose sensitive information to anyone with repo access.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

append_query_log records raw natural-language questions into a persistent repository log, truncating only by length and not by sensitivity. In a knowledge-base skill, user queries are especially likely to include confidential research questions, internal project names, URLs, or accidental secrets, so this behavior materially increases data leakage risk.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The script is presented as a read-only knowledge-base reader, but the --list flow can perform a write by appending user-supplied queries to a log. This mismatch is security-relevant because callers, reviewers, or higher-level agents may grant broader trust or permissions based on the read-only description, causing unexpected data retention and side effects.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

When --log_question is provided, the script writes user-controlled content to persistent storage without any warning or consent mechanism in this file. In an agent setting, this can silently capture sensitive prompts, proprietary data, or personal information that users may assume is only used transiently for retrieval.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The help text frames logging as a convenience while explicitly ensuring the user's query will be recorded, which increases the risk of covert or normalized surveillance behavior. In a knowledge-base skill, users are likely to submit research questions that may contain confidential project details, making undisclosed permanent logging more dangerous in context.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The natural-language docstring and user messages are entirely in Chinese, including usage guidance and error text, which implies a fixed language experience. There is no indication that users can choose another language or that the locale restriction is intentional and documented as a region-specific constraint.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The script extracts full document contents and writes them to a predictable sidecar file on disk (.extracted.txt) without any consent prompt, retention control, or warning to the caller. In a document-processing skill, this can expose sensitive material through leftover files, broader filesystem access by other components, backup/sync systems, or later reuse beyond the user's expectation.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.