T09 · Insecure Skill Coding Practices
- Location
env-example.txt:4- Finding
Site-Administrator Token and Private Repository Content Transmitted over Plaintext HTTP
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
The skill does what it says by storing research materials in a Gitea knowledge base, but it uses over-privileged Gitea admin credentials over plaintext HTTP and has broad file-read and persistence behaviors that need review.
Install only in a controlled environment after changing GITEA_URL to HTTPS, replacing the site-admin token with the narrowest possible service credential, pinning dependencies in a virtual environment, and deciding whether remote logs and uploaded source files are acceptable for your data.
env-example.txt:4Site-Administrator Token and Private Repository Content Transmitted over Plaintext HTTP
scripts/gitea_api.py:78Knowledge-Base Ingestion Requires Excessive Site-Administrator Privileges
requirements.txt:1Unpinned Dependencies Are Installed Directly into the System Python Environment
scripts/save_paper.py:66Unrestricted Local File Paths Can Be Read and Uploaded to Gitea
The example configuration explicitly instructs users to copy the file to .env and populate a site administrator token, creating a workflow centered on high-privilege credentials. Combined with the hardcoded external Gitea endpoint and the requirement that the token be a site admin token, this increases the risk of credential exposure, over-privileged deployment, and compromise of the Gitea instance if the .env file is mishandled, committed, or accessed by the skill runtime.
# paper-kb / ingest_paper 环境配置
# 复制本文件为 .env 并填入真实值
# Gitea 服务器地址(末尾不要带斜杠)
GITEA_URL=http://43.134.182.170:3000
Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.
gitea_api.py — paper-kb 系统的 Gitea API 封装(共用模块)
职责:
- 读取 .env 配置(GITEA_URL / GITEA_ADMIN_TOKEN)
- 封装所有 Gitea REST API 调用
- 提供 users.json(用户映射表)的读写,带并发冲突重试
Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.
gitea_api.py — paper-kb 系统的 Gitea API 封装(共用模块)
职责:
- 读取 .env 配置(GITEA_URL / GITEA_ADMIN_TOKEN)
- 封装所有 Gitea REST API 调用
- 提供 users.json(用户映射表)的读写,带并发冲突重试
Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.
gitea_api.py — paper-kb 系统的 Gitea API 封装(共用模块)
职责:
- 读取 .env 配置(GITEA_URL / GITEA_ADMIN_TOKEN)
- 封装所有 Gitea REST API 调用
- 提供 users.json(用户映射表)的读写,带并发冲突重试
Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.
gitea_api.py — paper-kb 系统的 Gitea API 封装(共用模块)
职责:
- 读取 .env 配置(GITEA_URL / GITEA_ADMIN_TOKEN)
- 封装所有 Gitea REST API 调用
- 提供 users.json(用户映射表)的读写,带并发冲突重试
Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.
gitea_api.py — paper-kb 系统的 Gitea API 封装(共用模块)
职责:
- 读取 .env 配置(GITEA_URL / GITEA_ADMIN_TOKEN)
- 封装所有 Gitea REST API 调用
- 提供 users.json(用户映射表)的读写,带并发冲突重试
Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.
gitea_api.py — paper-kb 系统的 Gitea API 封装(共用模块)
职责:
- 读取 .env 配置(GITEA_URL / GITEA_ADMIN_TOKEN)
- 封装所有 Gitea REST API 调用
- 提供 users.json(用户映射表)的读写,带并发冲突重试
Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.
except ImportError:
return
here = Path(__file__).resolve().parent
for candidate in (here / ".env", here.parent / ".env"):
if candidate.exists():
load_dotenv(candidate)
return
The skill is designed to send user-supplied documents and derived metadata into external systems, specifically a Gitea knowledge base and optionally Feishu, but the description does not present a clear upfront disclosure or consent boundary before storage/synchronization. This creates a real privacy and data-governance risk because users may submit sensitive research notes, meeting minutes, or internal files without realizing they will be persisted and shared to third-party services.
The instruction '全部中文' forces all generated output into Chinese regardless of user preference. This is a language policy constraint and no opt-in, language choice, or region-specific justification is provided in the file.
The manifest contains a Chinese-only natural-language description, which may imply a fixed language/locale without user opt-in. The file does not document that the skill is intended only for Chinese-speaking users or that language is configurable.
This code file contains natural-language documentation entirely in Chinese, including the module description and usage/output explanation. The file does not offer a language choice or document that the skill is intentionally limited to a Chinese-speaking context, which creates a locale policy violation under the rule for forced language without opt-in.
The module reads GITEA_ADMIN_TOKEN from the environment and places it into the Authorization header for all API calls. Access to sensitive credentials and authenticated transmission to a remote service are safety-relevant operations, but the code provides no user-facing notice that an admin token is required and will be used.
This code automatically creates the system-config repository and initializes users.json, which are file-writing operations that affect persistent user/system data. While the docstrings describe the behavior for developers, there is no user-facing disclosure, confirmation prompt, or visible logging in this code to warn that remote repository state will be created or modified.
The write_users function performs persistent remote file updates with retry logic, changing the shared users.json mapping in a repository. Although the function is documented for developers, the code contains no confirmation, print/log disclosure, or other user-facing warning that shared configuration data will be modified.
The module-level documentation and generated content strings are written entirely in Chinese, including headings such as knowledge-base index and operation log, with no indication that users may choose another language. This can violate a language/locale policy when the skill implicitly forces one locale without opt-in or clear justification.
This function writes catalog data to the remote repository via g.put_file, but the code provides no confirmation prompt, print/log disclosure to the user, or comment/docstring warning about modifying repository contents. Because this is a file write operation that changes persisted user data, the file lacks an explicit user disclosure in the code shown.
The function rebuilds and writes index.md to the repository, replacing existing content, yet there is no confirmation prompt or user-facing warning in the code. This is a persistent file modification affecting user-managed knowledge-base structure, so some explicit disclosure is expected.
This code automatically persists operation descriptions and user queries to log.md without any visible consent, redaction, or sensitivity filtering. Because queries can contain proprietary research topics, credentials pasted by mistake, or personal data, the logging creates a durable disclosure surface inside the repository and can expose sensitive information to anyone with repo access.
append_query_log records raw natural-language questions into a persistent repository log, truncating only by length and not by sensitivity. In a knowledge-base skill, user queries are especially likely to include confidential research questions, internal project names, URLs, or accidental secrets, so this behavior materially increases data leakage risk.
The script is presented as a read-only knowledge-base reader, but the --list flow can perform a write by appending user-supplied queries to a log. This mismatch is security-relevant because callers, reviewers, or higher-level agents may grant broader trust or permissions based on the read-only description, causing unexpected data retention and side effects.
When --log_question is provided, the script writes user-controlled content to persistent storage without any warning or consent mechanism in this file. In an agent setting, this can silently capture sensitive prompts, proprietary data, or personal information that users may assume is only used transiently for retrieval.
The help text frames logging as a convenience while explicitly ensuring the user's query will be recorded, which increases the risk of covert or normalized surveillance behavior. In a knowledge-base skill, users are likely to submit research questions that may contain confidential project details, making undisclosed permanent logging more dangerous in context.
The natural-language docstring and user messages are entirely in Chinese, including usage guidance and error text, which implies a fixed language experience. There is no indication that users can choose another language or that the locale restriction is intentional and documented as a region-specific constraint.
The script extracts full document contents and writes them to a predictable sidecar file on disk (.extracted.txt) without any consent prompt, retention control, or warning to the caller. In a document-processing skill, this can expose sensitive material through leftover files, broader filesystem access by other components, backup/sync systems, or later reuse beyond the user's expectation.
No suspicious patterns detected.