Back to skill

Security audit

Query Papers

Security checks for vulnerabilities and agentic risk

Overview

This skill is a plausible Gitea knowledge-base query tool, but it asks for an overpowered administrator token and can expose or persist sensitive data in unsafe ways.

Only install this after replacing the site-admin token with least-privilege credentials, requiring HTTPS, removing shell interpolation of user questions, making query logging explicit and optional, and installing pinned dependencies in an isolated virtual environment.

Vulnerability Patterns
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
Findings (4)

T09 · Insecure Skill Coding Practices

Error
Location
SKILL.md:34
Finding

User-Controlled Query Interpolated into a Shell Command

Content
View full analysis
--list all --log_question "<用户的原始问题>" ``` The surrounding instructions explicitly require the command to include the user's original question: ```text **务必带上 --log_question 参数**:它会在读目录的同时把本次查询写入 log.md。 ``` ### Technical Analysis The Skill instructs the Agent to place the original, untrusted user question directly inside a shell command. Enclosing the value in double quotes is not sufficient shell escaping. Characters such as embedded double quotes, command substitutions (`$(...)`), backticks, and some shell expansions may alter command interpretation before Python receives the argument. Although `kb_read.py` processes the resulting argument through `argparse`, that protection only applies after the shell has parsed the command. An attacker can therefore target the shell invocation itself. Exploitability depends on whether the Skill runtime executes the documented command through a shell and whether it performs independent argument escaping. The Skill instructions do not require a shell-free invocation or define a safe encoding mechanism. ### Attack Path 1. An attacker asks a knowledge-base question containing shell metacharacters or command substitution. 2. The Agent follows `SKILL.md` and inserts the original question into the documented command. 3. The runtime passes the constructed command to a shell. 4. The shell interprets the attacker-controlled syntax before launching `kb_read.py`. 5. The injected command executes with the permissions of the Agent or Skill process. 6. The attacker may use that access to read `.env`, recover the Gitea administrator token, modify local files, or send data to an external destination. ### Impact Assessment Successful exploitation provides arbitrary command execution ...[truncated 392 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
scripts/gitea_api.py:34
Finding

Gitea Administrator Token Can Be Transmitted over Plaintext HTTP

Content
View full analysis
Remediation
View remediation

T05 · Unauthorized Access and Privilege Escalation

Error
Location
env-example.txt:7
Finding

Query Skill Requires Excessive Site-Administrator Privileges

Content
View full analysis
bool: """检测当前 token 的账号是否为 Gitea 站点管理员。 管理员才能访问 /admin/users 列表;403 即非管理员。 """ code, _ = _api("GET", "/admin/users", params={"limit": 1}) return code == 200 ``` It also contains administrative repository-creation functionality: ```python code, data = _api("POST", f"/admin/users/{username}/repos", json_body={ "name": name, "private": True, "description": description, "auto_init": True, "default_branch": "main", }) ``` ### Technical Analysis The declared purpose of this Skill is to query one registered user's knowledge repository and append a query entry to that repository's `log.md`. That workflow needs only narrowly scoped read access to the relevant repository and limited write access to one log file. A site-administrator token exceeds those requirements. The common API module additionally exposes administrative endpoints and repository-creation operations that are not necessary for the query workflow. Sharing a broadly privileged token between initialization, ingestion, and querying expands the consequences of vulnerabilities in any one component. This is a least-privilege violation even when the administrative methods are not directly invoked during an ordinary query, because the powerful credential remains available to the process ...[truncated 1108 chars]
Remediation
View remediation

T08 · Insecure Dependencies

Warning
Location
requirements.txt:1
Finding

Unpinned Dependencies Are Installed into the Host Python Environment

Content
View full analysis
=2.28 python-dotenv>=1.0 ``` ```bash pip3 install -r requirements.txt --break-system-packages 2>/dev/null || pip3 install -r requirements.txt ``` ### Technical Analysis The dependency constraints contain no upper bounds or exact versions. A future package release satisfying `>=` can therefore be installed without being reviewed with the Skill. The dependency file also lacks cryptographic hashes, so package artifacts are not locked to known contents. The setup script first attempts installation with `--break-system-packages`, intentionally bypassing protections for an externally managed Python installation. This can modify the host's shared Python environment and cause dependency conflicts with other applications. The package names are legitimate and no malicious dependency was identified in the audited files. The risk arises from non-reproducible future resolution, package-index compromise, account compromise, or an incompatible future release. ### Attack Path 1. The Skill is installed or reinstalled at a later date. 2. `pip` resolves the open-ended constraints against the configured package index. 3. A compromised, malicious, or incompatible later release satisfies the `>=` constraint. 4. `pip` downloads and installs that release, potentially into the shared system Python environment. 5. Package installation hooks or imported runtime code execute with the installer or Skill process's permissions. 6. A malicious dependency could read `.env`, recover the Gitea token, modify other Python applications, or alter query results. ### Impact Assessment Successful supply-chain compromise can execute code with the permissions of the user running `setup.sh` or the Skill. Because installation may affect the shared Python environme ...[truncated 232 chars]
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
Findings (30)

Credential Access

High
Category
Privilege Escalation
Confidence
88% confidence
Finding

The example environment file instructs operators to create a .env containing a site administrator token for a bot account. Storing or encouraging use of a high-privilege admin token in a local .env increases the risk of credential leakage through source control, backups, logs, or host compromise, and compromise of this token would grant broad administrative control over the Gitea instance.

Content

Scanner excerpt · env-example.txt (reported line 2)May include surrounding context.

text
# paper-kb / query_papers 环境配置
# 复制本文件为 .env 并填入真实值

# Gitea 服务器地址(末尾不要带斜杠)
GITEA_URL=http://43.134.182.170:3000

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/gitea_api.py (reported line 6)May include surrounding context.

python
gitea_api.py — paper-kb 系统的 Gitea API 封装(共用模块)

职责:
  - 读取 .env 配置(GITEA_URL / GITEA_ADMIN_TOKEN)
  - 封装所有 Gitea REST API 调用
  - 提供 users.json(用户映射表)的读写,带并发冲突重试

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/gitea_api.py (reported line 27)May include surrounding context.

python
gitea_api.py — paper-kb 系统的 Gitea API 封装(共用模块)

职责:
  - 读取 .env 配置(GITEA_URL / GITEA_ADMIN_TOKEN)
  - 封装所有 Gitea REST API 调用
  - 提供 users.json(用户映射表)的读写,带并发冲突重试

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/gitea_api.py (reported line 33)May include surrounding context.

python
except ImportError:
        return
    here = Path(__file__).resolve().parent
    for candidate in (here / ".env", here.parent / ".env"):
        if candidate.exists():
            load_dotenv(candidate)
            return

Credential Access

High
Category
Privilege Escalation
Confidence
96% confidence
Finding

The script explicitly checks for a .env file and instructs the user to fill in a real GITEA_ADMIN_TOKEN, which is a privileged credential request during setup. Requesting an admin token for a seemingly unrelated skill increases the likelihood of unnecessary credential exposure and abuse if the skill or its dependencies are malicious or compromised.

Content

Scanner excerpt · setup.sh (reported line 9)May include surrounding context.

sh
echo "[1/3] 安装 Python 依赖..."
pip3 install -r requirements.txt --break-system-packages 2>/dev/null || pip3 install -r requirements.txt

echo "[2/3] 检查 .env 配置..."
if [ ! -f .env ]; then
  cp env-example.txt .env
  echo "  已生成 .env,请编辑填入真实的 GITEA_ADMIN_TOKEN!"

Credential Access

High
Category
Privilege Escalation
Confidence
96% confidence
Finding

Copying an example environment file into .env as part of setup normalizes local secret handling and prepares the user to place a privileged Gitea admin token in a file in the working directory. In the context of an untrusted skill, this increases the blast radius because later scripts or dependencies may read, transmit, or misuse that credential.

Content

Scanner excerpt · setup.sh (reported line 10)May include surrounding context.

sh
pip3 install -r requirements.txt --break-system-packages 2>/dev/null || pip3 install -r requirements.txt

echo "[2/3] 检查 .env 配置..."
if [ ! -f .env ]; then
  cp env-example.txt .env
  echo "  已生成 .env,请编辑填入真实的 GITEA_ADMIN_TOKEN!"
else

Credential Access

High
Category
Privilege Escalation
Confidence
98% confidence
Finding

The message instructing the user to enter a real GITEA_ADMIN_TOKEN is the clearest indicator of credential solicitation. In this skill context, asking for an administrative token is disproportionate to the stated 'query_papers' purpose and therefore materially increases suspicion of over-privileged access and potential credential theft or misuse.

Content

Scanner excerpt · setup.sh (reported line 11)May include surrounding context.

sh
echo "[2/3] 检查 .env 配置..."
if [ ! -f .env ]; then
  cp env-example.txt .env
  echo "  已生成 .env,请编辑填入真实的 GITEA_ADMIN_TOKEN!"
else
  echo "  .env 已存在,跳过。"

Credential Access

High
Category
Privilege Escalation
Confidence
88% confidence
Finding

The script's logic to preserve an existing .env file indicates continued reliance on a local secret-bearing configuration, which can perpetuate storage of privileged credentials in the skill directory. This is less severe than the explicit token prompt, but still contributes to insecure secret handling patterns in an untrusted installation flow.

Content

Scanner excerpt · setup.sh (reported line 14)May include surrounding context.

sh
cp env-example.txt .env
  echo "  已生成 .env,请编辑填入真实的 GITEA_ADMIN_TOKEN!"
else
  echo "  .env 已存在,跳过。"
fi

echo "[3/3] 自检..."

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill requires passing the user's raw question to --log_question, which records it in log.md without any notice, consent, retention policy, or minimization. User queries can contain sensitive personal, medical, business, or research information, so silent persistence creates a privacy and data-handling risk beyond what is necessary to answer the request.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The instruction '全部中文' forces a specific language for all responses. This is a natural-language policy concern because it does not offer the user a language preference or explain a justified locale restriction.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
83% confidence
Finding

This manifest file describes the skill as a general knowledge-base query capability for "全类型" queries, but it does not define specific trigger phrases, invocation boundaries, or exclusions. In a manifest, such broad wording can make activation scope ambiguous and increase the chance of unintended invocation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The description explicitly says the skill generates '中文回答' (Chinese answers). This is a natural-language locale constraint presented as part of the skill behavior, but there is no indication that users can choose another language or opt in to Chinese-only output.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

This module reads GITEA_ADMIN_TOKEN from the environment and immediately builds Authorization headers for all outbound Gitea API requests. Although comments describe the mechanics, there is no user-facing warning, confirmation, or disclosure that administrator credentials will be accessed and used for network operations.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
84% confidence
Finding

The module can create repositories and write or update remote files such as users.json and binary content through Gitea API calls. While function names and internal comments describe behavior for developers, there is no user-facing warning, confirmation, or disclosure in this file that remote state will be modified.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The module docstring and generated content strings such as "知识库索引", "研究方向", and "操作日志" are hard-coded in Chinese, indicating the skill's outputs are language-constrained. There is no natural-language indication that users may choose another language or that this locale restriction is intentional and justified for a region-specific tool.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This module performs multiple remote file writes via g.put_file, including catalog updates, index regeneration, and log/query-log appends. The file contains no confirmation prompt, print/log disclosure, or comments warning that user queries and repository content will be persisted, so these safety-relevant write operations lack visible user disclosure within the code.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

append_query_log stores raw user questions into a persistent repository-backed log by default, which can capture secrets, personal data, research topics, or credentials users accidentally type. In this skill context, a knowledge-base repository may be shared or retained long-term, so natural-language queries become a realistic confidentiality and data-retention risk.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The tool is presented as a knowledge-base reader, but the --log_question option causes a write to log.md during a list operation. This violates least surprise and can silently persist user-provided content, creating privacy, compliance, and audit-integrity risks if callers assume the operation is read-only.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The help text explicitly encourages recording the user's query so it is always logged, normalizing retention of user-supplied data without clear necessity or consent. In a skill context, this increases the chance that sensitive prompts, internal data, or personal information will be stored persistently and exposed to later readers of the repository or logs.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The script persists a user's query text together with an account identifier to storage via append_log without any visible notice, consent, or minimization beyond truncation. Search queries can contain sensitive personal, proprietary, or security-relevant information, so silent retention increases privacy and data-governance risk if logs are later accessed, reused, or breached.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The setup script claims to install a 'query_papers' skill but instead tells the user to populate a GITEA_ADMIN_TOKEN and executes a knowledge-base reader self-test, indicating functionality that does not match the stated purpose. That mismatch is dangerous because it can socially engineer operators into supplying high-privilege credentials unrelated to the advertised capability, expanding trust beyond what a paper-query tool should need.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The script's user-facing strings are entirely in Chinese, including installation progress and credential setup instructions. This imposes a specific language on users without opt-in or any indication that the skill is intentionally limited to a Chinese-speaking context, which matches the language/locale policy violation criteria.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
92% confidence
Finding

The natural-language content of the file is entirely in Chinese and does not offer any language choice or indicate that the skill is region-specific. Under the stated policy, forcing a specific language without user opt-in is a locale/language policy violation.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
97% confidence
Finding

The dependency is specified with a minimum version only, which makes builds non-reproducible and can result in installing unexpected or vulnerable releases over time. For a security-sensitive skill, leaving versions unpinned reduces supply-chain assurance and makes it harder to verify whether known vulnerabilities in requests are avoided.

Content

Scanner excerpt · requirements.txt (reported line 1)May include surrounding context.

text
requests>=2.28
python-dotenv>=1.0

Unverifiable Dependency: requests has 16 known advisory(ies) (CVE-2014-1830 (Exposure of Sensitive Information to an Unauthorized Actor in Requests); CVE-2024-47081 (Requests vulnerable to .netrc credentials leak via malicious URLs); CVE-2024-35195 (Requests `Session` object does not verify requests after making first request wi) +13 more), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
90% confidence
Finding

Requests has multiple known advisories, and because the manifest does not pin a version, it is impossible to determine from this file whether deployments will install a fixed or affected release. The issue is primarily a supply-chain verification gap: the skill context does not increase severity by itself, but it does leave network-facing functionality potentially exposed if an unsafe version is resolved.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.