Back to skill

Security audit

Scholar Search

Security checks for vulnerabilities and agentic risk

Overview

This academic search skill appears purpose-aligned, but its API-key setup uses unsafe command-line handling and persists the key in a local plaintext file.

Review before installing. Use a runtime secret or environment variable for `S2_API_KEY` instead of the documented command when possible, and do not paste untrusted strings into the API-key setup command. If you do persist the key, treat `scripts/.env` as a local plaintext secret and restrict access to it.

Vulnerability Patterns
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (3)

T09 · Insecure Skill Coding Practices

Error
Location
SKILL.md:17
Finding

Shell Injection Through Unsafe API-Key Interpolation

Content
View full analysis
" ``` ``` ### Technical Analysis The skill instructs the agent to interpolate an API key supplied through the conversation directly into a quoted shell command. Quotation alone does not safely encode an untrusted shell argument. A value containing a quotation mark followed by shell metacharacters can terminate the intended argument and introduce additional shell commands. The Python argument parser does not create this vulnerability by itself. The vulnerability occurs before Python starts, when a shell interprets the command assembled according to the skill instructions. For example, a malicious value shaped like the following could escape the quoted argument when substituted literally: ```text "; id > /tmp/injection-result; # ``` This issue is exploitable when the agent follows the documented workflow using a shell-based execution tool. ### Attack Path 1. An attacker presents a crafted string as a Semantic Scholar API key in the conversation. 2. The skill requires the agent to substitute that string into the documented `--api-key` command. 3. The agent sends the resulting command to a shell. 4. The injected quotation mark terminates the intended API-key argument. 5. Shell metacharacters introduce and execute an attacker-selected command. 6. The injected command runs with the same operating-system identity and permissions as the agent process. ### Impact Assessment Successful exploitation permits arbitrary command execution within the agent's local security context. The attacker could read files accessible to the agent, modify project content, steal environment var ...[truncated 246 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/set_s2_api_key.py:18
Finding

API Key Exposure Through Process Arguments and Plaintext Storage

Content
View full analysis
argparse.ArgumentParser: parser = argparse.ArgumentParser(description="Set or overwrite S2_API_KEY in .env") parser.add_argument("--api-key", required=True, help="Semantic Scholar API key") return parser ``` The stored value is subsequently loaded from the plaintext file: ```python def _resolve_s2_api_key(env_path: Path) -> Optional[str]: # Prefer the process environment variable so CI/platform deployments can inject it. env_value = (os.getenv("S2_API_KEY") or "").strip() if env_value: return env_value return _read_s2_api_key_from_dotenv(env_path) ``` ### Technical Analysis The configuration workflow passes the API key through the `--api-key` command-line option. Command-line arguments may be retained in shell history, captured by command logging or telemetry, and exposed through process-inspection facilities while the command is running. The script then writes the key unencrypted to `scripts/.env`. It uses `Path.write_text()` without explicitly enforcing owner-only permissions. The effective protection therefore depends on the existing file mode and process umask. If the file already has permissive permissions, rewriting it does not correct them. The repository's `.gitignore` excludes `.env` files, which reduces accidental Git commits but does not protect the local file from other users, processes, ...[truncated 1216 chars]
Remediation
View remediation

T08 · Insecure Dependencies

Note
Location
SKILL.md:11
Finding

Unpinned and Unverified Runtime Dependency Installation

Content
View full analysis
Remediation
View remediation
``` - Generate and enforce cryptographic hashes, using an installation command such as: ```bash python -m pip install --require-hashes -r requirements.txt ``` - Use a lock file generated by a dependency-management tool and review transitive dependencies. - Configure pip to use an approved HTTPS package index. - Install dependencies inside an isolated virtual environment with only the permissions required for the skill. - Add automated vulnerability and dependency-integrity scanning to the release process. - Periodically update pinned versions through a controlled review and testing workflow. ]]>
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (22)

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · .gitignore (reported line 1)May include surrounding context.

text
scripts/.env
**/.env
**/__pycache__/
*.pyc

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · .gitignore (reported line 2)May include surrounding context.

text
scripts/.env
**/.env
**/__pycache__/
*.pyc

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The declared purpose is academic search, but the skill also includes credential-management behavior that writes or overwrites API keys in local storage. That mismatch is security-relevant because users may consent to search functionality without realizing the skill can modify files and persist secrets, which broadens the trust boundary and raises the risk of credential leakage or unintended state changes.

Content

No source excerpt is available for this finding.

Hidden Instructions

High
Category
Prompt Injection
Confidence
60% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · references/arxiv-api-reference.md (reported line 1)May include surrounding context.

md
# Semantic Scholar API Reference (Paper Data + Author Data)

更新日期:2026-03-06  
调研范围:`Paper Data`、`Author Data`

Hidden Instructions

High
Category
Prompt Injection
Confidence
60% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · references/semantic-scholar-api-reference.md (reported line 1)May include surrounding context.

md
# Semantic Scholar API Reference (Paper Data + Author Data)

更新日期:2026-03-06  
调研范围:`Paper Data`、`Author Data`

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/scholar-search.py (reported line 333)May include surrounding context.

python
def build_parser() -> argparse.ArgumentParser:
    parser = argparse.ArgumentParser(description="Set or overwrite S2_API_KEY in .env")
    parser.add_argument("--api-key", required=True, help="Semantic Scholar API key")
    return parser

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/set_s2_api_key.py (reported line 30)May include surrounding context.

python
def build_parser() -> argparse.ArgumentParser:
    parser = argparse.ArgumentParser(description="Set or overwrite S2_API_KEY in .env")
    parser.add_argument("--api-key", required=True, help="Semantic Scholar API key")
    return parser

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · scripts/set_s2_api_key.py (reported line 39)May include surrounding context.

python
def build_parser() -> argparse.ArgumentParser:
    parser = argparse.ArgumentParser(description="Set or overwrite S2_API_KEY in .env")
    parser.add_argument("--api-key", required=True, help="Semantic Scholar API key")
    return parser

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
89% confidence
Finding

The skill advertises executable behavior that can read environment variables, write local files, and access the network, but it does not declare any explicit tool scope or permissions boundary. This increases the chance that an agent or user will invoke broader capabilities than expected, especially because the document also instructs persisting credentials to a local .env file.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The skill instructs writing a user-supplied Semantic Scholar API key to scripts/.env and explicitly allows overwriting an existing S2_API_KEY, but it does not provide a clear warning about local file modification, persistence, or the risks of storing secrets in plaintext. Persisting credentials this way can expose them to other tools, accidental commits, backups, logs, or later file reads within the workspace.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The workflow fixes 输出语言:中文 and states that in all other cases the default is Chinese. This imposes a specific language choice unless the query is purely English or explicitly requests Chinese, which is a language policy issue because it does not offer a neutral user choice or opt-in.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The file title and all instructional content are written exclusively in Chinese, which imposes a specific language on users. Under the policy criteria, a forced language choice without opt-in or documented regional justification is a natural-language policy violation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The entire reference document is written in Chinese, including headings, operational notes, and examples, with no indication that the user may choose another language or locale. This creates a language/locale policy issue because the skill documentation effectively mandates Chinese for comprehension and use.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · references/semantic-scholar-api-reference.md (reported line 7)May include surrounding context.

md
import requests

ARXIV_BASE_URL = "https://export.arxiv.org/api/query"
SEMANTIC_SCHOLAR_BASE_URL = "https://api.semanticscholar.org/graph/v1"
MISSING_VALUE = "缺失"
SUPPORTED_SOURCES = {"arxiv", "semantic_scholar"}
ARXIV_XML_NS = {

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · references/semantic-scholar-api-reference.md (reported line 8)May include surrounding context.

md
import requests

ARXIV_BASE_URL = "https://export.arxiv.org/api/query"
SEMANTIC_SCHOLAR_BASE_URL = "https://api.semanticscholar.org/graph/v1"
MISSING_VALUE = "缺失"
SUPPORTED_SOURCES = {"arxiv", "semantic_scholar"}
ARXIV_XML_NS = {

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · references/semantic-scholar-api-reference.md (reported line 9)May include surrounding context.

md
import requests

ARXIV_BASE_URL = "https://export.arxiv.org/api/query"
SEMANTIC_SCHOLAR_BASE_URL = "https://api.semanticscholar.org/graph/v1"
MISSING_VALUE = "缺失"
SUPPORTED_SOURCES = {"arxiv", "semantic_scholar"}
ARXIV_XML_NS = {

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · references/semantic-scholar-api-reference.md (reported line 15)May include surrounding context.

md
import requests

ARXIV_BASE_URL = "https://export.arxiv.org/api/query"
SEMANTIC_SCHOLAR_BASE_URL = "https://api.semanticscholar.org/graph/v1"
MISSING_VALUE = "缺失"
SUPPORTED_SOURCES = {"arxiv", "semantic_scholar"}
ARXIV_XML_NS = {

External Transmission

Medium
Category
Data Exfiltration
Confidence
60% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · scripts/scholar-search.py (reported line 18)May include surrounding context.

python
import requests

ARXIV_BASE_URL = "https://export.arxiv.org/api/query"
SEMANTIC_SCHOLAR_BASE_URL = "https://api.semanticscholar.org/graph/v1"
MISSING_VALUE = "缺失"
SUPPORTED_SOURCES = {"arxiv", "semantic_scholar"}
ARXIV_XML_NS = {

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The constant and multiple user-facing error messages in this file are written in Chinese, which forces a specific locale for outputs regardless of user preference. The file does not offer a language choice or document a justified region-specific constraint, so this is a natural-language policy violation under the language/locale rule.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Low
Category
Not specified by scanner
Confidence
80% confidence
Finding

The manifest describes a unified academic search skill over arXiv and Semantic Scholar, but does not mention local credential discovery or reading process environment/local secret files. While using an API key can support Semantic Scholar access, the added capability to inspect both environment variables and a colocated .env file expands the skill beyond pure search behavior and introduces access to local secret material.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
82% confidence
Finding

This code accesses the sensitive S2_API_KEY credential via os.getenv and also falls back to reading it from a local .env file. While the credential is used for the stated Semantic Scholar purpose, the file itself does not include a docstring, comment, or other user-facing warning that the skill will inspect environment or local secret material.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
97% confidence
Finding

The print statement emits a fixed Chinese-language status message (已写入) with no option to select language or locale. This is a natural-language policy issue because the skill imposes a specific language on all users without opt-in or justification.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.