T09 · Insecure Skill Coding Practices
- Location
scripts/research.py:561- Finding
Semantic Scholar API Key Exposed Through Command-Line Arguments
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
The skill matches its academic-search purpose, but it needs review because it handles an optional API key unsafely and installs unpinned Python dependencies.
Review before installing. Use an isolated virtual environment, pin or lock dependencies first, avoid passing --semantic-api-key on the command line, and only use PDF download or output paths you explicitly intend to write to. The skill does not appear malicious, but its credential and dependency practices should be fixed before broad or shared use.
scripts/research.py:561Semantic Scholar API Key Exposed Through Command-Line Arguments
scripts/requirements.txt:4Unpinned and Unhashed Third-Party Dependencies
Referenced artifact was not completely inspected
python scripts/research.py multi "transformer attention" -n 10
Referenced artifact was not completely inspected
python scripts/research.py multi "transformer attention" -n 10
Referenced artifact was not completely inspected
python scripts/research.py multi "transformer attention" -n 10
Referenced artifact was not completely inspected
python scripts/research.py multi "transformer attention" -n 10
Referenced artifact was not completely inspected
python scripts/research.py multi "transformer attention" -n 10
Referenced artifact was not completely inspected
python scripts/research.py multi "transformer attention" -n 10
Referenced artifact was not completely inspected
python scripts/research.py multi "transformer attention" -n 10
Referenced artifact was not completely inspected
python scripts/research.py multi "transformer attention" -n 10
Referenced artifact was not completely inspected
python scripts/research.py multi "transformer attention" -n 10
Referenced artifact was not completely inspected
pip install -r scripts/requirements.txt
The skill advertises capabilities that imply network access and local file creation/output, but it does not declare any explicit tool scope or permission boundaries. This can cause an agent runtime to invoke broader-than-expected tools for searches, downloads, or file writes, increasing the chance of unintended data access or local side effects.
The trigger scope is broad enough to activate on generic academic-help requests, not just explicit literature search tasks. Overbroad activation can cause the agent to make unnecessary network requests, retrieve external content, or offer file-producing actions in contexts where the user did not intend to invoke this skill.
The visible natural-language instructions and descriptions are entirely in Chinese, with no indication that users may choose another language or that the skill is intentionally limited to a Chinese-speaking audience. This can constitute a language/locale policy issue when no opt-in or justification is provided.
The line explicitly advertises "全中文输出" (Chinese-only output). Per the policy, forcing a specific language without offering user choice or documenting a justified locale constraint is a natural-language policy violation.
The module docstring, CLI description, help text, status messages, and formatted output are written exclusively in Chinese, which imposes a specific language on users. The file does not offer an opt-in language selection or explain that the skill is intentionally limited to a Chinese-speaking audience.
The skill is described as a literature search and citation helper, but it also supports downloading remote PDF files and saving them locally. This expands the capability from metadata retrieval into file acquisition and local persistence, which can create security and governance risks such as unwanted storage of untrusted files, policy bypass, or abuse for bulk content collection.
The code performs direct retrieval of URLs from external metadata and writes the response body to local files. Even though the filenames are mostly fixed-format, this still introduces a file write/download primitive unrelated to the core citation-search use case and increases risk from storing malicious or unexpected content from third-party URLs.
The markdown documents PDF download and output file behavior but does not clearly warn that the skill may create local files or write to user-specified paths. This can lead to unexpected disk writes, clutter, or accidental overwriting when an agent follows the documented behavior automatically.
L057 states that the skill does not read or store the S2_API_KEY environment variable and that the key is only passed via CLI arguments. Later, L189 says S2_API_KEY is configured by the user in ~/.bashrc, which directly conflicts with the earlier instruction and creates intent/documentation divergence.
This requirements file contains natural-language comments in Chinese identifying the skill as a Chinese-language assistant. Under the stated policy, forcing a specific language without user opt-in can be a locale/language policy concern, and there is no indication here of optional language selection or region-specific justification.
Using a lower-bound-only dependency specifier for requests allows future installs to pull in any newer release, making builds non-reproducible and potentially introducing vulnerable or breaking versions from the supply chain. In a skill that performs network retrieval from multiple external sources, HTTP client security and consistent behavior are important, so unpinned versions increase exposure.
# 学术论文检索小助手 - Python 依赖
# 核心检索
requests>=2.31.0
urllib3>=2.0.0
# Semantic Scholar(用于 S2 API 调用)
The finding is valid because requests has known advisories and the manifest does not pin a specific version, so there is no assurance that deployments avoid affected releases. In a network-heavy academic search skill, HTTP client flaws can impact credential handling, request verification, redirects, or transport security when talking to external services.
An unpinned urllib3 version permits installation of arbitrary future compatible releases, increasing the risk of supply-chain instability and accidental inclusion of versions with security regressions. Because this skill depends on external HTTP APIs, urllib3 sits in a sensitive path for TLS, redirects, proxies, and response handling.
# 核心检索
requests>=2.31.0
urllib3>=2.0.0
# Semantic Scholar(用于 S2 API 调用)
semanticscholar>=0.8.0
This is a real supply-chain risk: urllib3 has multiple advisories and the file does not establish an exact safe version, so installations may resolve to affected releases. Given the skill’s reliance on outbound HTTP and possible redirect/proxy/compression handling, vulnerable urllib3 versions could materially affect confidentiality or availability.
Leaving semanticscholar unpinned reduces build reproducibility and can introduce unexpected API behavior or vulnerable transitive dependencies over time. While this library is less security-critical than the core HTTP stack, it still processes remote data and therefore should be version-controlled.
urllib3>=2.0.0
# Semantic Scholar(用于 S2 API 调用)
semanticscholar>=0.8.0
# arXiv 预印本搜索
arxiv>=2.0.0
No suspicious patterns detected.