Back to skill

Security audit

知网论文下载

Security checks for vulnerabilities and agentic risk

Overview

This CNKI downloader mostly matches its purpose, but it needs review because it asks agents to run outside the sandbox, stores login sessions, installs dependencies, and handles web/download paths too broadly.

Review before installing. Use this only with a CNKI or institution URL you trust, avoid running it with broad host privileges, and clear the saved browser_data profile when you no longer need the login session. The skill should be fixed to narrow sandbox permissions, validate CNKI/proxy hosts and download destinations, sanitize output filenames, and use deterministic dependency installation.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (4)

T01 · Skill Instruction Hijacking

Error
Location
SKILL.md:67
Finding

Skill Instructions Require Execution Outside the Sandbox with Elevated Permissions

Content
View full analysis
" cwd: ``` ### Technical Analysis The Skill explicitly instructs the Agent to obtain elevated permission and execute outside its sandbox. A visible browser may legitimately require narrowly scoped GUI access, but the instruction mandates a blanket sandbox escape rather than requesting only the minimum capabilities needed for Edge and network access. This changes the Agent's security posture when the Skill is loaded and breaks least-privilege boundaries. It also increases the impact of every other defect in the package, including unsafe file writes, attacker-controlled navigation, and dependency installation. ### Attack Path 1. A user activates the CNKI download Skill. 2. The Agent reads `SKILL.md` and treats its instructions as mandatory. 3. The Agent requests or obtains elevated, out-of-sandbox execution. 4. The Agent runs `npm install` and the bundled Node.js scripts with host-level access. 5. Any compromised dependency, malicious configuration, unsafe remote page, or future package update can act with the expanded host permissions. ### Impact Assessment Successful exploitation can provide the executed scripts or dependencies with access beyond the normal Agent sandbox. The precise scope depends on the host's approval mechanism and operating-system account, but it may include broader filesystem access, network access, GUI automation, persistent browser data, and modification of user-writable host files. No direct operating-system administrator escalation mechanism was f ...[truncated 143 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
scripts/cnki_pipeline.js:373
Finding

Unrestricted Entry, Detail, and Download URLs Permit Attacker-Directed Navigation and Requests

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
scripts/pdf_downloader.js:80
Finding

Page-Controlled Document Titles Permit Download Path Traversal

Content
View full analysis
:"/\\|?*\0'; ``` The remote page title is used directly as a path component: ```js // 5) 构造文件名(用详情页标题 + .pdf) const title = (await detailPage.title()).trim(); const savePath = path.join(saveDirAbs, `${title}.pdf`); // 6) 写文件 await fs.writeFile(savePath, body); return savePath; ``` ### Technical Analysis A remote page controls its document title. The title is passed directly to `path.join` without removing path separators, traversal components, control characters, reserved Windows names, or trailing dots and spaces. The defined `INVALID_FN_CHARS` constant is unused. There is also no post-join verification that the resolved destination remains inside `saveDirAbs`. On platforms where the supplied title contains recognized separators and traversal segments, a title such as `../../target` can cause `fs.writeFile` to target a path outside the intended `scripts/download` directory. An absolute or specially formed platform path may produce similar unsafe behavior depending on path semantics. ### Attack Path 1. The workflow opens a detail page selected from the result page. 2. An attacker controls that detail page or causes navigation to attacker-controlled content. 3. The attacker sets the page title to a traversal value containing `..` and platform path separators. 4. `detailPage.title()` returns the malicious title. 5. `path.join(saveDirAbs, title + '.pdf')` constructs a path outside the download directory. 6. `fs.writeFile` writes the downloaded response to the escaped path. 7. If the destination already exists and is writable, it is overwritten. This pat ...[truncated 861 chars]
Remediation
View remediation

T08 · Insecure Dependencies

Warning
Location
scripts/package-lock.json:34
Finding

Automatic Dependency Installation Uses a Non-Official Package Registry Mirror

Content
View full analysis
process.exit(0)).catch(e=>process.exit(1))"`,exit 0 即通过。Node 会沿父目录向上找 `node_modules/`,所以 workspace 根目录已装的 playwright 也能复用。 4. **Playwright 不够/缺失**时:在 `scripts/` 目录下自己跑 `npm install`。**重要**:执行时设环境变量 `PLAYWRIGHT_SKIP_BROWSER_DOWNLOAD=1`(这个 skill 用的是系统 Edge,不需要再下载 playwright 自带的 chromium,能省几百 MB)。 ``` The manifest uses a broad compatible-version range: ```json "dependencies": { "playwright": "^1.45.0" } ``` The lockfile resolves Playwright through a non-official mirror: ```json "node_modules/playwright": { "version": "1.60.0", "resolved": "https://registry.npmmirror.com/playwright/-/playwright-1.60.0.tgz", "integrity": "sha512-hheHdokM8cdqCb0lcE3s+zT4t4W+vvjpGxsZlDnikarzx8tSzMebh3UiFtgqwFwnTnjYQcsyMF8ei2mCO/tpeA==", "license": "Apache-2.0", "dependencies": { "playwright-core": "1.60.0" }, "bin": { "playwright": "cli.js" }, "engines": { "node": ">=18" }, "optionalDependencies": { "fsevents": "2.3.2" } }, "node_modules/playwright-core": { "version": "1.60.0", "resolved": "https://registry.npmmirror.com/playwright-core/-/playwright-core-1.60.0.tgz", "integrity": "sha512-9bW6zvX/m0lEbgTKJ6YppOKx8H3VOPBMOCFh2irXFOT4BbHgrx5hPjwJYLT40Lu+4qtD36qKc/Hn56StUW57IA==", "license": "Apache-2.0", "bin": { "playwright-core": "cli.js" }, "engines": { "node": ">=18" } } ``` ### Technical Analysis The Skill automatically runs `npm install`, while its lockfile retrieves core executable dependencies from `registry.npmmirror.com` rather than the official npm registry. The included SHA-512 i ...[truncated 1790 chars]
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (18)

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The invocation description says users can simply say "知网下载" to start the skill, but it does not define scope, exclusions, or alternative phrasing boundaries. As written, the trigger guidance is minimal and could match broad conversational mentions of CNKI downloading rather than a clearly constrained command.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
83% confidence
Finding

The skill instructs the agent to inspect environment-dependent paths, check installed software, run shell commands, and manage local files, but it does not declare any explicit tool scope or allowed-tools boundary. In practice this increases the chance that an agent invokes broader file/system capabilities than intended, especially because the document contains operational instructions for process launch, package installation, and persistent state handling.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The trigger phrases include very generic requests like '帮我下几篇知网文献' and similar natural language variants, which can match ordinary conversation without confirming the user wants this specific automation. Because the skill performs high-impact actions such as opening browsers, preserving login state, and downloading files, broad activation increases the risk of unintended execution and accidental handling of authenticated sessions.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The skill description mentions that login state is retained in browser_data/ and cookies remain valid for hours, but it does not present this as a clear user-facing privacy/security warning with consent. Storing reusable authenticated session material locally creates risk of session theft, unintended reuse by other local users/processes, or accidental inclusion in backups or source control.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The instruction to 'directly use this skill' for broad CNKI-related utterances removes normal ambiguity handling and encourages immediate execution. In this context, that is risky because the workflow includes process spawning, possible npm install, browser automation, local persistence of cookies, and file writes, so accidental activation has more consequence than a read-only helper skill.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The config hard-codes Chinese keyword, sort field, and filter column/value labels such as "区块链", "发表时间", and "来源类别". This imposes a specific language/locale behavior in a way that is not presented as optional or justified as region-specific, which matches the language/locale policy-violation category.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The CLI parses a --headless flag and prints mode-dependent messaging, but launchPersistentContext is hard-coded with headless: false. This actively contradicts the documented CLI behavior and can mislead users about whether the script will open a visible browser window.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
85% confidence
Finding

The user-facing instructions are presented only in Chinese for the required login and search steps, while the script otherwise uses mixed-language messaging. This imposes a specific language for critical workflow steps without offering an opt-in or alternative locale, which matches the language/locale policy violation criteria.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The browser context is hard-coded with locale: 'zh-CN', which enforces a specific language/locale for all users. This is a natural-language policy concern because the file does not offer any user choice or document a justified region-specific requirement.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The package description is written entirely in Chinese, which signals a fixed language expectation for the skill without any indication that users may choose another language or that the skill is intentionally limited to a Chinese-speaking context. Under the policy, locale or language constraints should be opt-in or clearly justified.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

This code performs an authenticated HTTP GET using the browser context's shared cookie state and an explicit Referer header, which transmits session-linked request data to the remote site. Although comments describe the mechanism, there is no confirmation prompt or user-facing disclosure in the code about this network action.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

The module comments and public API descriptions require Chinese labels such as 筛选列名 and 按钮上的中文名, and the exported interfaces expect Chinese values like "来源类别" and "相关度". This is a natural-language locale constraint embedded in the skill without any opt-in, fallback, or justification that the skill is intentionally region-specific.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The script prints the full CNKI URL from configuration to stdout, and that URL may contain institution-specific entry points, embedded search parameters, or access tokens. In this skill's context, the URL is user-supplied and tied to authenticated academic access workflows, so logging it increases the chance of leaking sensitive browsing or access information into terminal history, agent logs, or telemetry.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
87% confidence
Finding

The README states that a user-provided CNKI entry URL is saved to scripts/user_config.json, but it does not clearly warn users that this value will persist on disk. Persistent storage of user-supplied URLs can expose institution-specific access endpoints or browsing preferences to other local users or later processes, especially when paired with retained browser session data.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
89% confidence
Finding

This markdown file states that downloads are saved to ./scripts/download/ and that the directory is automatically created, which is a filesystem-modifying behavior. Under the markdown-file criteria for SQP-2, the description should warn users about behaviors that affect local data or system state, but this is presented as implementation detail rather than a clear user warning.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
84% confidence
Finding

The dependency uses a caret range (^1.45.0), which allows automatic installation of newer minor/patch versions. In a skill that automates browser login flows, preserves session state, and downloads content, an unexpected upstream dependency change could introduce supply-chain risk or break security assumptions without review.

Content

Scanner excerpt · scripts/package.json (reported line 12)May include surrounding context.

json
"demo": "node search_cnki.js"
  },
  "dependencies": {
    "playwright": "^1.45.0"
  },
  "engines": {
    "node": ">=18.0.0"

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
83% confidence
Finding

The code creates a directory and later saves the downloaded file locally, which affects user storage. While comments explain the implementation, there is no visible user disclosure such as a prompt, log message announcing the save location, or documented warning in this file.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
94% confidence
Finding

Natural-language strings throughout the header and CLI help are Chinese-only, including usage instructions and error/help text. This imposes a fixed language on users without opt-in or an alternative locale, which matches the policy's language/locale violation criterion.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.