Back to skill

Security audit

知乎抓取.SKILL

Security checks for vulnerabilities and agentic risk

Overview

This Zhihu scraping skill is mostly purpose-aligned, but it needs review because it stores session cookies and can expose them or use them too broadly while writing local files.

Install only after reviewing the cookie-handling risks. Use a dedicated workspace, protect or delete zhihu_cookies.json when finished, avoid untrusted batch JSON or untrusted article content until image-cookie leakage is fixed, and back up any Obsidian vault before running export commands.

Vulnerability Patterns
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
Findings (5)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/zhihu_fetch/fetch/batch.py:126
Finding

Zhihu Session Cookies Are Disclosed to Arbitrary Image Hosts

Content
View full analysis
Remediation
View remediation

T05 · Unauthorized Access and Privilege Escalation

Error
Location
scripts/zhihu_fetch/auth/login.py:23
Finding

Unrestricted Authenticated-Browser Navigation Enables SSRF-Like Requests

Content
View full analysis
1: raw = sys.argv[1].strip() if not raw: return None if not (raw.startswith("http://") or raw.startswith("https://")): print( f"\n[!] Optional verification requires a complete URL; ignored: {raw!r}" ) return None return raw ``` That value is opened in the persistent browser: ```python if logged_in: verify_url = optional_verify_url() if verify_url: print(f"\nOptional verification: opening {verify_url}") await page.goto( verify_url, wait_until="domcontentloaded", timeout=30000 ) await asyncio.sleep(3) content = await page.evaluate("() => document.body.innerText") if "Please log in to view" not in content: print("The current session is considered able to access the page.") ``` Batch input URLs are also accepted without destination validation: ```python with open(list_file, 'r', encoding='utf-8') as f: collection = json.load(f) items = collection.get('items', []) empty_urls = [item for item in items if not (item.get('url') or '').strip()] if empty_urls: items = [item for item in items if (item.get('url') or '').strip()] ``` They are later opened in the persistent browser context: ```python for i, item in enumerate(items): url = item.get('url', '') title = item.get('title', f'Article{i+1}') if n ...[truncated 2414 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
scripts/zhihu_fetch/core/url.py:33
Finding

Substring-Based Host Validation Accepts Attacker-Controlled Lookalike Domains

Content
View full analysis
= 2 and parts[0] == "p" and parts[1].isdigit(): return ZhihuTarget( kind="article", url=url, article_id=parts[1] ) if parts and parts[0] not in {"p", "api"}: return ZhihuTarget( kind="column", url=url, column_id=parts[0] ) return ZhihuTarget(kind="unknown", url=url) if "zhihu.com" not in host: return ZhihuTarget(kind="unknown", url=url) ``` The resulting target is forwarded into fetch pipelines: ```python if kind == "article": _run("zhihu_fetch.fetch.single", [target.url or raw, *rest]) if kind == "answer": _run("zhihu_fetch.fetch.single", [target.url or raw, *rest]) if kind == "collection": _run("zhihu_fetch.fetch.collection", [target.url or raw, *rest]) ... if kind == "column": _run("zhihu_fetch.fetch.columns", [target.url or raw, *rest]) ``` ### Technical Analysis The classifier uses substring containment instead of validating hostname boundaries. The following attacker-controlled hosts therefore satisfy the checks: - `zhihu.com.attacker.example` - `notzhihu.com` - `zhuanlan.zhihu.com.attacker.example` In addition, the code derives `host` from `parsed.netloc`, which can contain user information and a port, rather than using the normalized `parsed.hostname` property. An attacker can combine a lookalike hostname with a recognized path to have the URL classified and routed as a ...[truncated 1042 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/zhihu_fetch/fetch/batch.py:49
Finding

Authentication Cookies Are Persisted Without Restrictive File Protections

Content
View full analysis
Remediation
View remediation

T08 · Insecure Dependencies

Warning
Location
scripts/requirements.txt:1
Finding

Runtime Dependencies Are Unpinned and Lack Integrity Verification

Content
View full analysis
=2.28.0 beautifulsoup4>=4.12.0 playwright>=1.40.0 pytest>=8.0.0 pytest>=8.0.0 ``` The Skill documentation instructs users to install these dependencies: ```text pip install -r requirements.txt playwright install chromium ``` ### Technical Analysis All packages use open-ended minimum versions. A future installation can therefore resolve to dependency releases that were not reviewed with the Skill. No lockfile, package hashes, or index restrictions are supplied. `pytest` is duplicated and appears in the runtime dependency list despite being a development/testing tool. This unnecessarily increases the installed package set and supply-chain attack surface. No evidence was found that the listed package names are intentionally malicious or typosquatted. The risk arises from unconstrained future resolution and missing integrity guarantees. ### Attack Path 1. A user follows the documented installation procedure. 2. `pip` resolves the newest package versions that satisfy the lower bounds. 3. A future compromised, malicious, or incompatible release is selected without requiring a Skill update. 4. Package installation hooks or imported runtime code execute with the user’s privileges. 5. The dependency may access the same filesystem, network, browser profile, and saved cookie data available to the Skill. ### Impact Assessment A compromised dependency could execute code with the privileges of the user running the installation or Skill. This could expose saved authentication cookies, modify output files, access the Obsidian vault, or perform arbitrary network activity. Even without malicious compromise, unconstrained versions can cause non-reproducible behavior and security regressions. ]]>
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (57)

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The code is related to the general declared domain of Zhihu article scraping, but it implements a much narrower capability set than described. It only supports interactive fetching of a single article URL via Playwright with persistent browser context, manual captcha handling, basic text extraction, and saving to a TXT file. The declared description promises broader features such as favorites/collection scraping, batch retrieval, fallback between API and Playwright, image handling, resumable progress, and optional Obsidian export, none of which appear in this code chunk. Therefore the description does not accurately represent what this specific code chunk actually does.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
94% confidence
Finding

The declared description presents a broader workflow-oriented scraper centered on Zhihu collections/favorites, batch retrieval, persistence, and export features. This code chunk instead implements only one narrow component: fetching a single article via Playwright in stealth mode and writing extracted text to a local TXT file. While article content scraping is related to the declared domain, the implemented behavior omits most of the advertised capabilities and adds a notable anti-detection/stealth automation behavior that is not described. Therefore the description does not accurately represent what this code chunk actually does.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The description emphasizes fetching/scraping Zhihu collections and articles, including network-oriented capabilities like API/Playwright fallback, cookie management, image fetching, and resumable collection. This code chunk instead operates purely on local files already present in an Obsidian vault. It parses mirrored article metadata, extracts summaries/quotes, writes new Markdown notes, and updates local author/question index files. While 'optional Obsidian export' is mentioned in the declaration, this chunk’s primary behavior is specifically note generation from existing mirrored content, not scraping. Therefore the supplied code chunk materially differs from the declared purpose and lacks most of the declared core capabilities.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
93% confidence
Finding

The declared description presents a broader and somewhat different capability set centered on Zhihu 收藏夹 (favorites/collections) and full article content scraping with multi-level fallback, cookie persistence/keepalive, image capture, resume, and optional Obsidian writing. The supplied code chunk instead targets Zhihu 专栏/columns: it enumerates a user's columns, filters them, fetches per-column article lists from Zhihu JSON APIs, saves structured JSON, and records URLs for incremental fetching. There is partial overlap with Zhihu scraping and batch article retrieval, and there is some support for incremental/resume-like behavior via seen URLs and since-last filtering. However, the primary purpose and major advertised capabilities do not match this chunk: it does not scrape favorites/collections, does not fetch full article bodies or images, does not use Playwright fallback, and does not export to Obsidian. Therefore this chunk materially differs from the declared description.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The declared description centers on scraping Zhihu collections/favorites and article contents, with fallback mechanisms, cookie handling, resume support, and optional Obsidian export. This code chunk does not implement those features. Instead, it acts as a command-line entry point for a personal-profile 'follow' mode: it extracts a people slug, forces incremental fetching unless disabled, and dispatches to column and post fetch modules for that user. While this is still broadly Zhihu scraping, the primary purpose of this specific code is materially different from the declared collection/article-content workflow, so it should be flagged as a mismatch.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The declared description emphasizes 收藏夹与文章内容抓取 (collection scraping and article content extraction), including multi-level API/Playwright fallback, cookie persistence, batch body/image retrieval, resume, and optional Obsidian writing. The actual code does include some overlapping support features such as Playwright fallback behavior, cookie loading, and checkpoint/resume. However, its substantive behavior is different: it visits a Zhihu profile, calls/captures the /api/v3/moments/.../activities feed, filters activities whose action text begins with 赞同了/喜欢了/收藏了, and outputs item URLs plus metadata like title, author, vote count, interaction time, and action type. It does not download the actual article/answer正文, does not fetch images, does not scrape collection/favorite folders themselves, and does not write to Obsidian. Therefore the primary purpose of this chunk is activity-history extraction, not collection/content scraping as declared.

Content

No source excerpt is available for this finding.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 19)May include surrounding context.

md
- **技能根目录**:本仓库根目录(含 `SKILL.md` 与 `scripts/`)。下文命令均从该目录执行,写作 `python scripts/...`。

Lp1

High
Category
MCP Least Privilege
Confidence
75% confidence
Finding

The skill uses 'env' capability that is not listed in its permissions. This may indicate deceptive intent or missing permission declarations.

Content

No source excerpt is available for this finding.

Env Variable Harvesting

High
Category
Data Exfiltration
Confidence
60% confidence
Finding

Code enumerates, copies, or searches environment variables for secrets. Bulk environment access can collect credentials unrelated to the skill's stated purpose.

Content

Scanner excerpt · scripts/zhihu_fetch/fetch/batch.py (reported line 939)May include surrounding context.

python
'--retry-failed',
                '--no-auto-retry',
            ]
            result = subprocess.run(cmd, cwd=os.getcwd(), env=os.environ.copy())
            if result.returncode != 0:
                print(f"[!] 自动重试进程退出码: {result.returncode}")
                break

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The description advertises 'Cookie 持久化与保活', and later documents persistent browser context usage, but does not warn users that authenticated session state will be stored locally. Because cookies and browser session data are sensitive authentication material, the markdown should clearly disclose this behavior and its privacy/security implications.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
84% confidence
Finding

This markdown file instructs the user to run commands that write article content and generated notes into a Vault path, which affects local user data. Although the commands are part of the skill's purpose, the README does not include a clear user-facing warning that these operations will create or update files in the specified Obsidian vault.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The trigger conditions are broad enough to match ordinary conversation about Zhihu, cookies, CAPTCHA, or Obsidian, which can cause unintended activation of a powerful skill with Bash, WebFetch, Read, Write, and Edit access. In this context, accidental invocation is more dangerous because the skill is designed to scrape authenticated content, persist cookies, and write to local files or an Obsidian vault, increasing the chance of privacy-impacting or unexpected side effects.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The module docstring and all user-facing help output are written only in Chinese, which imposes a specific language on users without any opt-in or alternative locale. The policy explicitly flags language or locale constraints when the skill does not offer a language choice or document a justified regional limitation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

This code hard-codes locale='zh-CN' and timezone_id='Asia/Shanghai', which enforces a specific language/locale behavior for all users. The policy allows locale constraints only when the user is offered a choice or when the restriction is clearly documented and justified as region-specific.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The script injects a stealth payload that hides Playwright automation signals by modifying navigator.webdriver and adding fake Chrome properties. In a login helper whose stated purpose is simply to let a user sign in and save cookies, anti-detection behavior increases the capability to evade platform bot defenses and can facilitate unauthorized or policy-violating scraping, making the skill context more concerning rather than less.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
86% confidence
Finding

Launching Chromium with --no-sandbox disables an important browser security boundary and is not required for ordinary content fetching on a user workstation. If malicious content is loaded or a browser exploit is encountered during login or scraping, the reduced isolation can increase the severity of compromise.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The code writes authentication cookies, including the login-validating z_c0 token, to disk in JSON form without access controls, encryption, or an explicit warning to the user. In this skill context, persisted Zhihu session cookies can be reused by anyone with filesystem access to impersonate the account and access private authenticated content.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The script extracts all browser cookies after login and writes them in plaintext JSON to a local file, including the sensitive Zhihu authentication token z_c0. Anyone with filesystem access, backup access, or malware on the host could reuse these cookies to hijack the logged-in session without needing the user's password or MFA. In this skill's context, cookie persistence is functional for scraping, but it materially increases risk because the whole purpose is to maintain authenticated access over time.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The code automatically loads persisted Zhihu cookies and attaches them to outbound requests, which can transmit authenticated session material without explicit user confirmation or visibility at the point of use. In a scraping/export skill, this increases the risk of unintended account-context disclosure, misuse of privileged access, or sending sensitive cookies to requests the user did not realize would be authenticated.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The browser context is hard-coded to use locale='zh-CN' and timezone_id='Asia/Shanghai'. This is a natural-language/locale policy issue because the skill imposes a specific locale on all users without offering opt-in or documenting that it is restricted to a China-specific use case.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

This code forces locale='zh-CN', timezone_id='Asia/Shanghai', and Accept-Language values favoring Chinese, which is a natural-language/locale policy choice applied silently to all users. The file does not provide opt-in, fallback, or documentation that the skill is intended only for a China-specific compliance or regional use case.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

This code prints all user-visible summary labels in Chinese, such as the heading and status fields, with no mechanism for selecting or opting into another language. That is a natural-language locale policy concern because the skill forces a specific language in its output rather than offering a choice or documenting a justified locale restriction.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The usage message printed to users is hard-coded in Chinese, and the generated note path/title also use Chinese-only labels elsewhere in the file. For a code file, this is a natural-language locale policy concern because the skill does not offer user opt-in or indicate that it is intentionally limited to Chinese-language environments.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The function is presented as a non-destructive 'sync', but it uses shutil.move() to transfer images into the vault, deleting the originals from the source directory. This can cause silent data loss, break resumable workflows, and violate user expectations, especially in a scraping/export tool where source artifacts may be needed for retry or audit.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The script performs destructive operations on user data by moving image files and deleting article files without explicit confirmation or a clearly advertised destructive mode. In the context of a content-export skill with claims like 'sync' and 'write', this mismatch makes accidental data loss more likely and increases operational risk.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.