T09 · Insecure Skill Coding Practices
- Location
scripts/zhihu_fetch/fetch/batch.py:126- Finding
Zhihu Session Cookies Are Disclosed to Arbitrary Image Hosts
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This Zhihu scraping skill is mostly purpose-aligned, but it needs review because it stores session cookies and can expose them or use them too broadly while writing local files.
Install only after reviewing the cookie-handling risks. Use a dedicated workspace, protect or delete zhihu_cookies.json when finished, avoid untrusted batch JSON or untrusted article content until image-cookie leakage is fixed, and back up any Obsidian vault before running export commands.
scripts/zhihu_fetch/fetch/batch.py:126Zhihu Session Cookies Are Disclosed to Arbitrary Image Hosts
scripts/zhihu_fetch/auth/login.py:23Unrestricted Authenticated-Browser Navigation Enables SSRF-Like Requests
scripts/zhihu_fetch/core/url.py:33Substring-Based Host Validation Accepts Attacker-Controlled Lookalike Domains
scripts/zhihu_fetch/fetch/batch.py:49Authentication Cookies Are Persisted Without Restrictive File Protections
scripts/requirements.txt:1Runtime Dependencies Are Unpinned and Lack Integrity Verification
The code is related to the general declared domain of Zhihu article scraping, but it implements a much narrower capability set than described. It only supports interactive fetching of a single article URL via Playwright with persistent browser context, manual captcha handling, basic text extraction, and saving to a TXT file. The declared description promises broader features such as favorites/collection scraping, batch retrieval, fallback between API and Playwright, image handling, resumable progress, and optional Obsidian export, none of which appear in this code chunk. Therefore the description does not accurately represent what this specific code chunk actually does.
The declared description presents a broader workflow-oriented scraper centered on Zhihu collections/favorites, batch retrieval, persistence, and export features. This code chunk instead implements only one narrow component: fetching a single article via Playwright in stealth mode and writing extracted text to a local TXT file. While article content scraping is related to the declared domain, the implemented behavior omits most of the advertised capabilities and adds a notable anti-detection/stealth automation behavior that is not described. Therefore the description does not accurately represent what this code chunk actually does.
The description emphasizes fetching/scraping Zhihu collections and articles, including network-oriented capabilities like API/Playwright fallback, cookie management, image fetching, and resumable collection. This code chunk instead operates purely on local files already present in an Obsidian vault. It parses mirrored article metadata, extracts summaries/quotes, writes new Markdown notes, and updates local author/question index files. While 'optional Obsidian export' is mentioned in the declaration, this chunk’s primary behavior is specifically note generation from existing mirrored content, not scraping. Therefore the supplied code chunk materially differs from the declared purpose and lacks most of the declared core capabilities.
The declared description presents a broader and somewhat different capability set centered on Zhihu 收藏夹 (favorites/collections) and full article content scraping with multi-level fallback, cookie persistence/keepalive, image capture, resume, and optional Obsidian writing. The supplied code chunk instead targets Zhihu 专栏/columns: it enumerates a user's columns, filters them, fetches per-column article lists from Zhihu JSON APIs, saves structured JSON, and records URLs for incremental fetching. There is partial overlap with Zhihu scraping and batch article retrieval, and there is some support for incremental/resume-like behavior via seen URLs and since-last filtering. However, the primary purpose and major advertised capabilities do not match this chunk: it does not scrape favorites/collections, does not fetch full article bodies or images, does not use Playwright fallback, and does not export to Obsidian. Therefore this chunk materially differs from the declared description.
The declared description centers on scraping Zhihu collections/favorites and article contents, with fallback mechanisms, cookie handling, resume support, and optional Obsidian export. This code chunk does not implement those features. Instead, it acts as a command-line entry point for a personal-profile 'follow' mode: it extracts a people slug, forces incremental fetching unless disabled, and dispatches to column and post fetch modules for that user. While this is still broadly Zhihu scraping, the primary purpose of this specific code is materially different from the declared collection/article-content workflow, so it should be flagged as a mismatch.
The declared description emphasizes 收藏夹与文章内容抓取 (collection scraping and article content extraction), including multi-level API/Playwright fallback, cookie persistence, batch body/image retrieval, resume, and optional Obsidian writing. The actual code does include some overlapping support features such as Playwright fallback behavior, cookie loading, and checkpoint/resume. However, its substantive behavior is different: it visits a Zhihu profile, calls/captures the /api/v3/moments/.../activities feed, filters activities whose action text begins with 赞同了/喜欢了/收藏了, and outputs item URLs plus metadata like title, author, vote count, interaction time, and action type. It does not download the actual article/answer正文, does not fetch images, does not scrape collection/favorite folders themselves, and does not write to Obsidian. Therefore the primary purpose of this chunk is activity-history extraction, not collection/content scraping as declared.
Referenced artifact was not completely inspected
- **技能根目录**:本仓库根目录(含 `SKILL.md` 与 `scripts/`)。下文命令均从该目录执行,写作 `python scripts/...`。
The skill uses 'env' capability that is not listed in its permissions. This may indicate deceptive intent or missing permission declarations.
Code enumerates, copies, or searches environment variables for secrets. Bulk environment access can collect credentials unrelated to the skill's stated purpose.
'--retry-failed',
'--no-auto-retry',
]
result = subprocess.run(cmd, cwd=os.getcwd(), env=os.environ.copy())
if result.returncode != 0:
print(f"[!] 自动重试进程退出码: {result.returncode}")
break
The description advertises 'Cookie 持久化与保活', and later documents persistent browser context usage, but does not warn users that authenticated session state will be stored locally. Because cookies and browser session data are sensitive authentication material, the markdown should clearly disclose this behavior and its privacy/security implications.
This markdown file instructs the user to run commands that write article content and generated notes into a Vault path, which affects local user data. Although the commands are part of the skill's purpose, the README does not include a clear user-facing warning that these operations will create or update files in the specified Obsidian vault.
The trigger conditions are broad enough to match ordinary conversation about Zhihu, cookies, CAPTCHA, or Obsidian, which can cause unintended activation of a powerful skill with Bash, WebFetch, Read, Write, and Edit access. In this context, accidental invocation is more dangerous because the skill is designed to scrape authenticated content, persist cookies, and write to local files or an Obsidian vault, increasing the chance of privacy-impacting or unexpected side effects.
The module docstring and all user-facing help output are written only in Chinese, which imposes a specific language on users without any opt-in or alternative locale. The policy explicitly flags language or locale constraints when the skill does not offer a language choice or document a justified regional limitation.
This code hard-codes locale='zh-CN' and timezone_id='Asia/Shanghai', which enforces a specific language/locale behavior for all users. The policy allows locale constraints only when the user is offered a choice or when the restriction is clearly documented and justified as region-specific.
The script injects a stealth payload that hides Playwright automation signals by modifying navigator.webdriver and adding fake Chrome properties. In a login helper whose stated purpose is simply to let a user sign in and save cookies, anti-detection behavior increases the capability to evade platform bot defenses and can facilitate unauthorized or policy-violating scraping, making the skill context more concerning rather than less.
Launching Chromium with --no-sandbox disables an important browser security boundary and is not required for ordinary content fetching on a user workstation. If malicious content is loaded or a browser exploit is encountered during login or scraping, the reduced isolation can increase the severity of compromise.
The code writes authentication cookies, including the login-validating z_c0 token, to disk in JSON form without access controls, encryption, or an explicit warning to the user. In this skill context, persisted Zhihu session cookies can be reused by anyone with filesystem access to impersonate the account and access private authenticated content.
The script extracts all browser cookies after login and writes them in plaintext JSON to a local file, including the sensitive Zhihu authentication token z_c0. Anyone with filesystem access, backup access, or malware on the host could reuse these cookies to hijack the logged-in session without needing the user's password or MFA. In this skill's context, cookie persistence is functional for scraping, but it materially increases risk because the whole purpose is to maintain authenticated access over time.
The code automatically loads persisted Zhihu cookies and attaches them to outbound requests, which can transmit authenticated session material without explicit user confirmation or visibility at the point of use. In a scraping/export skill, this increases the risk of unintended account-context disclosure, misuse of privileged access, or sending sensitive cookies to requests the user did not realize would be authenticated.
The browser context is hard-coded to use locale='zh-CN' and timezone_id='Asia/Shanghai'. This is a natural-language/locale policy issue because the skill imposes a specific locale on all users without offering opt-in or documenting that it is restricted to a China-specific use case.
This code forces locale='zh-CN', timezone_id='Asia/Shanghai', and Accept-Language values favoring Chinese, which is a natural-language/locale policy choice applied silently to all users. The file does not provide opt-in, fallback, or documentation that the skill is intended only for a China-specific compliance or regional use case.
This code prints all user-visible summary labels in Chinese, such as the heading and status fields, with no mechanism for selecting or opting into another language. That is a natural-language locale policy concern because the skill forces a specific language in its output rather than offering a choice or documenting a justified locale restriction.
The usage message printed to users is hard-coded in Chinese, and the generated note path/title also use Chinese-only labels elsewhere in the file. For a code file, this is a natural-language locale policy concern because the skill does not offer user opt-in or indicate that it is intentionally limited to Chinese-language environments.
The function is presented as a non-destructive 'sync', but it uses shutil.move() to transfer images into the vault, deleting the originals from the source directory. This can cause silent data loss, break resumable workflows, and violate user expectations, especially in a scraping/export tool where source artifacts may be needed for retry or audit.
The script performs destructive operations on user data by moving image files and deleting article files without explicit confirmation or a clearly advertised destructive mode. In the context of a content-export skill with claims like 'sync' and 'write', this mismatch makes accidental data loss more likely and increases operational risk.
No suspicious patterns detected.