Back to skill

Security audit

知乎数据获取 | Zhihu Data Fetcher

Security checks for vulnerabilities and agentic risk

Overview

This skill mostly fetches and stores Zhihu hot-list data, but it ships sensitive cookie material and includes under-scoped diagnostic scripts that users should review before installing.

Install only if you are comfortable reviewing and sanitizing the package first. Remove the bundled cookie values, do not commit or share your own Zhihu cookies, avoid running the anti-crawl/browser-research snippets unless you understand the console/log exposure, and treat generated HTML reports as unsafe if they include untrusted fetched content.

Vulnerability Patterns
  • Tool Hijacking and SpoofingModifies or replaces tools so legitimate-looking calls execute attacker logic
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (4)

T09 · Insecure Skill Coding Practices

Error
Location
config/fallback-sources.json:6
Finding

Committed Zhihu Session Credentials in Plaintext Configuration

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
snippets/cookie-manager.js:103
Finding

Authentication Cookie Material Disclosed in Process Logs

Content
View full analysis
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
scripts/generate_html.py:414
Finding

Stored HTML and JavaScript Injection in Generated Reports

Content
View full analysis
`
${a.rank}
${a.title}
${formatHeat(a.heat)} ${a.date} 采集时间: ${a.fetched_at}
`).join(''); container.innerHTML = html; ``` The report also embeds externally sourced records directly into a script block: ```python const articles = {json.dumps(articles, ensure_ascii=False)}; ``` ### Technical Analysis Article titles and URLs originate from Zhihu or a remotely fetched GitHub Markdown fallback and are stored without output sanitization. The report then inserts those fields into an HTML string and assigns it to `innerHTML`. A malicious title can inject HTML event handlers or other active markup. A crafted `` sequence can also terminate the generated script block because JSON serialization does not inherently make untrusted strings safe for direct embedding in HTML script content. A malicious URL can inject attributes or use a dangerous scheme. The configured GitHub fallback is not executed as code when downloaded, so this is not remote payload retrieval and execution by itself. However, compromise of that content can provide attacker-controlled records that later become executable when the generated report is opened. The generated links also use `target="_blank"` without `rel="noopener noreferrer"`, allowing the opened destination to manipulate the opener in browser environments that do not apply implicit opener isolation. ### At ...[truncated 1055 chars]
Remediation
View remediation
`, `&`, and script-closing sequences before insertion. 6. Consider using a templating engine with context-aware auto-escaping. 7. Add a restrictive Content Security Policy that disallows inline scripts and limits navigation and network destinations. 8. Add `rel="noopener noreferrer"` to links opened with `target="_blank"`. 9. Sanitize existing database content before regenerating reports. ]]>

T07 · Tool Hijacking and Spoofing

Warning
Location
snippets/browser-research.js:12
Finding

Browser Fetch API Hijacking and Sensitive Session Reconnaissance

Content
View full analysis
c.trim()).filter(c => c.startsWith('_xsrf') || c.startsWith('_zap') || c.startsWith('d_c0') || c.startsWith('SESSIONID') ), userAgent: navigator.userAgent, screen: { width: screen.width, height: screen.height, colorDepth: screen.colorDepth }, timezone: Intl.DateTimeFormat().resolvedOptions().timeZone, language: navigator.language, platform: navigator.platform, windowKeys: Object.keys(window).filter(k => k.toLowerCase().includes('zhihu') || k.toLowerCase().includes('anti') || k.toLowerCase().includes('captcha') ).slice(0, 20) }; console.log('📊 当前页面特征:', result); return result; } function monitorNetwork() { const originalFetch = window.fetch; window.fetch = async (...args) => { const [url, options] = args; console.log('📡 Fetch 请求:', { url: url.toString().slice(0, 100), headers: options?.headers, timestamp: new Date().toISOString() }); try { const response = await originalFetch(...args); console.log('📥 响应:', { status: response.status, headers: Object.fromEntries(response.headers.entries()) }); return response; } catch (error) { console.error('❌ 请求失败:', error); throw error; } }; console.log('✅ 网络监听已启动'); } ``` ### Technical Analysis The browser research utility reads JavaScript-accessible Zhihu cookie values and browser fingerprint attributes, then prints them to the developer console. It also replaces the page-wide `window.fetch` function and logs headers associated with subsequent reque ...[truncated 1584 chars]
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (59)

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

A script that only loads file-based cookies while claiming broader auth fallback conceals the real trust boundary: sensitive credentials stored on disk are the primary auth mechanism. This matters because users may not realize the skill depends on local secret material and lacks a safe degraded mode if those credentials are absent or mishandled.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

A script that only loads file-based cookies while claiming broader auth fallback conceals the real trust boundary: sensitive credentials stored on disk are the primary auth mechanism. This matters because users may not realize the skill depends on local secret material and lacks a safe degraded mode if those credentials are absent or mishandled.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
94% confidence
Finding

A script that only loads file-based cookies while claiming broader auth fallback conceals the real trust boundary: sensitive credentials stored on disk are the primary auth mechanism. This matters because users may not realize the skill depends on local secret material and lacks a safe degraded mode if those credentials are absent or mishandled.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

A script that only loads file-based cookies while claiming broader auth fallback conceals the real trust boundary: sensitive credentials stored on disk are the primary auth mechanism. This matters because users may not realize the skill depends on local secret material and lacks a safe degraded mode if those credentials are absent or mishandled.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

A script that only loads file-based cookies while claiming broader auth fallback conceals the real trust boundary: sensitive credentials stored on disk are the primary auth mechanism. This matters because users may not realize the skill depends on local secret material and lacks a safe degraded mode if those credentials are absent or mishandled.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

A script that only loads file-based cookies while claiming broader auth fallback conceals the real trust boundary: sensitive credentials stored on disk are the primary auth mechanism. This matters because users may not realize the skill depends on local secret material and lacks a safe degraded mode if those credentials are absent or mishandled.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

A script that only loads file-based cookies while claiming broader auth fallback conceals the real trust boundary: sensitive credentials stored on disk are the primary auth mechanism. This matters because users may not realize the skill depends on local secret material and lacks a safe degraded mode if those credentials are absent or mishandled.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

A script that only loads file-based cookies while claiming broader auth fallback conceals the real trust boundary: sensitive credentials stored on disk are the primary auth mechanism. This matters because users may not realize the skill depends on local secret material and lacks a safe degraded mode if those credentials are absent or mishandled.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

A script that only loads file-based cookies while claiming broader auth fallback conceals the real trust boundary: sensitive credentials stored on disk are the primary auth mechanism. This matters because users may not realize the skill depends on local secret material and lacks a safe degraded mode if those credentials are absent or mishandled.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

A script that only loads file-based cookies while claiming broader auth fallback conceals the real trust boundary: sensitive credentials stored on disk are the primary auth mechanism. This matters because users may not realize the skill depends on local secret material and lacks a safe degraded mode if those credentials are absent or mishandled.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

A script that only loads file-based cookies while claiming broader auth fallback conceals the real trust boundary: sensitive credentials stored on disk are the primary auth mechanism. This matters because users may not realize the skill depends on local secret material and lacks a safe degraded mode if those credentials are absent or mishandled.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
94% confidence
Finding

A script that only loads file-based cookies while claiming broader auth fallback conceals the real trust boundary: sensitive credentials stored on disk are the primary auth mechanism. This matters because users may not realize the skill depends on local secret material and lacks a safe degraded mode if those credentials are absent or mishandled.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

A script that only loads file-based cookies while claiming broader auth fallback conceals the real trust boundary: sensitive credentials stored on disk are the primary auth mechanism. This matters because users may not realize the skill depends on local secret material and lacks a safe degraded mode if those credentials are absent or mishandled.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
94% confidence
Finding

A script that only loads file-based cookies while claiming broader auth fallback conceals the real trust boundary: sensitive credentials stored on disk are the primary auth mechanism. This matters because users may not realize the skill depends on local secret material and lacks a safe degraded mode if those credentials are absent or mishandled.

Content

No source excerpt is available for this finding.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 336)May include surrounding context.

md
node snippets/test-simple.js

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The file contains a hardcoded Zhihu authenticated cookie set, including session and anti-CSRF-related values, which are effectively secrets embedded in source-controlled configuration. If valid, these credentials can be reused by anyone with access to the skill to impersonate the account, scrape private or rate-limited data, or trigger account abuse; even if expired, embedding real-looking auth material normalizes unsafe secret handling and may indicate credential leakage from a real browser session.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The script is explicitly framed as an anti-crawl research tool and includes logic for testing request-header combinations and analyzing platform defenses, which materially exceeds the stated purpose of a minimalist Zhihu data fetcher. In the context of an agent skill, this creates dual-use capability for reconnaissance against a third-party service’s anti-abuse controls and increases the chance the skill will be used to bypass access restrictions.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
95% confidence
Finding

The predefined header sets emulate browser-originated traffic and are designed to compare which combinations succeed against the target API, including fetch-related headers commonly used to mimic same-origin browser requests. That behavior is not necessary for ordinary data retrieval and can be used to tune requests to evade anti-bot filtering, making the skill more dangerous in context.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The rate-limit test intentionally sends repeated requests to detect throttling thresholds, which is classic reconnaissance for service-abuse tuning. In a fetcher skill, this is especially concerning because it operationalizes how to measure and adapt to platform defenses rather than simply retrieving user-requested content.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
86% confidence
Finding

The skill documents use of browser state, local files, shell commands, network retrieval, and database/file generation, but it declares no explicit tool scope or permissions. In an agent setting, this creates an over-privileged and under-specified capability boundary, increasing the risk of unintended file access, credential handling, or network use without clear user consent.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The manifest frames the skill as minimalist data retrieval, but the documentation shows persistent storage, querying, and HTML report generation. This hidden expansion of data lifecycle and outputs is dangerous because it affects privacy, retention, and local file-write exposure in ways users may not anticipate from a simple fetcher.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill instructs users to manually copy live authentication cookies into a local config file, which is sensitive credential material. In an agent or shared workstation context, this is particularly dangerous because plaintext cookie storage can enable session hijacking if the file is exposed through backups, logs, repo commits, or over-broad file permissions.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
50% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · SKILL.md (reported line 299)May include surrounding context.

md
},
    {
      "name": "another-api",
      "url": "https://api.example.com/zhihu-hot.json",
      "type": "json",
      "priority": 2
    }

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The file header and all user-facing strings indicate the tool is designed to operate exclusively in Chinese, and there is no natural-language indication that users may choose another language or locale. Under the policy, forcing a specific language without opt-in is a natural-language policy violation unless the locale restriction is explicitly justified.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/save_to_db.py (reported line 30)May include surrounding context.

python
try:
        # 运行Node.js脚本获取数据
        result = subprocess.run(
            ['node', str(fetch_script), str(limit)],
            capture_output=True,
            text=True,

Static analysis

No suspicious patterns detected.