Back to skill

Security audit

Zhihu Search

Security checks for vulnerabilities and agentic risk

Overview

This appears to be a real Zhihu content-fetching skill, but it handles live account cookies and includes unsafe persistence and process-control behavior that should be reviewed before use.

Install only if you are comfortable giving the skill a live Zhihu session cookie. Prefer a dedicated low-risk account, do not run install-cron or keepalive until the cron quoting and process-kill behavior are fixed, keep cookie and state files outside shared or committed directories with owner-only permissions, and verify no sensitive data is added to git.

Vulnerability Patterns
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (4)

T09 · Insecure Skill Coding Practices

Error
Location
zhihu-fetch.py:49
Finding

Authenticated Zhihu Cookie May Be Exposed Through Redirect Following

Content
View full analysis
str: if not COOKIE_FILE.exists(): sys.exit(f"Cookie file not found: {COOKIE_FILE}") return COOKIE_FILE.read_text().strip() def curl_json(url: str, cookie: str) -> dict: r = subprocess.run( ["curl", "-sL", "-A", UA, "-H", f"Cookie: {cookie}", url], capture_output=True, text=True, timeout=30 ) try: return json.loads(r.stdout) except json.JSONDecodeError as e: sys.exit(f"Invalid JSON response: {e}") def curl_text(url: str, cookie: str, timeout: int = 30) -> str: r = subprocess.run( ["curl", "-sL", "-A", UA, "-H", f"Cookie: {cookie}", url], capture_output=True, text=True, timeout=timeout ) return r.stdout ``` ### Technical Analysis The Skill loads the complete authenticated Zhihu cookie string and manually adds it as an HTTP `Cookie` header. The `-L` option instructs curl to follow redirects. All initial request URLs are internally constructed from fixed HTTPS Zhihu endpoints. Therefore, the behavior is directly related to the declared functionality, and no intentional third-party exfiltration destination was identified. However, the code does not independently validate the destination of each redirect or restrict redirects to an explicit host allowlist. Security consequently depends on curl's version-specific handling of sensitive custom headers during cross-origin redirects. This is unsafe for a credential capable of representing the user's authenticated Zhihu session. ### Attack Path 1. The user places an authenticated Zhihu cookie in the configured cookie file. 2. The Skill calls a fixed Zhihu API or article endpoint. 3. The endpoint, a compromised upstream component, or another network-layer condition returns an HTTP redirect. 4. Curl f ...[truncated 722 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
paths.py:34
Finding

Sensitive Cookie and Browser-State Storage Does Not Enforce Restrictive Permissions

Content
View full analysis
Path: p.mkdir(parents=True, exist_ok=True) return p DATA_DIR = _resolve_data_dir() COOKIE_RAW = DATA_DIR / "cookies-raw.txt" COOKIE_FILE = _resolve_cookie_file(DATA_DIR) STATE_DIR = _ensure_dir(DATA_DIR / "state") STATE_FILE = STATE_DIR / "zhihu.state.json" EXPORTS_DIR = _ensure_dir(DATA_DIR / "exports") ANSWERS_DIR = _ensure_dir(DATA_DIR / "answers") ARTICLES_DIR = _ensure_dir(DATA_DIR / "articles") COLUMNS_DIR = _ensure_dir(DATA_DIR / "columns") COMMENTS_DIR = _ensure_dir(DATA_DIR / "comments") TMP_DIR = Path("/tmp/zhihu") _ensure_dir(TMP_DIR) ``` Browser state is subsequently written without an explicit permission correction: ```python run(["ab", "state", "save", str(STATE_FILE)]) ``` The documentation states that runtime data is excluded from Git, but the audited directory structure contains no `.gitignore` file. ### Technical Analysis The data and state directories are created using the process's default umask. The code does not explicitly enforce mode `0700` for directories or mode `0600` for the cookie, raw-cookie, browser-state, and log files. The browser-state file may contain authenticated browser data comparable in sensitivity to the cookie file. The default storage location is also inside the project directory. Although the documentation tells users to apply `chmod 600` to `cookies.txt`, this is a manual instruction rather than an enforced security control. The missing `.gitignore` conflicts with the documentation's assertion that `data/` is excluded. This increases the risk that credentials, state, or fetched data will be accidentally committed. ### Attack Path 1. The Skill is run under a permissive umask or in a shared project directory. 2. Runtime directories and browser-state files are created without ...[truncated 962 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
keepalive.py:151
Finding

Unquoted Environment-Derived Paths Permit Persistent Cron Command Injection

Content
View full analysis
>{LOG_FILE} 2>&1' r = run(["crontab", "-l"], check=False) existing = r.stdout if r.returncode == 0 and r.stdout else "" if "keepalive.py refresh" in existing: return new_content = ( existing.rstrip("\n") + "\n" + cron_line + "\n" ) if existing.strip() else (cron_line + "\n") proc = subprocess.Popen( ["crontab", "-"], stdin=subprocess.PIPE, text=True ) proc.communicate(new_content) ``` `LOG_FILE` is derived from an environment variable: ```python LOG_FILE = Path(os.environ.get("ZHIHU_LOG", str(STATE_DIR / "cron.log"))) ``` ### Technical Analysis Cron executes each scheduled command through a shell. The generated command directly interpolates `python_bin`, `SELF_PATH`, and `LOG_FILE` without shell quoting. `ZHIHU_LOG` is explicitly environment-controlled. The source installation path can also contain spaces or shell metacharacters. Consequently, these values are treated as shell syntax rather than strictly as path arguments when cron executes the entry. Cron installation is user-triggered, documented, and functionally related to maintaining the Zhihu session. It is not a hidden backdoor. Nevertheless, it creates cross-session persistence, and the unsafe interpolation can turn the documented persistence feature into persistent arbitrary command execution. ### Attack Path 1. An attacker influences the environment or deployment configuration used to invoke the Skill, particularly `ZHIHU_LOG`, or ...[truncated 912 chars]
Remediation
View remediation

T05 · Unauthorized Access and Privilege Escalation

Warning
Location
keepalive.py:51
Finding

Broad Process Matching Can Force-Terminate Unrelated Agent Browser Sessions

Content
View full analysis
bool: r = run(["pgrep", "-f", AB_PROCESS]) pids = [ p for p in r.stdout.split() if p not in (str(os.getpid()), str(os.getppid())) ] return bool(pids) def kill_daemon() -> None: r = run(["pgrep", "-f", AB_PROCESS]) pids = [ p for p in r.stdout.split() if p not in (str(os.getpid()), str(os.getppid())) ] for pid in pids: run(["kill", "-9", pid]) time.sleep(1) ``` The function is invoked whenever the Skill loads and opens browser state: ```python def ab_load_and_open(url: str = TARGET_DEFAULT) -> int: if not STATE_FILE.exists(): return 1 kill_daemon() r = run(["ab", "--state", str(STATE_FILE), "open", url], capture=False) return r.returncode ``` ### Technical Analysis The Skill uses `pgrep -f` to select every process whose complete command line contains `agent-browser-linux`. It excludes only the current process and its parent. There is no Skill-specific PID file, session identifier, process start-time check, ownership marker, or verification that a matched process was created by this Skill. Every match is terminated using `SIGKILL`, which prevents graceful shutdown and state cleanup. This behavior exceeds the minimum privileges required for read-only Zhihu retrieval. It can disrupt unrelated Agent Browser tasks running under the same operating-system account. ### Attack Path 1. Another agent, Skill, or user process starts an `agent-browser-linux` daemon under the same account. 2. This Skill executes `open`, `refresh`, or `check`. 3. `pgrep -f agent-browser-linux` returns the unrelated daemon's PID. 4. `kill -9` immediately terminates that process. 5. The unrelated browser operation loses its active session and may lose ...[truncated 648 chars]
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Tool MisuseTool Parameter Abuse, Chaining Abuse, Unsafe Defaults
  • Rogue AgentSelf-Modification, Session Persistence
Findings (32)

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The declared purpose is content fetching, but the documentation also includes session persistence, browser-driven login validation, daemon/process management, and cron-based keepalive behavior. This mismatch is security-relevant because users and agents may authorize a seemingly read-only scraping skill while it actually performs persistent account automation and local system state changes.

Content

No source excerpt is available for this finding.

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
85% confidence
Finding

The documented cleanup command uses 'rm -rf' with environment-variable-expanded paths. In context it is intended to delete the skill's own cache directories, but if $SKILL is unset, malformed, or unexpectedly points elsewhere, an agent or user could remove unintended data; shell-based destructive cleanup is therefore a real safety issue even if not overtly malicious.

Content

Scanner excerpt · SKILL.md (reported line 205)May include surrounding context.

md
## 数据保留与清理

`$SKILL/data/` 是 git 排除目录(仅在本地存在)。如果想长期保留某个 topic 的抓取结果,建议在调用时显式加 `--md-out /path/to/keep.md`,把摘要拷到自己的目录。原 JSON 缓存可以随时 `rm -rf $SKILL/data/answers $SKILL/data/articles $SKILL/data/columns $SKILL/data/comments` 清理。

## 死胡同

YARA rule 'backdoor_persistence': Backdoor persistence with malicious payloads (shell commands, SSH key injection, hidden root users) [malware]

High
Category
YARA Match
Confidence
75% confidence
Finding

YARA rule matched a known malware signature (reverse shell, backdoor, ransomware, C2 framework, or info stealer).

Content

Scanner excerpt · keepalive.py (reported line 173)May include surrounding context.

python
n line:
                print(f"  {line}")
        return

    # 追加 (用 stdin 而非 crontab 文件,避免权限问题)
    new_content = (existing.rstrip("\n") + "\n" + cron_line + "\n") if existing.strip() else (cron_line + "\n")
    proc = subprocess.Popen(["crontab", "-"], stdin=subprocess.PIPE, text=True)
    proc.communicate(new_content)
    if proc.returncode != 0:
        print(f"✗ crontab 安装失败 (rc={proc.returncode})", file=sys.stderr)
        sys.exit(1)
    print(f"✓ 已安装 cron: 每天 9:00 触发 z_c0 续期")
    print(f"  日志: {LOG_FILE}")

def cmd_status(args):
    """看 state / daemon / cron 状态"""
    print("=== state 文件 ===")
    if STATE_FILE.exists():
        st = STATE_FILE.stat()
        age_h = (time.time() - st.st_mtime) / 3600
        print(f"  {STATE_FILE}")
        print(f"  {st.st_size/1024:.1f} KB, 距今 {age_h:.1f} 小时")
    else:
        print("  ✗ 不存在")

    print("\n=== ab daemon 状态 ===")
    if is_daemon_r

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The contribution guide instructs users to place live Zhihu authentication cookies into a local plaintext file for testing, but gives no warning that these tokens are sensitive session credentials. If that file is exposed through backups, accidental commits, shell history, shared workstations, or permissive file handling elsewhere, an attacker could reuse the cookies to access the contributor's account session.

Content

No source excerpt is available for this finding.

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
80% confidence
Finding

Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.

Content

Scanner excerpt · CONTRIBUTING.md (reported line 14)May include surrounding context.

md
# 准备 cookie (跟正式使用一样)
mkdir -p data
echo "_xsrf=xxx; d_c0=xxx; z_c0=xxx; ..." > data/cookies.txt
chmod 600 data/cookies.txt

# 跑测试 (随便抓一个)
python3 zhihu-fetch.py hotlist --limit 3

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
80% confidence
Finding

Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.

Content

Scanner excerpt · README.md (reported line 114)May include surrounding context.

md
# 准备 cookie (跟正式使用一样)
mkdir -p data
echo "_xsrf=xxx; d_c0=xxx; z_c0=xxx; ..." > data/cookies.txt
chmod 600 data/cookies.txt

# 跑测试 (随便抓一个)
python3 zhihu-fetch.py hotlist --limit 3

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
80% confidence
Finding

Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.

Content

Scanner excerpt · SKILL.md (reported line 79)May include surrounding context.

md
# 准备 cookie (跟正式使用一样)
mkdir -p data
echo "_xsrf=xxx; d_c0=xxx; z_c0=xxx; ..." > data/cookies.txt
chmod 600 data/cookies.txt

# 跑测试 (随便抓一个)
python3 zhihu-fetch.py hotlist --limit 3

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
80% confidence
Finding

Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.

Content

Scanner excerpt · SKILL.md (reported line 89)May include surrounding context.

md
# 准备 cookie (跟正式使用一样)
mkdir -p data
echo "_xsrf=xxx; d_c0=xxx; z_c0=xxx; ..." > data/cookies.txt
chmod 600 data/cookies.txt

# 跑测试 (随便抓一个)
python3 zhihu-fetch.py hotlist --limit 3

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
80% confidence
Finding

Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.

Content

Scanner excerpt · SKILL.md (reported line 159)May include surrounding context.

md
# 准备 cookie (跟正式使用一样)
mkdir -p data
echo "_xsrf=xxx; d_c0=xxx; z_c0=xxx; ..." > data/cookies.txt
chmod 600 data/cookies.txt

# 跑测试 (随便抓一个)
python3 zhihu-fetch.py hotlist --limit 3

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
84% confidence
Finding

The skill description and usage guidance are presented entirely in Chinese, and the title explicitly frames the skill as a Zhihu content tool without any opt-in or alternative language support. Under the policy, a forced language/locale experience is a finding unless the locale constraint is clearly documented and justified or the user is offered a choice.

Content

No source excerpt is available for this finding.

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
80% confidence
Finding

Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.

Content

Scanner excerpt · README.md (reported line 40)May include surrounding context.

或者一次性命令行:

bash
echo "_xsrf=xxx; d_c0=xxx; z_c0=xxx; ..." > ~/.pi/agent/skills/zhihu-search/data/cookies.txt
chmod 600 ~/.pi/agent/skills/zhihu-search/data/cookies.txt

2. 验证(3 秒)

Sudo/Root Execution

Medium
Category
Privilege Escalation
Confidence
80% confidence
Finding

Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.

Content

Scanner excerpt · README.md (reported line 138)May include surrounding context.

或者一次性命令行:

bash
echo "_xsrf=xxx; d_c0=xxx; z_c0=xxx; ..." > ~/.pi/agent/skills/zhihu-search/data/cookies.txt
chmod 600 ~/.pi/agent/skills/zhihu-search/data/cookies.txt

2. 验证(3 秒)

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
93% confidence
Finding

The skill advertises substantial capabilities—environment-variable use, filesystem reads/writes, shell execution, and network access—without declaring an explicit tool scope or permissions boundary. That makes the operational trust model unclear and increases the chance an agent will grant or use broader privileges than the user expects, especially because the skill also handles authenticated cookies and persistence data.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill instructs the user to inject authenticated Zhihu cookies, which are bearer-style session secrets, but provides no prominent warning about credential sensitivity, reuse risk, exfiltration risk, or account takeover implications. In this context, the missing guidance materially increases the chance of unsafe handling of live authentication tokens.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The documentation recommends browser/account automation and daily cookie keepalive without a clear warning that these actions may affect the user account, prolong active sessions, or violate user expectations and platform policies. Because the skill operates with authenticated session material, normalizing silent persistence makes the risk more significant than in a purely local or anonymous workflow.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · keepalive.py (reported line 48)May include surrounding context.

python
# ===== 底层工具 =====
def run(cmd: list[str], check: bool = False, capture: bool = True) -> subprocess.CompletedProcess:
    """统一运行 shell 命令"""
    return subprocess.run(cmd, capture_output=capture, text=True, check=check)

def is_daemon_running() -> bool:
    """检查 ab daemon 是否在跑 (排除自己 + 父进程)"""

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The script enumerates and force-kills local processes matching the browser daemon name, which is an OS-level process-management capability not necessary for simple Zhihu content retrieval. This can disrupt unrelated user activity and is especially risky in an agent skill because it affects processes outside the immediate task boundary.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

Using SIGKILL without warning or confirmation immediately terminates matching browser daemon processes and bypasses graceful shutdown. In a user environment this can cause data loss, break active sessions, and unexpectedly affect other workflows using the same agent-browser daemon.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The save flow persists browser login state to disk, which may include highly sensitive authenticated session material, without an explicit warning about what is being stored or the security implications. In the context of a Zhihu login helper, this increases the chance of accidental credential exposure to other local users, backups, or later compromise of the host.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The check command is described as read-only, but it calls ab_load_and_open(), which first kills the running daemon and opens a browser session. That mismatch between documented behavior and actual side effects can mislead users into invoking a disruptive action they would not otherwise permit.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The helper installs a persistent daily cron job that survives the current session and modifies the user's environment beyond the core content-fetching purpose of the skill. In the context of a scraping skill, persistence is higher risk because it creates ongoing execution and continued access to stored authenticated state without requiring the user to re-authorize each run.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The script modifies the user's crontab to create persistent scheduled execution without requiring confirmation or prominently warning that a recurring task will be added. This is dangerous because it establishes silent ongoing behavior that continues to refresh and preserve authenticated state after the user may have forgotten about it.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
90% confidence
Finding

The cron-management logic supports ongoing session persistence by regularly refreshing state associated with an authenticated Zhihu session. In a skill whose purpose is content fetching, maintaining long-lived authenticated access increases exposure if the machine or stored state is later compromised.

Content

Scanner excerpt · keepalive.py (reported line 157)May include surrounding context.

python
python_bin = sys.executable if Path(sys.executable).is_absolute() else "/usr/bin/python3"
    cron_line = f'0 9 * * * {python_bin} {SELF_PATH} refresh >>{LOG_FILE} 2>&1'

    # 读现有 crontab (没有 crontab 时 crontab -l 返回非 0)
    r = run(["crontab", "-l"], check=False)
    existing = r.stdout if r.returncode == 0 and r.stdout else ""

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · keepalive.py (reported line 170)May include surrounding context.

python
# 追加 (用 stdin 而非 crontab 文件,避免权限问题)
    new_content = (existing.rstrip("\n") + "\n" + cron_line + "\n") if existing.strip() else (cron_line + "\n")
    proc = subprocess.Popen(["crontab", "-"], stdin=subprocess.PIPE, text=True)
    proc.communicate(new_content)
    if proc.returncode != 0:
        print(f"✗ crontab 安装失败 (rc={proc.returncode})", file=sys.stderr)

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The natural-language documentation and comments in this file are entirely in Chinese, which imposes a language constraint without any opt-in or alternative. The policy specifically calls for flagging language or locale requirements when the skill forces a specific language without user choice.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.