T09 · Insecure Skill Coding Practices
- Location
xhs-keepalive.py:127- Finding
Authentication cookies exposed through child-process command-line arguments
- Content
View full analysis
. ...[truncated 895 chars]- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This skill is a disclosed Xiaohongshu scraper, but it handles live login cookies and persistent token/session data in ways users should review carefully before installing.
Install only if you are comfortable giving the skill access to an authenticated Xiaohongshu session. Do not use it on shared, CI, or heavily monitored machines; keep the data directory private, avoid publishing harvest outputs, rotate or revoke cookies after use, and review whether bulk scraping, comment collection, token caching, proxy/IP changes, and captcha-avoidance behavior comply with your obligations and Xiaohongshu's rules.
xhs-keepalive.py:127Authentication cookies exposed through child-process command-line arguments
xhs-harvest.py:214Access tokens and authenticated browser state stored without explicit restrictive permissions
docs/pitfalls.md:94Optional diagnostic command discloses public IP and network metadata to a third-party service
The documented behavior goes beyond simple content reading: it instructs users to export authenticated cookies, persist them locally, and use helper scripts to inject and validate session state. That mismatch is security-relevant because operators may trust the skill as a read-only scraper while it actually manages sensitive credentials and browser session material, creating elevated risk of credential leakage or unintended account use.
Remote code is downloaded and executed. This bypasses code review and could introduce malicious code.
# 1. 看 IP / 城市
curl -s "https://ipinfo.io/json" | python3 -m json.tool | head -10
# 2. 看 agent-browser 当前 cookies
agent-browser cookies get
The contribution instructions are written almost entirely in Chinese and assume contributors can read that language, with no English alternative or opt-in language choice. This is a natural-language locale policy concern because it imposes a specific language on users without explicit justification in the file.
The README explicitly instructs users to export live Xiaohongshu authentication cookies, including session-bearing tokens, into a local plaintext file. Those cookies appear sufficient to impersonate the user for scraping and may grant account access until expiry, so storing them this way creates real credential-exposure risk if the host, repo directory, logs, shell history, backups, or other local processes are compromised. The skill context makes this more dangerous, not less, because the skill requires strong login state and repeatedly tells users to handle short-lived but privileged auth material.
The skill explicitly relies on shell execution, network access, environment variables, and writing cookie/session data to disk, yet it declares no tool scope or permission boundaries. In a skill that handles authenticated scraping with injected login cookies, the lack of explicit allowed-tools/permissions increases the attack surface and makes it harder for a runtime or reviewer to constrain misuse.
SQP-3 applies to all file types and covers language or locale policy violations. This markdown file begins with and continues in Chinese-only instructions, with no indication that users can opt into another language or that the skill is intentionally region-specific in a documented way.
subprocess module calls execute external commands. Without careful input validation, this enables command injection.
def run(cmd, timeout=30, check=True):
"""跑 agent-browser 命令"""
r = subprocess.run(cmd, capture_output=True, text=True, timeout=timeout)
if check and r.returncode != 0:
print(f" stdout: {r.stdout[:300]}")
print(f" stderr: {r.stderr[:300]}")
The script automatically checks for and loads authentication cookies into agent-browser without an explicit runtime warning or consent gate. In a skill that requires login cookies, silent injection increases the chance that users expose an authenticated session to automated scraping actions without understanding the privacy, account, or policy consequences.
subprocess module calls execute external commands. Without careful input validation, this enables command injection.
r = run(['agent-browser', 'cookies', 'get'], timeout=10)
if not r.stdout.strip() or 'web_session' not in r.stdout:
print("Loading cookies into agent-browser...")
subprocess.run(
['python3', str(Path(__file__).parent / 'xhs-keepalive.py'), 'load'],
timeout=30
)
The search flow makes authenticated requests through agent-browser and supports writing scraped results to disk, but does not clearly disclose that protected session context may be used or that third-party content may be persisted locally. This creates a meaningful privacy and consent risk, especially because the skill is designed to operate with injected login cookies and can collect user-generated content at scale.
The note-detail path can collect note text, metadata, and comments from third parties and save them to disk without a clear user-facing warning. Because the skill operates under an authenticated session and is explicitly built to bypass some access controls using xsec_token flows, the privacy and policy implications are elevated beyond ordinary local processing.
The skill reads the logged-in user's own Xiaohongshu user ID from document.cookie to exclude that identity during DOM candidate filtering. Even though it is not exfiltrated directly here, accessing authenticated cookie-derived identity data exceeds the stated fetch-only purpose and establishes a pattern of harvesting session-linked metadata unnecessarily.
The script advertises a 'harvest' workflow that stores harvested notes, user data, comments, search results, and reports under persistent directories, but it does not present an explicit privacy or data-handling warning at the point of use. In this skill context, silent bulk persistence of third-party content and comments is security-relevant because it facilitates unauthorized retention, redistribution, and secondary analysis of personal/content data.
subprocess module calls execute external commands. Without careful input validation, this enables command injection.
def run_fetch(args, timeout=60):
"""subprocess 调 xhs-fetch.py, 捕获 stdout (JSON 路径在 ok 行里有)"""
cmd = ['python3', str(FETCH_SCRIPT)] + args
return subprocess.run(cmd, capture_output=True, text=True, timeout=timeout)
def safe_int(s, default=0):
The hot workflow expands a nominal search/read capability into bulk harvesting, persistent storage, ranking, and report generation for many notes and comments. In context, this materially increases privacy, compliance, and misuse risk because the skill is explicitly designed to collect and retain datasets rather than just transiently view content.
The user workflow claims to 'view a user' but actually enumerates many works, resolves tokens, fetches note details, collects comments, caches tokens, and writes a structured local archive and report. That broader behavior makes large-scale profile/content aggregation easier and more dangerous than the manifest suggests, especially because it targets a single person's full body of work.
Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.
注意:
- 不依赖任何第三方库 (除 agent-browser 自身)
- cookies 必须 chmod 600 (含 web_session + id_token 等敏感字段)
- xhs cookie 6-12 小时会失效 (web_session 短效),需要定期重导
"""
Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.
注意:
- 不依赖任何第三方库 (除 agent-browser 自身)
- cookies 必须 chmod 600 (含 web_session + id_token 等敏感字段)
- xhs cookie 6-12 小时会失效 (web_session 短效),需要定期重导
"""
Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.
注意:
- 不依赖任何第三方库 (除 agent-browser 自身)
- cookies 必须 chmod 600 (含 web_session + id_token 等敏感字段)
- xhs cookie 6-12 小时会失效 (web_session 短效),需要定期重导
"""
Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.
注意:
- 不依赖任何第三方库 (除 agent-browser 自身)
- cookies 必须 chmod 600 (含 web_session + id_token 等敏感字段)
- xhs cookie 6-12 小时会失效 (web_session 短效),需要定期重导
"""
Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.
注意:
- 不依赖任何第三方库 (除 agent-browser 自身)
- cookies 必须 chmod 600 (含 web_session + id_token 等敏感字段)
- xhs cookie 6-12 小时会失效 (web_session 短效),需要定期重导
"""
subprocess module calls execute external commands. Without careful input validation, this enables command injection.
return 1
# 清掉旧的
subprocess.run(['agent-browser', 'cookies', 'clear'],
capture_output=True, text=True)
# 读 Netscape
subprocess module calls execute external commands. Without careful input validation, this enables command injection.
continue
domain, include_subdomain, path, secure, expires, name, value = parts[:7]
flags = _cookie_to_agent_browser_flags(name)
r = subprocess.run(
['agent-browser', 'cookies', 'set', name, value] + flags,
capture_output=True, text=True
)
subprocess module calls execute external commands. Without careful input validation, this enables command injection.
# 打开一个低风险页面测试
print("Probing https://www.xiaohongshu.com ...")
r = subprocess.run(
['agent-browser', 'open', 'https://www.xiaohongshu.com/explore'],
capture_output=True, text=True, timeout=30
)
subprocess module calls execute external commands. Without careful input validation, this enables command injection.
return 1
# 检查 title
r = subprocess.run(
['agent-browser', 'eval', 'document.title'],
capture_output=True, text=True, timeout=10
)
No suspicious patterns detected.