T09 · Insecure Skill Coding Practices
- Location
scripts/extract_post.py:23- Finding
Authenticated cookies can be disclosed to an attacker-controlled host
- Content
View full analysis
tuple[str, str]: """Extract note ID and xsec_token from URL.""" # Pattern: xiaohongshu.com/explore/xxx or xiaohongshu.com/discovery/item/xxx patterns = [ r'xiaohongshu\.com/explore/([a-f0-9]{24})', r'xiaohongshu\.com/discovery/item/([a-f0-9]{24})', ] for p in patterns: m = re.search(p, url) if m: post_id = m.group(1) xsec_m = re.search(r'xsec_token=([a-f0-9]+)', url) xsec = xsec_m.group(1) if xsec_m else '' return post_id, xsec raise ValueError(f"Cannot parse post ID from URL: {url}") ``` ```python req = urllib.request.Request(url) req.add_header('Cookie', cookie_str) req.add_header('User-Agent', 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36') req.add_header('Referer', 'https://www.xiaohongshu.com/') ``` ```python cookie_str = load_cookies(cookie_path) post_id, xsec = extract_post_id(args.url) url_with_token = args.url if xsec: url_with_token = args.url if 'xsec_token' in args.url else f"{args.url}&xsec_token={xsec}" try: note = fetch_post(url_with_token, cookie_str) ``` ### Technical Analysis The validation logic uses an unanchored regular-expression search against the complete URL. It verifies only that the text contains a Xiaohongshu-looking path; it does not parse or validate the actual hostname, scheme, port, or user-information component. For example, this attacker-controlled URL passes the post-ID check: ```text https://attacker.example/xiaohongshu.com/explore/0123456789abcdef01234567 ``` The complete Xiaohongshu cookie string is ...[truncated 1907 chars]- Remediation
View remediation
