T09 · Insecure Skill Coding Practices
- Location
scripts/mama_crawler.py:40- Finding
Disabled TLS Verification Permits Persistent Knowledge-Base Poisoning
- Content
View full analysis
Vulnerability Details
File Location:
scripts/mama_crawler.py:40-49, with persistent storage occurring atscripts/mama_crawler.py:121-146
Vulnerability Type: Improper TLS certificate validation and persistent storage of untrusted remote content
Risk Level: HighVulnerable Code
python def get(url, timeout=15): """发送 HTTP GET 请求(使用 curl 避免 SSL 问题)""" import subprocess result = subprocess.run( ["curl", "-s", "--max-time", str(timeout), "-A", UA, "-k", # -k: 不验证 SSL 证书(妈妈网证书问题) url], capture_output=True, text=True ) return result.stdoutThe remotely retrieved content is subsequently converted into Markdown and stored persistently:
python def html_to_markdown(title, content, url, source, pub_date, category): """转换为 Markdown 格式""" md = f"""# {title} **分类**: {category} | **来源**: {source} | **日期**: {pub_date} **链接**: {url} --- {content} --- *本文由妈妈网爬虫自动采集,存入御知库* """ return md def save_article(md_content, title, category): """保存文章到本地""" # 清理文件名 safe_title = re.sub(r'[\\/:*?"<>|]', '', title)[:50] if not safe_title.strip(): safe_title = f"article_{int(time.time())}" filename = CRAWL_DIR / category / f"{safe_title}.md" filename.parent.mkdir(parents=True, exist_ok=True) filename.write_text(md_content, encoding="utf-8") return filenameTechnical Analysis
The
-koption instructscurlto skip TLS certificate and hostname validation. HTTPS encryption therefore provides no reliable server authentication. An attacker able to intercept or manipulate the network connection can impersonatewww.mama.cnand return forged article pages.Content extracted from those pages is placed directly into Markdown and written beneath
~/.yuzhi/crawls/mama_cn/, which the Skill identifies as a knowledge repository. Removing HTML tags does not neu ...[truncated 1964 chars]- Remediation
View remediation
Remediation Suggestions
- Remove the
-koption and retain curl's default certificate-chain and hostname verification. - If the remote site has a certificate problem, repair or update the local CA trust store rather than bypassing validation. If operationally appropriate, use a narrowly scoped custom CA bundle through
--cacert. - Add
--fail-with-bodyand inspectresult.returncodeso certificate failures, HTTP errors, and transport failures cause the request to fail closed. - Set an explicit redirect policy. If redirects are enabled, validate every final destination and restrict it to expected HTTPS hosts.
- Treat all downloaded article text as untrusted external data. Store explicit provenance metadata and ensure downstream Agent workflows quote or isolate article content from system and user instructions.
- Apply downstream prompt-injection defenses before retrieval, including clear data/instruction separation and a policy that commands found inside crawled documents must never be executed.
- Consider adding content hashes, crawl timestamps, and verification status so unverified or modified articles can be detected and excluded.
A hardened request implementation should resemble:
python def get(url, timeout=15): import subprocess result = subprocess.run( [ "curl", "--silent", "--show-error", "--fail-with-body", "--max-time", str(timeout), "--user-agent", UA, url, ], capture_output=True, text=True, check=False, ) if result.returncode != 0: raise RuntimeError(f"HTTPS request failed: {result.stderr.strip()}") return result.stdout- Remove the
