Back to skill

Security audit

妈妈网爬虫

Security checks for vulnerabilities and agentic risk

Overview

The skill is a disclosed mama.cn crawler, but it turns off HTTPS certificate checks while saving fetched articles into a persistent local knowledge folder.

Review before installing. Only use this on a trusted network, treat saved articles as untrusted web content, and avoid feeding the saved Markdown back into an agent as instructions. The safer fix is to remove curl -k, fail closed on HTTPS errors, and make crawling plus persistent writes require explicit user approval.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/mama_crawler.py:40
Finding

Disabled TLS Verification Permits Persistent Knowledge-Base Poisoning

Content
View full analysis

Vulnerability Details

File Location: scripts/mama_crawler.py:40-49, with persistent storage occurring at scripts/mama_crawler.py:121-146
Vulnerability Type: Improper TLS certificate validation and persistent storage of untrusted remote content
Risk Level: High

Vulnerable Code

python
def get(url, timeout=15):
    """发送 HTTP GET 请求(使用 curl 避免 SSL 问题)"""
    import subprocess
    result = subprocess.run(
        ["curl", "-s", "--max-time", str(timeout),
         "-A", UA, "-k",  # -k: 不验证 SSL 证书(妈妈网证书问题)
         url],
        capture_output=True, text=True
    )
    return result.stdout

The remotely retrieved content is subsequently converted into Markdown and stored persistently:

python
def html_to_markdown(title, content, url, source, pub_date, category):
    """转换为 Markdown 格式"""
    md = f"""# {title}

**分类**: {category} | **来源**: {source} | **日期**: {pub_date}
**链接**: {url}

---

{content}

---
*本文由妈妈网爬虫自动采集,存入御知库*
"""
    return md

def save_article(md_content, title, category):
    """保存文章到本地"""
    # 清理文件名
    safe_title = re.sub(r'[\\/:*?"<>|]', '', title)[:50]
    if not safe_title.strip():
        safe_title = f"article_{int(time.time())}"
    filename = CRAWL_DIR / category / f"{safe_title}.md"
    filename.parent.mkdir(parents=True, exist_ok=True)
    filename.write_text(md_content, encoding="utf-8")
    return filename

Technical Analysis

The -k option instructs curl to skip TLS certificate and hostname validation. HTTPS encryption therefore provides no reliable server authentication. An attacker able to intercept or manipulate the network connection can impersonate www.mama.cn and return forged article pages.

Content extracted from those pages is placed directly into Markdown and written beneath ~/.yuzhi/crawls/mama_cn/, which the Skill identifies as a knowledge repository. Removing HTML tags does not neu ...[truncated 1964 chars]

Remediation
View remediation

Remediation Suggestions

  1. Remove the -k option and retain curl's default certificate-chain and hostname verification.
  2. If the remote site has a certificate problem, repair or update the local CA trust store rather than bypassing validation. If operationally appropriate, use a narrowly scoped custom CA bundle through --cacert.
  3. Add --fail-with-body and inspect result.returncode so certificate failures, HTTP errors, and transport failures cause the request to fail closed.
  4. Set an explicit redirect policy. If redirects are enabled, validate every final destination and restrict it to expected HTTPS hosts.
  5. Treat all downloaded article text as untrusted external data. Store explicit provenance metadata and ensure downstream Agent workflows quote or isolate article content from system and user instructions.
  6. Apply downstream prompt-injection defenses before retrieval, including clear data/instruction separation and a policy that commands found inside crawled documents must never be executed.
  7. Consider adding content hashes, crawl timestamps, and verification status so unverified or modified articles can be detected and excluded.

A hardened request implementation should resemble:

python
def get(url, timeout=15):
    import subprocess

    result = subprocess.run(
        [
            "curl",
            "--silent",
            "--show-error",
            "--fail-with-body",
            "--max-time", str(timeout),
            "--user-agent", UA,
            url,
        ],
        capture_output=True,
        text=True,
        check=False,
    )
    if result.returncode != 0:
        raise RuntimeError(f"HTTPS request failed: {result.stderr.strip()}")
    return result.stdout
Vulnerability Patterns
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Tool MisuseTool Parameter Abuse, Chaining Abuse, Unsafe Defaults
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Behavioral ASTexec() Call, eval() Call, Dynamic Import
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (10)

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
78% confidence
Finding

This duplicate finding points to the same unsafe curl configuration: disabling certificate checks removes server authenticity guarantees. In a crawler that ingests remote content into a local repository, that makes malicious content substitution easier on compromised or untrusted networks.

Content

Scanner excerpt · scripts/mama_crawler.py (reported line 41)May include surrounding context.

python
SEARCH_KEYWORDS = ["育儿", "怀孕", "备孕", "新生儿", "孕妇", "月子", "亲子教育"]

def get(url, timeout=15):
    """发送 HTTP GET 请求(使用 curl 避免 SSL 问题)"""
    import subprocess
    result = subprocess.run(
        ["curl", "-s", "--max-time", str(timeout),

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
78% confidence
Finding

This duplicate finding points to the same unsafe curl configuration: disabling certificate checks removes server authenticity guarantees. In a crawler that ingests remote content into a local repository, that makes malicious content substitution easier on compromised or untrusted networks.

Content

Scanner excerpt · scripts/mama_crawler.py (reported line 41)May include surrounding context.

python
SEARCH_KEYWORDS = ["育儿", "怀孕", "备孕", "新生儿", "孕妇", "月子", "亲子教育"]

def get(url, timeout=15):
    """发送 HTTP GET 请求(使用 curl 避免 SSL 问题)"""
    import subprocess
    result = subprocess.run(
        ["curl", "-s", "--max-time", str(timeout),

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
91% confidence
Finding

The skill advertises shell, network, and file-write capable behavior but does not declare any tool scope or permissions boundaries. That makes the skill easier to invoke with broader-than-necessary capabilities and reduces reviewability, increasing the chance of unintended file modification or network access when the skill is triggered.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The trigger conditions are broad enough that ordinary mentions of crawling or parenting knowledge collection could invoke the skill unintentionally. Because the skill performs network access and persistent local writes, accidental activation can cause unwanted scraping activity, unnecessary resource consumption, and silent data accumulation.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

Explicitly disabling SSL certificate validation exposes the crawler to man-in-the-middle attacks, allowing an attacker on the network path to tamper with or inject article content. In this skill, that could poison the stored markdown corpus, mislead downstream users, or introduce malicious content into local knowledge storage.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The crawler explicitly disables TLS certificate verification via curl -k, which allows a man-in-the-middle attacker or hostile network to spoof mama.cn and feed arbitrary content to the crawler. In this skill's context, that content is then written into the local knowledge store, making silent data poisoning and trust-boundary bypass plausible.

Content

No source excerpt is available for this finding.

subprocess module call

Medium
Category
Dangerous Code Execution
Confidence
70% confidence
Finding

subprocess module calls execute external commands. Without careful input validation, this enables command injection.

Content

Scanner excerpt · scripts/mama_crawler.py (reported line 43)May include surrounding context.

python
def get(url, timeout=15):
    """发送 HTTP GET 请求(使用 curl 避免 SSL 问题)"""
    import subprocess
    result = subprocess.run(
        ["curl", "-s", "--max-time", str(timeout),
         "-A", UA, "-k",  # -k: 不验证 SSL 证书(妈妈网证书问题)
         url],

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
79% confidence
Finding

The description notes that content will be stored under a persistent local directory, but it does not prominently warn users about retention, disk usage, or overwrite/duplication behavior. This can lead to users unknowingly persisting scraped data on disk, which is a security and privacy concern in environments where local storage should be tightly controlled.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
67% confidence
Finding

The natural-language description and trigger phrases are entirely in Chinese and imply activation through Chinese commands only, without stating that the skill is region-specific or offering an opt-in language choice. Under the stated policy, a fixed language expectation can be a locale-policy issue unless explicitly justified.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

The file description and built-in categories/keywords indicate the skill is fixed to Chinese parenting content, with no user opt-in or configurable language/locale behavior. Under the policy, forcing a specific language or locale without user choice is a natural-language policy concern unless clearly justified as region-specific.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.