Back to skill

Security audit

topic-research-report

Security checks for vulnerabilities and agentic risk

Overview

This skill is purpose-aligned but needs review because its no-save control is misleading and remote report attachments can still be written locally.

Install only if you are comfortable sending research queries to EastMoney using EM_API_KEY and allowing generated DOCX/PDF attachments to be saved locally. Do not rely on --no-save for non-persistent operation until the implementation is fixed, and avoid opening generated attachments unless you trust the API source.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/get_data.py:154
Finding

Unbounded Remote Response and Attachment Writes

Content
View full analysis

Vulnerability Details

File Location: scripts/get_data.py, lines 154–178 and 199
Vulnerability Type: Unbounded processing and storage of remotely supplied data
Risk Level: Medium

Vulnerable Code

python
def _decode_attachment_base64(data: Dict[str, Any], output_dir: Path) -> List[Dict[str, str]]:
    attachments: List[Dict[str, str]] = []
    output_dir.mkdir(parents=True, exist_ok=True)

    file_map = [
        ("wordBase64", "docx", "DOCX"),
        ("pdfBase64", "pdf", "PDF"),
    ]
    safe_article_id = _safe_article_id(data.get("articleId"))

    for key, ext, ftype in file_map:
        b64_value = data.get(key)
        b64_str = b64_value.strip() if isinstance(b64_value, str) else ""
        if not b64_str:
            continue
        try:
            raw = base64.b64decode(b64_str)
        except Exception:
            continue
        file_name = "{0}_{1}.{2}".format(safe_article_id, ftype.lower(), ext)
        file_path = output_dir / file_name
        file_path.write_bytes(raw)
        attachments.append({"type": ftype, "path": str(file_path)})
    return attachments
python
with urllib_request.urlopen(req, timeout=TIMEOUT_SECONDS) as resp:
    raw_body = resp.read().decode("utf-8", errors="replace")

Technical Analysis

The implementation reads the entire HTTP response into memory without enforcing a maximum response size. It subsequently accepts arbitrarily large Base64 attachment fields, decodes each field entirely in memory, and writes the resulting bytes without a decoded-size limit or available-space check.

The declared attachment type is also trusted without validating file signatures or structure. Consequently, arbitrary remote bytes may be stored under .docx or .pdf extensions. Although the upstream service uses HTTPS, a compromised or malicious API endpoint could still provide hostile content.

Attack Path

  1. An attacker compromises the configured API service or otherwise causes it to return ...[truncated 1160 chars]
Remediation
View remediation

Remediation Suggestions

  • Stream the HTTP response and reject it when a strict maximum byte count is exceeded.
  • Check Content-Length when present, while still enforcing the limit during streaming because the header may be absent or inaccurate.
  • Enforce encoded and decoded attachment-size limits before allocating or writing data.
  • Decode Base64 incrementally rather than creating a complete second in-memory copy.
  • Check available disk capacity before writing attachments.
  • Validate PDF magic bytes and structure before using a .pdf extension.
  • Validate that DOCX output is a well-formed ZIP container with the expected Office document entries.
  • Use strict Base64 validation, such as base64.b64decode(value, validate=True).
  • Write to a temporary file with restrictive permissions and atomically rename it only after successful validation.
  • Surface invalid or oversized attachments as explicit errors instead of silently accepting or discarding them.

T09 · Insecure Skill Coding Practices

Note
Location
scripts/get_data.py:235
Finding

The No-Save Option Does Not Prevent Filesystem Writes

Content
View full analysis

Vulnerability Details

File Location: scripts/get_data.py, lines 235–244, 272–274, and 299
Vulnerability Type: Security-relevant option is accepted but not enforced
Risk Level: Low

Vulnerable Code

python
async def generate_topic_research_report(
    query: str,
    output_dir: Optional[Path] = None,
    save_to_file: bool = True,
) -> Dict[str, Any]:
    query = (query or "").strip()
    if not query:
        return {
            "query": "",
            "title": "",
            "article_id": "",
            "share_url": "",
            "content": "",
            "attachments": [],
            "raw": None,
            "error": "query is empty",
        }

    out_dir = Path(output_dir or DEFAULT_OUTPUT_DIR)
    out_dir.mkdir(parents=True, exist_ok=True)
python
if _has_valid_report(raw):
    attachments = _decode_attachment_base64(data, out_dir)
    if attachments:
        result["attachments"] = attachments
python
result = await generate_topic_research_report(
    query=query,
    save_to_file=not args.no_save
)

Technical Analysis

The command-line layer converts --no-save into save_to_file=False, but generate_topic_research_report() never checks that argument. It creates the output directory unconditionally and invokes _decode_attachment_base64(), which writes returned attachments to disk.

This violates the caller-visible no-save contract and causes externally supplied documents to persist even when the caller explicitly requests no filesystem output. The behavior is documented as a compatibility limitation in SKILL.md, but the option's name and programmatic API still create a reasonable expectation that writes will be suppressed.

Attack Path

  1. A caller invokes the script with --no-save or calls the function with save_to_file=False.
  2. The caller expects execution not to modify the filesystem.
  3. The remote API returns a valid report containing Base64 attachments.
  4. The function unconditi ...[truncated 751 chars]
Remediation
View remediation

Remediation Suggestions

  • Guard directory creation and attachment decoding with if save_to_file:.
  • When save_to_file is false, do not invoke any function that writes to disk.
  • Decide whether no-save mode should omit attachments or return validated attachment bytes in memory, and document that behavior explicitly.
  • Make --no-save help text accurately match the implementation.
  • Add automated tests that snapshot the relevant filesystem before and after no-save execution and assert that no directory or file is created.
  • Add a programmatic test covering generate_topic_research_report(..., save_to_file=False).
  • Consider removing the option entirely if the intended product behavior is always to persist attachments, rather than retaining a misleading control.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Taint TrackingDirect Taint Flow, Variable-Mediated Taint Flow, Credential Exfiltration Chain
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (5)

Tainted flow: 'req' from os.environ.get (line 191, credential/environment) → urllib.request.urlopen (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.

Content

Scanner excerpt · scripts/get_data.py (reported line 202)May include surrounding context.

python
)

    try:
        with urllib_request.urlopen(req, timeout=TIMEOUT_SECONDS) as resp:
            raw_body = resp.read().decode("utf-8", errors="replace")
    except urllib_error.HTTPError as exc:
        err_body = exc.read().decode("utf-8", errors="replace") if exc.fp else ""

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
93% confidence
Finding

The skill uses sensitive capabilities—environment variable access, network access to an external API, and local file writes—without declaring an explicit tool scope or allowlist. This weakens least-privilege controls and makes it easier for the skill, or future modifications to it, to access more resources than intended without clear operator visibility.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The manifest description is entirely in Chinese and defines the skill's purpose and activation conditions only in that language, which can impose a language/locale constraint on users without opt-in. The file does not state that the skill is China-region-specific or that users may choose another language.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This code sends the user-provided natural-language query to an external HTTPS endpoint using the EM_API_KEY, but there is no user-facing disclosure at the point of execution that input will leave the local system. The surrounding CLI output only reports results and errors, so users are not warned about the network transmission of potentially sensitive research queries.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The CLI advertises that '--no-save' disables local file saving, but the flag is ignored by generate_topic_research_report and attachment files are still written whenever the remote API returns base64-encoded documents. This can cause unexpected persistence of potentially sensitive report contents to disk, violating user expectations and creating privacy/data-handling risk.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.