Back to skill

Security audit

医药情报简报(Pharma Intelligence Brief)

Security checks for vulnerabilities and agentic risk

Overview

The skill has a coherent pharmaceutical report purpose, but its HTML renderer can turn untrusted source links into executable browser content.

Install only if you are comfortable reviewing or patching the HTML renderer first. Reports should be generated from trusted data and opened in an isolated context until URL escaping and safe-scheme validation are added, especially if reports will be hosted inside an authenticated internal site.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/generate_report.py:74
Finding

Stored HTML and Script Injection Through Untrusted Patent Links

Content
View full analysis
str: """Render a patent number as a clickable link.""" if not pn: return pn or "" if url: href = url elif patent_id: href = f"https://eureka.zhihuiya.com/view/#/fullText?patentId={patent_id}" else: import urllib.parse href = f"https://eureka.zhihuiya.com/patent-search/search?q={urllib.parse.quote(pn)}" return ( f'{esc(pn)}' ) ``` ### Technical Analysis The `url` and `patent_id` fields originate from report input data and are incorporated into the quoted `href` attribute without HTML attribute escaping. Although the displayed patent number is escaped with `esc(pn)`, the actual link target is not. An attacker who can influence a retrieved patent record or the normalized input JSON can supply an attribute-breaking value such as: ```text https://example.invalid/" onmouseover="alert(document.domain) ``` This produces attacker-controlled HTML attributes in the generated report. The renderer also does not restrict URL schemes. Consequently, a value using a browser-executable scheme such as `javascript:` could execute when a reader follows the link. HTML escaping alone is insufficient for that second case; explicit scheme validation is required. Other report sections escape URL attribute values, but they also lack a protocol allowlist. The patent renderer is more directly exploitable because it does not perform even attribute escaping. The base64 operation elsewhere in this script is not part of this vulnerability. It only embeds an RDKit-generated molecular S ...[truncated 1582 chars]
Remediation
View remediation
' f'{esc(pn)}' ) ``` 2. **Allowlist safe URL schemes** Use `urllib.parse.urlsplit()` and accept only `https` and, if operationally necessary, `http`. Reject `javascript:`, `data:`, `file:`, and other schemes. Invalid links should be rendered as escaped plain text rather than as anchors. ```python from urllib.parse import urlsplit def safe_web_url(value: str) -> str: value = (value or "").strip() try: parsed = urlsplit(value) except ValueError: return "" if parsed.scheme.lower() not in {"https", "http"}: return "" if not parsed.netloc: return "" return value ``` 3. **Validate patent identifiers** Validate `patent_id` against the identifier format expected by the destination service before including it in a generated URL. Encode it with `urllib.parse.quote()` even after format validation. 4. **Apply the same URL policy globally** Centralize URL rendering and apply the safe-scheme allowlist to patent, paper, clinical-trial, deal, news, and source links. Existing HTML escaping in those sections does not prevent executable URL schemes. 5. **Add security regression tests** Test at least the following inputs: - Quotes that attempt to break out of `href` - Injected event-handler attributes - `javascript:` URLs - `data:` URLs - Malformed URLs - Unexpected patent identifiers - Valid HTTP and HTTPS links 6. **Add defense-in-depth controls** When reports are served over HTTP, apply a restrictive Content Security Policy that blocks inline script and limits navigation and resource loading to approve ...[truncated 138 chars]
Vulnerability Patterns
  • Taint TrackingDirect Taint Flow, Variable-Mediated Taint Flow, Credential Exfiltration Chain
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (10)

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 36)May include surrounding context.

md
10. 验证 JSON 数据约定,并使用 `scripts/generate_report.py` 渲染。

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 59)May include surrounding context.

md
10. 验证 JSON 数据约定,并使用 `scripts/generate_report.py` 渲染。

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
87% confidence
Finding

The skill instructs the agent to read local reference files and invoke a local rendering script, which implies file access and code-execution-adjacent behavior, yet it declares no explicit tool scope or permissions boundary. Without an allowlist, an agent/runtime may grant broader file, environment, or network access than is actually needed, increasing the blast radius if the skill is misused or prompt-injected.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The file forces a specific language/locale for the skill's user-facing description and operating instructions. Under the policy, language constraints should not be imposed without explicitly offering the user a language choice or documenting a justified locale-specific scope.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The manifest description is written as a Chinese-only skill description and presents the skill as producing a pharmaceutical intelligence brief for Chinese-speaking users, with no indication that users may choose another language. This creates a locale/language policy concern because the skill appears to impose a specific language without explicit user opt-in.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
87% confidence
Finding

The operational instructions, triggers, parameters, and report requirements are all specified in Chinese, and the file does not document that the skill is China-region-specific or that the user can opt into Chinese output. For a general-purpose skill, this is a natural-language policy issue because it effectively constrains interaction to one language by default.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The HTML output sets lang="zh-CN", which forces a specific language/locale in the generated report. In this file there is no user opt-in, configuration switch, or documented justification that the skill is intended only for a Chinese-language or China-specific compliance context.

Content

No source excerpt is available for this finding.

Tainted flow: 'out_path' from os.environ.get (line 729, credential/environment) → open (file write)

Medium
Category
Data Flow
Confidence
89% confidence
Finding

The script writes to a path derived from the EUREKA_PYTHON_OUTPUT_DIR environment variable without validating that the destination stays within an expected safe directory. If an attacker can influence the process environment or invoke the script in a privileged context, they could redirect output to unintended filesystem locations and overwrite arbitrary writable files.

Content

Scanner excerpt · scripts/generate_report.py (reported line 731)May include surrounding context.

python
today = datetime.now().strftime("%Y%m%d")
        out_path = os.path.join(out_dir, f"pharma_intel_brief_{today}.html")

    with open(out_path, "w", encoding="utf-8") as f:
        f.write(html_content)

    print(f"[OK] 报告已生成:{out_path}")

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The description is written entirely in Chinese and specifies the skill behavior in that locale without any indication that users may choose another language. Under the policy, forcing a specific language without user opt-in is a natural-language policy violation unless the locale restriction is documented and justified.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
91% confidence
Finding

The file is entirely written as a mandatory reference table in Chinese and uses prescriptive language such as '必须优先查阅本表' and domain-specific Chinese terminology without any indication that another language is supported. Under the policy, forcing a specific language without user opt-in is a natural-language policy concern unless the locale restriction is explicitly documented and justified.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.