Back to skill

Security audit

patent-family-analyzer

Security checks for vulnerabilities and agentic risk

Overview

The skill is purpose-aligned for patent report generation, but its HTML report generator can execute injected browser code from untrusted patent or JSON data.

Review before installing or using with untrusted patent data. The skill requires PatSnap/Eureka MCP account authorization and creates local HTML reports. Do not open reports generated from untrusted or attacker-controlled JSON/MCP data until the renderer escapes HTML, validates URLs, avoids innerHTML and inline event handlers, and safely embeds JSON.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/analyze_family.py:337
Finding

Stored HTML and JavaScript Injection in Generated Patent Reports

Content
View full analysis
{{ const fill = n.depth === 0 ? "#1a73e8" : (n.depth === 1 ? "#34a853" : "#fbbc04"); const textColor = n.depth <= 1 ? "white" : "#202124"; svgContent += ` ${{n.label}} ${{n.country || ""}} ${{n.date || ""}} `; }}); svg.innerHTML = svgContent; ``` Other input fields are directly interpolated into HTML attributes and element bodies: ```python items.append(f'''

🔗 {p.get("pn","")}

``` The attacker-controlled data source is loaded without validation: ```python with open(args.data_json, "r", encoding="utf-8") as f: data = json.load(f) html = build_html(data) ``` ### Technical Analysis The report generator treats values loaded from the input JSON as trusted markup. P ...[truncated 2698 chars]
Remediation
View remediation
... ``` Before embedding, encode `<` as `\u003c` so that `` cannot terminate the element. Parse the content using `JSON.parse(document.getElementById("tree-data").textContent)`. 6. Validate the loaded JSON against a strict schema: - Enforce expected scalar types. - Limit string lengths and collection sizes. - Reject unexpected properties where practical. - Validate dates, patent numbers, legal-status values, and URLs. 7. Add a restrictive Content Security Policy to the generated report. Prefer external or nonce-based scripts and prohibit inline script execution. This should be defense in depth rather than a replacement for output encoding. 8. Add automated tests using payloads in every rendered field, including HTML tags, quote characters, template-literal delimiters, ``, SVG event handlers, and dangerous URL schemes. ]]>
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (7)

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

声明描述的是一个端到端“同族专利深度分析工具”,核心能力包括:以专利号为输入、自动检索同族、抓取全文、分析技术内容并生成报告。但代码的实际作用是一个HTML报告生成器/渲染器:命令行参数要求的是 data_json 文件,而非专利号;代码只读取本地JSON并调用一组 make_* 函数拼接HTML;树状图也是基于传入的 tree 数据前端渲染;不存在任何网络请求、专利数据库访问、全文抓取、AI调用或同族检索逻辑。虽然文件头注释提到这些步骤会由 Eureka Agent / MCP 完成,但依据所给代码块本身,实际行为明显小于且不同于声明的完整自动分析能力,因此属于描述与代码行为不符。

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
83% confidence
Finding

The skill declares behavior that includes reading inputs and writing an HTML report, but it does not explicitly constrain tool scope via permissions or allowed-tools. In an agent environment, missing scope declarations can allow broader-than-necessary file access or create ambiguity about what the skill is permitted to read/write, increasing the chance of unintended filesystem interaction.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The manifest description is entirely in Chinese and presents the skill behavior and outputs in that language, while the file does not state that the skill is region-specific or provide an opt-in choice of output language. Under the policy, forcing a specific language without user choice is a natural-language locale violation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

This code file contains user-facing natural language in the module docstring and report content that is exclusively Chinese. Under the stated policy, forcing a specific language without user opt-in is a locale-policy violation unless the tool clearly offers a choice or documents a justified region-specific scope, which this file does not.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

模块说明在 L008-L014 将脚本描述为会执行数据检索和全文获取流程,但实际实现中 main 仅在 L530-L538 读取已有 JSON 文件并写出 HTML,完全没有任何 MCP 调用、检索或分析执行逻辑。这会让人误判该脚本具备自动采集与分析能力,而它实际上只是报告渲染器。

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
87% confidence
Finding

The skill specifies creating an HTML report in session storage without explicitly warning the user that a file will be written. While the write target appears limited to session storage, silent file creation can still surprise users, create persistence they did not expect, and become riskier if report content includes untrusted data rendered in HTML/JS.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
99% confidence
Finding

文档在 L005-L006 写明用法为传入 patent_number,但 argparse 在 L525-L527 实际要求的位置参数是 data_json 文件路径。这里不是简单省略细节,而是对输入类型和执行方式作了直接相反的说明。

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.