Back to skill

Security audit

云启智联AI服务

Security checks for vulnerabilities and agentic risk

Overview

This skill appears to be a real financial document OCR client, but it sends sensitive documents to an external service and automatically saves unsafe local previews containing financial data.

Review before installing. Only use this skill with financial files you are allowed to send to Yunqi Zhilian's service, avoid casual auto-use from keyword mentions, do not pass API keys on the command line in shared systems, and treat generated HTML/JSON previews as sensitive records. Avoid opening previews from untrusted documents until the HTML escaping issues are fixed.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (4)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/api_client.py:843
Finding

Stored script injection in receipt preview

Content
View full analysis
${f.label}${val}`; document.getElementById('receiptDataPanel').innerHTML = html; ``` ### Technical Analysis The receipt preview embeds API-controlled OCR results directly inside a `` terminates the surrounding script element regardless of whether that sequence occurs inside a JavaScript string. An attacker-controlled receipt field could therefore contain a value such as: ```html ``` When the generated preview is opened, the browser interprets the injected element as executable script. There are also secondary DOM injection sinks because OCR values are interpolated into HTML strings and assigned to `innerHTML` without applying the available `escapeHtml()` function. The sour ...[truncated 1760 chars]
Remediation
View remediation
` element. Escape at least `<`, `>`, `&`, U+2028, and U+2029 before embedding: ```python def safe_script_json(value): return ( json.dumps(value, ensure_ascii=False) .replace("&", "\\u0026") .replace("<", "\\u003c") .replace(">", "\\u003e") .replace("\u2028", "\\u2028") .replace("\u2029", "\\u2029") ) ``` 2. Prefer storing serialized data in a non-executable ``, ``, quotes, backticks, and malformed URLs. ]]>

T09 · Insecure Skill Coding Practices

Error
Location
scripts/api_client.py:978
Finding

Stored HTML injection in statement preview

Content
View full analysis
' + "".join( '
{}
{}
'.format(k, v) for k, v in global_items ) + '' ``` Transaction fields are also inserted directly into table markup: ```python rows_html += """ {date} {cp_name}{cp_acc_html} {cp_acc} {credit} {debit} {balance} {abstract} {trans_no} """.format( date=item.get("tradeDate", ""), cp_name=cp_name, cp_acc_html='
' + cp_acc + '' if cp_acc else "", cp_acc=cp_acc, credit=fmt(credit) if credit else "-", debit=fmt(debit) if debit else "-", balance=fmt(balance) if balance is not None else "-", abstract=item.get("abstract", ""), trans_no=item.get("transNo", ""), ) ``` The generated fragment is then inserted into the final document: ```python html = html_template.replace("{page_count}", str(page_count)) \ .replace("{record_count}", str(len(records))) \ .replace("{global_items_html}", global_items_html) \ .replace("{total_credit}", fmt(total_credit)) \ .replace("{total_debit}", fmt(total_debit)) \ .replace("{net_change_style}", net_change_style) \ .replace("{net_change}", fmt(net_change)) \ .replace("{rows_html}", rows_html) ``` ### Technical Analysis Account names, bank names, account numbers, dates, counterparties, transaction descriptions, and transaction numbers originate from an external OCR response and are concatenated directly into HTML. No HTML escaping is performed. An atta ...[truncated 1467 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
scripts/voucher_generator.py:1237
Finding

Stored HTML injection in voucher preview

Content
View full analysis
{code} {name} {debit} {credit} {remark} """.format( code=e.account_code, name=e.account_name, d_cls=d_cls, c_cls=c_cls, debit=_fmt_money(e.debit) if e.debit > 0 else "-", credit=_fmt_money(e.credit) if e.credit > 0 else "-", remark=e.remark or "", ) ``` Review details and voucher descriptions are also inserted without escaping: ```python flags_html += '
{icon} {text}
'.format( color=flag_color, icon=flag_icon, text=flag.get("suggestion", "") ) ``` ```python return """
Voucher {index}   {badge}
Date: {date} Source: {source} Confidence: {conf} Status: {status}
Description: {description}
{entries_html} {review_html}
""".format( index=index + 1, badge=badge, date=v.date, source=v.source_type, conf=conf_text, status=v.status.value, description=v.description, entries_html=entries_html, review_html=review_html, ) ``` ### Technical Analysis Voucher data is derived from OCR results or from a user-supplied parsing-result JSON file. Fields such as descriptions, remarks, account labels, review summaries, alternatives, and suggestions are inserted directly into HTML without contextual encoding. A malicious source field propagated into `v.descriptio ...[truncated 1447 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/config_manager.py:33
Finding

API key encryption uses a predictable locally reproducible key

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Taint TrackingDirect Taint Flow, Variable-Mediated Taint Flow, Credential Exfiltration Chain
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (24)

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

This code does not implement the declared end-user functionality of the skill. Instead of calling document analysis APIs or processing files, it only manages API key storage and retrieval via environment variables and an encrypted local file. While API key management could be a supporting utility for such a service, this chunk's actual primary purpose is configuration/credential handling, which is materially different from the declared OCR parsing and voucher-generation service behavior. Therefore, the description does not accurately represent what this supplied code chunk actually does.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

声明描述的是一个完整的AI文档解析服务/技能,重点在OCR识别、解析接口、任务查询与凭证生成的全链路能力;而给定代码仅覆盖其中“根据已有解析结果生成记账凭证”这一后处理环节。代码没有网络请求、没有调用外部OCR服务、没有文件上传解析、没有任务ID查询,也没有ping接口,因此其主要行为明显窄于且不同于声明的核心能力。虽然声明中提到支持智能记账凭证生成,与代码部分一致,但整体描述会让人认为该技能可直接完成文档识别与相关接口调用,而实际代码并不具备这些能力,因此属于实质性描述与行为不匹配。

Content

No source excerpt is available for this finding.

Missing User Warnings

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The skill encourages uploading financial documents and submitting file paths or URLs to a remote OCR service, but it does not prominently warn that those documents and referenced URLs will leave the local environment for third-party processing. Because the content involves bank receipts, statements, invoices, and accounting artifacts, the missing disclosure materially increases privacy, confidentiality, and compliance risk.

Content

No source excerpt is available for this finding.

Tainted flow: 'files' from open (line 257, file read) → requests.post (network output)

High
Category
Data Flow
Confidence
80% confidence
Finding

File contents flow to a network sink. This may indicate data exfiltration of sensitive files.

Content

Scanner excerpt · scripts/api_client.py (reported line 314)May include surrounding context.

python
resp = requests.get(url, headers=headers, timeout=30)
        else:
            if files:
                resp = requests.post(url, data=data, files=files, headers=headers, timeout=120)
            elif send_json:
                resp = requests.post(url, json=data, headers=headers, timeout=30)
            else:

Tainted flow: 'timeout' from open (line 381, file read) → urllib.request.urlopen (network output)

High
Category
Data Flow
Confidence
80% confidence
Finding

File contents flow to a network sink. This may indicate data exfiltration of sensitive files.

Content

Scanner excerpt · scripts/api_client.py (reported line 382)May include surrounding context.

python
try:
        timeout = 120 if (method != "GET" and files) else 30
        with urllib.request.urlopen(req, timeout=timeout) as resp:
            raw = resp.read()
            status = getattr(resp, "status", resp.getcode())
        return _json.loads(raw.decode("utf-8"))

Tainted flow: 'timeout' from open (line 381, file read) → urllib.request.urlopen (network output)

High
Category
Data Flow
Confidence
80% confidence
Finding

File contents flow to a network sink. This may indicate data exfiltration of sensitive files.

Content

Scanner excerpt · scripts/api_client.py (reported line 464)May include surrounding context.

python
try:
        timeout = 120 if (method != "GET" and files) else 30
        with urllib.request.urlopen(req, timeout=timeout) as resp:
            raw = resp.read()
            status = getattr(resp, "status", resp.getcode())
        return _json.loads(raw.decode("utf-8"))

Tainted flow: 'timeout' from open (line 381, file read) → requests.get (network output)

High
Category
Data Flow
Confidence
80% confidence
Finding

File contents flow to a network sink. This may indicate data exfiltration of sensitive files.

Content

Scanner excerpt · scripts/api_client.py (reported line 459)May include surrounding context.

python
def _http_get(url, timeout=15):
    """GET 请求返回 (content_bytes, content_type),requests 缺失时回退 urllib"""
    if REQUESTS_AVAILABLE and requests is not None:
        resp = requests.get(url, timeout=timeout, stream=False)
        resp.raise_for_status()
        return resp.content, resp.headers.get("Content-Type", "image/jpeg")
    import urllib.request

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
87% confidence
Finding

The skill instructs the agent to use shell, local file reads/writes, environment-backed credential handling, and network calls, but it does not declare any explicit tool scope or permission boundary. That makes the effective privilege surface larger than users and hosts can easily audit, increasing the chance of unintended command execution, file creation, or remote data transmission during normal use.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The top-level description says the skill should auto-call interfaces whenever common finance/accounting keywords appear, including broad terms like '文件解析', 'ping', '做账', and '体验馆'. Overbroad triggers can cause the agent to invoke shell/network actions during ordinary conversation without sufficiently clear user intent, leading to unintended remote requests or local side effects.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The trigger section maps simple keyword presence directly to API or command invocation, with no scope limits, confirmation step, or disambiguation rules. In an agent setting, this increases the risk of accidental execution, unintended file handling, or external transmission of user-provided documents based on ambiguous conversational context.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The HTML preview behavior can expose remote image URLs in generated local files and may write preview artifacts to fallback paths including system temporary directories, but this risk is not clearly surfaced to users. That can leak document-related metadata, create unexpected persistence on disk, and surprise users in sensitive environments handling financial records.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
80% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · scripts/api_client.py (reported line 316)May include surrounding context.

python
if files:
                resp = requests.post(url, data=data, files=files, headers=headers, timeout=120)
            elif send_json:
                resp = requests.post(url, json=data, headers=headers, timeout=30)
            else:
                resp = requests.post(url, data=data, headers=headers, timeout=30)

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The helper fetches arbitrary image_url values and can download attacker-controlled remote URLs into the local environment. If parsing results or upstream service output are compromised, this enables SSRF-style access to internal resources, metadata endpoints, or local-network hosts from the runtime generating previews.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The code writes bank receipt parsing results into local HTML files, which may contain sensitive financial data such as account numbers, counterparties, transaction amounts, and images. Persisting this data without an explicit warning or opt-in increases the risk of unintended disclosure to other local users, backups, sync tools, or forensic recovery.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The same persistence issue exists for bank statement previews, which can include highly sensitive account and transaction histories. Automatic local storage broadens exposure beyond the original API response and may violate least-retention expectations for financial records.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

Generated voucher JSON is saved locally and may contain detailed accounting records derived from financial documents. This creates a durable sensitive-data artifact without an explicit warning at the save point, increasing the chance of accidental disclosure through shared disks, endpoint monitoring, or backup systems.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The CLI accepts the API key as a positional argument, which exposes the secret to shell history, process listings, and audit logs on multi-user systems or managed environments. Even though the script later encrypts the key at rest, the exposure happens before encryption and can allow credential theft by local users, monitoring tools, or support/logging infrastructure.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
84% confidence
Finding

The module description and all user-facing strings are fixed in Chinese, including the stated accounting standard and generated output text, with no indication that language is configurable or intentionally restricted to a China-specific deployment context. Under the language/locale policy, forcing a specific language without user opt-in can be a natural-language policy violation unless the locale limitation is clearly documented and justified.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The HTML preview feature writes parsed accounting and voucher data to local disk, including potentially sensitive financial information such as counterparties, amounts, remarks, and invoice details. In an agent/service context, creating local files expands data persistence beyond the core parsing/generation function and can expose sensitive records through predictable directories, shared workspaces, backups, or later local access.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

When HTML generation is enabled, the code automatically chooses a writable directory and stores a human-readable file containing financial voucher contents without a prominent warning about where sensitive data will persist. This can lead to unintended disclosure because users may not realize bank/accounting data has been saved under home, current working, or temp directories that other processes or users may access.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
83% confidence
Finding

L181-L205 明确将系统临时目录中的文件创建描述为“严禁”行为,并强调这是必须严格遵守的规则;但 L309-L310 又说明 HTML 预览文件在默认目录不可写时会自动回退到系统临时目录。虽然前者主要针对 .py 桥接脚本、后者是 HTML 预览文件,但当前文档表述容易让执行约束与实际行为产生冲突。

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
80% confidence
Finding

The module docstring presents the client description entirely in Chinese, and the rest of the user-facing strings in the file also assume Chinese output. For a general-purpose skill, forcing a specific language without user opt-in can be a natural-language policy issue.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Low
Category
Not specified by scanner
Confidence
83% confidence
Finding

The manifest emphasizes remote OCR parsing, async query, and voucher generation APIs, but this client also creates local output directories and persists HTML preview pages and voucher JSON files on disk. Local artifact generation may be useful, yet it is a material behavior not reflected in the manifest's service description.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Low
Category
Not specified by scanner
Confidence
77% confidence
Finding

The manifest frames this skill as an AI service that processes parsing results and returns voucher outputs when relevant intents are mentioned. The CLI introduces a generic local-file processing mode that reads arbitrary JSON paths from disk, which is not part of the stated service capability and is unnecessary for the core intent of voucher generation from already supplied data.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.