Back to skill

Security audit

Museum Guide

Security checks for vulnerabilities and agentic risk

Overview

This skill is a coherent museum itinerary helper that uses local reference data and configured search/LLM services, with no evidence of hidden persistence, destructive behavior, or credential theft.

Install only if you are comfortable configuring an LLM API key and having museum requests, preferences, and generated itinerary context processed by that provider. Prefer offline reference data for supported museums, and verify hours, tickets, and important exhibit claims with official museum sources because online search results and LLM-generated details can be wrong or manipulated.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/search_artifacts.py:55
Finding

Indirect Prompt Injection Through Untrusted Search Results

Content
View full analysis

Vulnerability Details

File Location: scripts/search_artifacts.py, lines 55–91
Vulnerability Type: Indirect prompt injection caused by untrusted search content being embedded directly into an LLM prompt
Risk Level: Medium

Vulnerable Code

python
combined_content = "\n\n".join([
    f"--- 搜索结果 {i+1} ---\n{result[:1500]}"
    for i, result in enumerate(search_results[:12])
])

prompt = f"""
你是一位博物馆文物专家。请从以下{museum_name}的搜索结果中提取文物信息。

要求:
1. 只提取该博物馆的**常设馆藏文物**,排除借展、巡展、复制品、仿品
2. 提取文物名称、所属展馆、所属时期、文物种类、是否是镇馆之宝、简要描述
2. 时期必须从以下列表中选择:远古时期、夏商西周、春秋战国、秦汉、三国两晋南北朝、隋唐五代、辽宋夏金元、明清
3. 文物种类必须从以下列表中选择:青铜器、陶器、瓷器、漆器、玉器宝石、石器石刻、书画古籍、服饰、砖瓦、钱币、化石、金银器、其他
4. 关注的领域从以下列表中选择:农耕、狩猎、饮食、建筑、人物、武器、文房四宝、牌章证件、货币、书法、绘画、雕像、服装、饰品、仪器、佛教、乐器、纹饰、花瓶、礼制、古生物、新石器、旧石器、陈设品、科技、其他
5. 如果搜索结果中未提及展馆,请使用"待确认"
6. 尽可能多地提取文物(至少20件)

请返回JSON数组格式:
[
    {{
        "name": "文物名称",
        "hall": "展馆名称",
        "period": "时期",
        "type": "文物种类",
        "is_treasure": true/false,
        "description": "简要描述",
        "domains": ["领域1", "领域2"],
        "child_friendly": true/false
    }}
]

搜索结果:
{combined_content}
"""

try:
    result = call_llm_api(prompt)

Technical Analysis

The application obtains passages from ProSearch and concatenates them into combined_content. These passages are then interpolated directly into an instruction-bearing prompt and submitted to the configured LLM.

Search results are attacker-influenceable data. A malicious or compromised indexed page can contain text framed as model instructions, such as directions to disregard the extraction rules, fabricate artifacts, return promotional content, or place attacker-controlled text in output fields. The prompt does not clearly isolate search passages as untrusted data, explicitly prohibit following instructions found within them, or identify their provenance.

The response is checked only to determine whether it is a list. Indivi ...[truncated 2027 chars]

Remediation
View remediation

Remediation Suggestions

  1. Establish an explicit trust boundary

    • Mark every search passage as untrusted quoted data.
    • Instruct the model that text inside search-result delimiters is evidence only and that any instructions contained there must be ignored.
    • Use clear per-document delimiters and preserve source metadata separately.
  2. Use structured model inputs where supported

    • Place application instructions in a system or developer message.
    • Place search passages in a separate user message or structured field rather than concatenating instructions and retrieved text into one prompt.
  3. Validate every returned record

    • Require an exact JSON schema.
    • Reject unknown properties and incorrect data types.
    • Enforce enumerations for period, type, and domains.
    • Set limits for field length and total record count.
    • Reject records containing instruction-like language, unexpected Markdown, scripts, or URLs where those values are not required.
  4. Verify provenance

    • Prefer official museum websites and other allowlisted sources.
    • Require source citations for extracted artifacts.
    • Cross-check important claims against multiple independent sources before marking an artifact as a museum treasure or permanent exhibit.
  5. Reduce attacker-controlled context

    • Deduplicate results and remove irrelevant content.
    • Strip page navigation, advertisements, and obvious prompt-injection patterns before invoking the model.
    • Do not treat filtering alone as the primary defense, because malicious instructions can be phrased in many ways.
  6. Fail safely

    • If validation or provenance checks fail, discard the generated records rather than rendering them.
    • Clearly label unverified online information and direct users to official museum sources.
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Taint TrackingDirect Taint Flow, Variable-Mediated Taint Flow, Credential Exfiltration Chain
Findings (21)

Tainted flow: 'url' from os.environ.get (line 39, credential/environment) → requests.post (network output)

Critical
Category
Data Flow
Confidence
90% confidence
Finding

Credentials or environment variables flow to a network sink. This is a high-confidence indicator of credential exfiltration.

Content

Scanner excerpt · scripts/search_artifacts.py (reported line 45)May include surrounding context.

python
headers = {"Content-Type": "application/json"}
    
    try:
        response = requests.post(url, json=payload, headers=headers, timeout=30)
        response.raise_for_status()
        return response.json()
    except Exception as e:

Hidden Instructions

High
Category
Prompt Injection
Confidence
60% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · references/上海博物馆.csv (reported line 1)May include surrounding context.

text
文物名称,展厅,时期,参观顺序,是否镇馆之宝,文物种类,所属领域,人员类别,,,,,,,,,
金累丝龙纹嵌珍珠宝石帽顶,珠光宝气,明清,1,f,金银器,"['饰品', '礼制']",0,,,,,,,,,
点翠嵌珍珠宝石金龙凤冠,珠光宝气,明清,2,f,金银器,"['饰品', '礼制']",0,,,,,,,,,
貂皮嵌珠皇后冬朝冠,珠光宝气,明清,3,f,服饰,"['服装', '饰品', '礼制']",0,,,,,,,,,

Hidden Instructions

High
Category
Prompt Injection
Confidence
60% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · references/中国国家博物馆.csv (reported line 1)May include surrounding context.

text
文物名称,展厅,时期,参观顺序,是否镇馆之宝,文物种类,所属领域,人员类别,,,,,,,,,
金累丝龙纹嵌珍珠宝石帽顶,珠光宝气,明清,1,f,金银器,"['饰品', '礼制']",0,,,,,,,,,
点翠嵌珍珠宝石金龙凤冠,珠光宝气,明清,2,f,金银器,"['饰品', '礼制']",0,,,,,,,,,
貂皮嵌珠皇后冬朝冠,珠光宝气,明清,3,f,服饰,"['服装', '饰品', '礼制']",0,,,,,,,,,

Hidden Instructions

High
Category
Prompt Injection
Confidence
60% confidence
Finding

Hidden instructions were detected in comments or invisible text. These could contain malicious directives. Manual review is recommended.

Content

Scanner excerpt · references/故宫珍宝馆.csv (reported line 1)May include surrounding context.

text
文物名称,展厅,时期,参观顺序,是否镇馆之宝,文物种类,所属领域,人员类别,,,,,,,,,
金累丝龙纹嵌珍珠宝石帽顶,珠光宝气,明清,1,f,金银器,"['饰品', '礼制']",0,,,,,,,,,
点翠嵌珍珠宝石金龙凤冠,珠光宝气,明清,2,f,金银器,"['饰品', '礼制']",0,,,,,,,,,
貂皮嵌珠皇后冬朝冠,珠光宝气,明清,3,f,服饰,"['服装', '饰品', '礼制']",0,,,,,,,,,

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
93% confidence
Finding

The skill instructs the agent to run local scripts, use environment-supplied configuration, read local reference files, write profile JSON files, and optionally access networked search, but it declares no explicit tool scope or allowlist. This creates an authorization gap where the runtime may permit broader-than-intended file, env, or network access, increasing the blast radius if the skill is invoked or modified maliciously.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The trigger list includes broad generic terms like '博物馆', 'museum', and '路线规划', which can cause the skill to activate for loosely related conversations rather than explicit user intent. Overbroad invocation is risky here because the skill can lead the agent into running scripts and using file/network capabilities unnecessarily, expanding exposure from an otherwise simple chat interaction.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

Natural-language strings throughout the file, including the module description and prompt template, require Chinese interaction and output formats with no indication that users may choose another language. This can violate language/locale policy where skills must not force a specific language without user opt-in.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
97% confidence
Finding

The function performs an outbound HTTP request to a configurable external endpoint and includes the full prompt payload, which contains user-supplied content and possibly prior profile data. In this skill context, the transmitted content can include personal preference data and whether the user is traveling with children, so unbounded external transmission increases privacy exposure and trust risk.

Content

Scanner excerpt · scripts/extract_profile.py (reported line 156)May include surrounding context.

python
}

    try:
        response = requests.post(config['api_base'], headers=headers, json=data, timeout=120)
        response.raise_for_status()
        result = response.json()

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The code sends raw user input to an external LLM service for processing without any visible notice, consent flow, or data-minimization at the call site. Because user prompts may include personal preferences, travel plans, or information about children, this creates a real privacy and data-governance risk if users do not expect third-party transmission.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

This code file contains natural-language instructions and output text exclusively in Chinese, including the main module documentation and runtime messages. Under the language/locale policy, forcing a specific language without user opt-in or clear justification is a policy concern.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

This Python file contains its primary docstring usage instructions entirely in Chinese, including the invocation example and phase descriptions. Because the skill does not offer an explicit language selection or indicate that it is limited to Chinese-speaking users, it creates a locale-policy issue by forcing one language implicitly.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The function sends artifact metadata and user-profile-derived museum preferences to an external LLM service without any visible user notice, consent, or minimization in this file. Even if the data is not highly sensitive in every case, profiles can reveal interests, family status (with_children), and visit preferences, creating unnecessary third-party data exposure.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The embedded prompt requires the model to return recommendation reasons in Chinese and does not provide any opt-in or alternative locale handling. This is a language-policy issue because the skill imposes a specific language on generated user-facing content without documenting a user choice or justified locale constraint.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

Manifest描述聚焦于根据参观时长、兴趣偏好和是否带儿童来生成必看文物清单、镇馆之宝路线与参观顺序。这里代码另外通过LLM获取并输出位置、开放时间、门票、交通、馆内设施和附近景点信息,属于旅行/场馆资讯生成,而非路线排序本身的直接实现细节。

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

一个博物馆参观路线规划助手,合理能力是基于文物数据进行筛选、排序和生成推荐理由。这里额外构造提示词让LLM生成场馆位置、门票、交通和周边设施等综合信息,形成了与核心规划逻辑分离的通用信息检索/生成能力,且manifest未说明该扩展能力。

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The prompt instructs the model to return user-facing museum information using Chinese headings and JSON keys only. There is no indication that the user can choose another language or that the tool is explicitly restricted to a Chinese-language context.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

This script performs online search and then forwards aggregated search-result content to an LLM without clearly documenting that network transmission occurs. In this skill, the query is usually just a museum name, so sensitivity is limited, but the lack of disclosure and explicit consent still creates a privacy and transparency risk for users and operators.

Content

No source excerpt is available for this finding.

External Transmission

Medium
Category
Data Exfiltration
Confidence
80% confidence
Finding

Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.

Content

Scanner excerpt · scripts/search_artifacts.py (reported line 45)May include surrounding context.

python
headers = {"Content-Type": "application/json"}
    
    try:
        response = requests.post(url, json=payload, headers=headers, timeout=30)
        response.raise_for_status()
        return response.json()
    except Exception as e:

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
80% confidence
Finding

The description is written as a Chinese-only assistant description while also advertising English trigger phrases like "museum itinerary," but the document does not state that users may choose their preferred language. This can be interpreted as a locale/language policy issue because the skill appears to assume a specific language without opt-in or explicit multilingual support guidance.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

The museum lookup sends the selected museum name to an external LLM without explicit disclosure in this file. While a museum name is generally low sensitivity, silent third-party transmission still creates a privacy and transparency issue, especially when combined with other user context elsewhere in the workflow.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
93% confidence
Finding

The natural-language instructions, usage text, and embedded LLM prompt are entirely Chinese and constrain outputs to Chinese labels and categories, but the file does not state that the skill is intentionally China/Chinese-only or allow user language selection. That creates a locale-policy concern because the skill implicitly forces one language without opt-in.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.