Back to skill

Security audit

unisound-missing-tumor-name

Security checks for vulnerabilities and agentic risk

Overview

This skill performs the advertised medical-record quality check, but it can send sensitive records and the API key to a configurable external endpoint and has inconsistent privacy claims about local persistence.

Review before installing. Use this only with de-identified records and an approved medical-LLM endpoint. Do not pass an untrusted --base value, because the current code will send the appkey and record content to that destination. Treat output files and any --save-prepared files as sensitive medical data and store or delete them according to your privacy policy.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/emr_qc_impl.py:53
Finding

Configurable API Endpoint Can Expose the AppKey and Medical Records

Content
View full analysis
str: payload = {"model": model, "messages": messages, "temperature": 0} resp = _http_post(url, payload, headers, timeout=timeout) ``` The destination is supplied directly through the command line: ```python parser.add_argument( "--base", default=DEFAULT_LLM_BASE, help=f"大模型 base URL(默认:{DEFAULT_LLM_BASE})。", ) ``` ### Technical Analysis The `--base` argument is accepted without validating the URL scheme, hostname, port, or resolved IP address. The resulting URL receives: - The supplied AppKey in the `Authorization: Bearer` header. - Medical-record fields embedded in the request messages. - Any other prompt data generated by the quality-control workflow. Consequently, an invocation wrapper, configuration injection, or operator trick can change `--base` to an attacker-controlled server. The implementation does not require HTTPS or restrict the destination to the documented HiVoice MaaS service. The same functionality can also be abused as a limited server-side request primitive when the program runs in an environment that can reach internal services. Because requests use a fixed POST path and JSON body, this is not unrestricted SSRF, but it can still disclose connectivity information or transmit credentials and medical information to internal or external destinations. ### Attack Path ...[truncated 1224 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/emr_qc_impl.py:117
Finding

Medical-Record Content Can Inject Instructions into the LLM Quality-Control Workflow

Content
View full analysis
标记,请参考 现病史:未用药,否认口干、多饮、多尿、体重下降 既往史:乳腺恶性肿瘤术后,现靶向治疗。否认胰腺炎。 有 现病史:甲状腺乳头状癌术后2年。术后病理:甲状腺乳头状癌,经典型及滤泡亚型。现优甲乐100ug,否认心慌、手抖、多汗、体重下降。 既往史:2014诊断甲亢,赛治治疗,现已停药。 有 现病史:少痰,不烧复查CT,伴关节疼痛 既往史:肺部结节病史,过敏史:无。传染病史:无。 无 现在请判断下面的门诊病历 现病史:{hpi} 既往史:{pmh}""" )]) if "无" in has_tumor and "有" not in has_tumor: return "无缺陷" ``` The final stage uses the same pattern and accepts the resulting text without structural validation: ```python return llm([user_msg( f"""你是一位病历质控专家,请对门诊病历进行质控,如果没有缺陷请直接回答"无缺陷",不要进行任何分析;如果有缺陷请先回答"有缺陷"再另起一行分析原因 缺陷描述: 1.既往史或现病史中记录患者有肿瘤或者癌症病史,没有记录具体的肿瘤或者癌症名称 2.任何带有部位的肿瘤或者癌症,都认为是有肿瘤名称,请直接回答"无缺陷",比如"甲状腺Ca"、"肺癌"、"肺部肿瘤"、"结肠癌"、"淋巴瘤"、"脑膜瘤"、"脑淋巴瘤"、"乳腺癌"、"直肠癌"、"直肠肿瘤"等等 以下是一些示例,用标记,请参考 【病历】 现病史:体重下降 既往史:结肠癌术后,磺胺药敏史 【质控结果】 无缺陷 【病历】 现病史:未用药,否认口干、多饮、多尿、体重下降 既往史:乳腺恶性肿瘤术后,现靶向治疗。否认胰腺炎。 【质控结果】 无缺陷 【病历】 现病史:不烧,自觉胸背部不适,偶有咳痰。 既往史:肾功能不全,肿瘤病史,过敏史:无。传染病史:无。 【质控结果】 有缺陷 既往史中记录"肿瘤病史",但是没有记录具体的肿瘤名称 现在请对下面的门诊病历进行质控 【病历】 现病史:{hpi} 既往史:{pmh} 【质控结果】""" )]) ``` ### Technical Analysis The `hpi` and `pmh` values originate from input medical records and are therefore untrusted. They are interpolated into the same natural-language message that defines the model's task. There is no reliable trust boundary separating task instructions from record data. A malicious record can contain model-directed text instructing the model to ignore the preceding rule, emit a chosen classification, or g ...[truncated 1933 chars]
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (13)

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
96% confidence
Finding

The skill explicitly describes capabilities to read local files, write output files, and send medical record content to a remote API, but it does not declare any tool/permission scope. In a system that relies on manifest-declared permissions, this creates a governance gap where sensitive EMR data may be processed or exfiltrated without explicit least-privilege constraints or operator visibility.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
78% confidence
Finding

The markdown specifies that input must be outpatient medical records using Chinese field labels such as “主诉”, “现病史”, and related tags, and does not mention any user opt-in or alternative locale support. Under the policy, enforcing a specific language/locale in natural-language instructions without choice or explicit justification is a reportable issue.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

User-facing natural-language strings in the module docstring and CLI argument descriptions are entirely in Chinese, with no indication that another language is supported or that the user can choose a locale. Per the policy, forcing a specific language without opt-in is a natural-language policy violation.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The skill sends patient medical record content to an external LLM API endpoint for classification without any explicit disclosure or consent mechanism. In a medical context, this is especially sensitive because it can expose protected health information to a third-party service, creating privacy, compliance, and data-governance risk.

Content

No source excerpt is available for this finding.

Ssd 1

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The first-stage prompt interpolates untrusted patient record text directly into a single user message with instructions and examples. If the record contains adversarial text such as instruction-like phrases, the external model may follow those embedded directions instead of the intended classification task, producing manipulated QC outcomes or causing unintended data handling.

Content

No source excerpt is available for this finding.

Ssd 1

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The second-stage QC prompt mixes policy, examples, and raw medical record text in one prompt, making it vulnerable to semantic prompt injection and behavior redirection. Because the model's output determines defect findings in a medical QA workflow, manipulated record text could suppress defects, fabricate reasons, or degrade reliability of compliance-related review.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The skill contains a local file-read helper with fallback path traversal through parent directories to a data folder, which exceeds the stated text-only QC behavior. This broadens the attack surface by allowing local file access patterns in an agent skill that should only process supplied text, and may enable unintended access to sensitive records if the path is user-influenced.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill writes QC output to a local file even though its stated purpose is to analyze provided record text and return a result. Because the QC result is derived from patient record content, this creates an unnecessary persistence channel for potentially sensitive medical data and increases exposure on shared hosts or agent environments.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

When --save-prepared is used, the script writes preprocessed outpatient medical record text directly to disk, which can include highly sensitive personal and health information. If the output path is shared, insecurely permissioned, or retained longer than intended, this creates a real confidentiality and compliance risk even though it is framed as a debugging feature.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The skill sends full patient record text to an LLM service using user-supplied base/model parameters, but there is no in-code guardrail, minimization, consent check, or disclosure about transmitting regulated medical data. In a medical-record QC context, this is especially sensitive because the data likely contains PHI, and misconfiguration or use of an unintended endpoint could expose records outside the intended trust boundary.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
87% confidence
Finding

Natural-language instructions, prompts, outputs, and status messages in this file are all fixed to Chinese. Under the stated policy, forcing a specific language without user opt-in is a language/locale policy violation unless the constraint is explicitly documented and justified.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
79% confidence
Finding

The function accepts a bearer-token style appkey and uses it in an Authorization header for outbound requests. The code lacks nearby comments or user-facing disclosure explaining that credentials are being used to access an external service.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
93% confidence
Finding

The module docstring and CLI help text are entirely in Chinese, which effectively forces a specific language for users. The file does not offer a language/locale option or state that the tool is intentionally limited to a Chinese-speaking context.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.