Back to skill

Security audit

unisound-pe-missing-negative

Security checks for vulnerabilities and agentic risk

Overview

The skill appears purpose-built for medical record quality control, but it handles sensitive clinical text with an externally configurable LLM endpoint and some under-disclosed persistence risk.

Review before installing in any real clinical workflow. Use only de-identified records, pin the LLM base URL to an approved HTTPS service, protect the app key, avoid --save-prepared unless you intentionally want a plaintext copy of the record, and treat outputs as clinician-review-required rather than authoritative.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/emr_qc_impl.py:57
Finding

Caller-Controlled LLM Endpoint Can Expose API Credentials and Medical Records

Content
View full analysis
str: payload = {"model": model, "messages": messages, "temperature": 0} resp = _http_post(url, payload, headers, timeout=timeout) ``` ### Technical Analysis The command-line `--base` argument is accepted without validating its scheme, hostname, port, resolved address, or relationship to the intended HiVoice service. The resulting URL receives an HTTP request containing: - The API credential in the `Authorization: Bearer ...` header. - The model request payload, including EMR-derived physical-examination and diagnosis data. Consequently, anyone able to influence invocation arguments can redirect the request to an attacker-controlled service. Permitting arbitrary destinations can also enable server-side requests to internal or link-local services from environments where the Skill has network access. No restriction guarantees HTTPS or prevents redirects to an untrusted destination. A non-HTTPS endpoint could additionally expose credentials and medical information to network interception. ### Attac ...[truncated 1387 chars]
Remediation
View remediation

T01 · Skill Instruction Hijacking

Warning
Location
scripts/emr_qc_impl.py:113
Finding

EMR Fields Can Inject Instructions into the LLM Quality-Control Prompt

Content
View full analysis
标记,请参考 【病历】 体格检查:T:37℃ P:84次/分 R:20次/分 BP:114/64mmHg 诊断:腹腔积液,高血压病3级(很高危) 【质控结果】 有缺陷 体格检查中未提及"移动性浊音,液波震颤" 【病历】 体格检查:T:37℃ P:84次/分 R:20次/分 BP:114/64mmHg 诊断:慢性阻塞性肺疾病急性发作 【质控结果】 有缺陷 体格检查中未提及"杵状指,语音震颤情况,双肺叩诊情况" 【病历】 体格检查:双肺未闻及啰音 诊断:肺部感染肺部阴影 【质控结果】 无缺陷 现在请对下面的门诊病历进行质控,诊断类型仅限于以下几种,如果不是,请直接回答"无缺陷" 1.肺部感染、社区获得性肺炎、肺炎、支气管扩张伴感染,体格检查中未提及"呼吸音(呼吸音、啰音、叩诊)" 2.慢性阻塞性肺疾病,体格检查中未提及"呼吸音(呼吸音、啰音、叩诊、杵状指)" 3.心律失常、室性期前收缩、室性早搏、心房颤动、阵发性心房颤动、持续性心房颤动、房性期前收缩,体格检查中未提及"心音(心率、心律、心音、杂音或额外心音)" 4.脾大,体格检查中未提及"脾脏触诊" 5.肾结石,体格检查中未提及"(肾区叩击痛)" 6.腰椎间盘突出,体格检查中未提及"股神经牵拉试验,直腿抬高试验,状肌试验,直腿抬高加强试验(直腿抬高试验、脊柱生理弯曲或脊柱无畸形、棘突压痛或脊柱压痛)" 7.腹水、腹腔积液,体格检查中未提及"移动性浊音,液波震颤" 8.胸腔积液,体格检查中未提及"语音震颤减弱,叩诊浊音,异常呼吸音(语音震颤或语音共振、叩诊、呼吸音、胸膜摩擦感、胸膜摩擦音)" 9.白血病,体格检查中未提及"全身皮肤有无出血点,瘀斑瘀点(胸骨压痛,皮肤黏膜出血点或瘀斑瘀点)" 10.低蛋白血症,体格检查中未提及"浮肿(眼睑有无浮肿,双下肢水肿)" 11.二尖瓣狭窄,体格检查中未提及"二尖瓣面容(面容、心率、心律、心音、杂音或额外心音、震颤)" 12.肝硬化,体格检查中未提及"腹水,腹壁静脉曲张(肝掌、蜘蛛痣、皮肤黏膜有无黄染、腹壁静脉曲张、肝脏触诊)" 【病历】 体格检查:{pe} 诊断:{dx} 【质控结果】""" )]) ``` ### Technical Analysis The untrusted `pe` and `dx` fields are directly interpolated into the same user message that contains the operational QC instructions and examples. The implementation does not: - Place the controlling policy in a higher-priority system message. - Encode record fields in a structured data format with clear trust boundaries. - Instruct the model to treat all record content as inert data. - Detect instruction-like con ...[truncated 1846 chars]
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (11)

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
93% confidence
Finding

The skill documentation advertises file read, file write, and network-capable code paths, but it does not declare any explicit tool scope or permissions boundary. In an agent ecosystem, this increases the chance the skill will run with broader-than-necessary ambient privileges, enabling unintended data exfiltration or filesystem access if the implementation is invoked on sensitive medical records.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
86% confidence
Finding

The markdown specifies that input records use Chinese field labels and the output is fixed as Chinese strings such as 无缺陷 and 有缺陷. This imposes a specific language/locale behavior, and the file does not state that users can opt into another language or that the restriction is a justified regional requirement.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The manifest describes a skill focused on one QC rule: detecting missing diagnosis-related physical signs. However, the actual LLM prompt embeds a large catalog of specific diagnoses and required findings, plus examples for multiple unrelated conditions, effectively implementing a broader diagnosis-specific QC policy engine. This goes beyond a narrow generic 'missing related signs' checker and changes the skill's operational scope.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The skill sends outpatient record content, which may contain protected health information, to an external LLM service without explicit consent, minimization, or clear disclosure in the workflow. This creates a real confidentiality and compliance risk because sensitive medical data leaves the local trust boundary and is exposed to a third-party processor.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The function accepts a caller-supplied path and reads arbitrary local files, while the skill's stated purpose is to analyze provided outpatient record text. In a broader agent environment, this can be abused to access unrelated local files and potentially disclose their contents if later processed, logged, or sent onward.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

When --save-prepared is enabled, the script writes the full preprocessed medical record to disk in plaintext. Because the data is clinical/PHI and the code provides only a convenience/debug message without privacy warning, access control, redaction, or secure storage handling, this can lead to unauthorized local disclosure, persistence of sensitive data, and accidental inclusion in backups or shared directories.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The script sends full medical record text to an LLM service via run_qc using user-supplied appkey/base/model, but the CLI does not provide an explicit user-facing disclosure or consent flow for transmitting sensitive clinical data. In a medical context, this creates material privacy and compliance risk, especially if the endpoint is remote, configurable, or misdirected to an unintended service.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
94% confidence
Finding

The module docstring is written as a Chinese-only user-facing description, and the CLI description/help text throughout the file is likewise fixed in Chinese. Under the policy, language constraints should not be forced unless the skill offers a language choice or clearly documents a justified locale restriction, which this file does not do.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
89% confidence
Finding

Natural-language strings throughout the module, including the module description and prompt instructions, assume Chinese-only operation and output. The file does not offer a language/locale choice or explain that the skill is intentionally limited to a Chinese-language clinical context.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Low
Category
Not specified by scanner
Confidence
91% confidence
Finding

The description says the skill takes outpatient record text, calls an internal medical LLM, and outputs whether defects exist and why. In addition to producing that result, the implementation creates directories and writes the result to a local file. Persistent file output is extra behavior not stated in the manifest description.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
79% confidence
Finding

The script's docstring, argument descriptions, and user-facing messages are entirely in Chinese, which imposes a specific language/locale on users. There is no opt-in, alternative language support, or documented justification that this tool is limited to a Chinese-language environment.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.