Back to skill

Security audit

unisound-primary-diagnosis-surgery-selection

Security checks for vulnerabilities and agentic risk

Overview

The skill has a coherent medical-coding purpose, but it can send medical records and an API key to any caller-supplied LLM URL and can optionally persist prepared medical text despite saying it does not persist data.

Review before installing. Use this only in an environment where medical data is already approved for the configured LLM service, avoid real identifiers unless de-identified, do not allow untrusted users or wrappers to set --base, and avoid --save-prepared unless the output directory is protected and retention is acceptable.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/run.py:219
Finding

Caller-Controlled LLM Endpoint Exposes Bearer Credentials and Medical Data

Content
View full analysis
str: url = f"{base.rstrip('/')}/chat/completions" headers = {"Authorization": f"Bearer {appkey}"} if appkey else {} payload = { "model": model, "messages": [{"role": "user", "content": prompt}], "temperature": 0, } response = _http_post(url, payload, headers, timeout=timeout) ``` The destination is exposed directly as a command-line option: ```python parser.add_argument("--base", default=DEFAULT_LLM_BASE, help=f"Internal LLM base URL (default: {DEFAULT_LLM_BASE}).") ``` The caller-controlled value is passed to the network request without validation: ```python response = run( payload, base=args.base, model=args.model, appkey=args.appkey, timeout=args.timeout, ) ``` ### Technical Analysis The `--base` argument accepts an unrestricted URL. `call_llm()` appends `/chat/completions` to that value and sends both the complete medical prompt and the supplied API credential in an `Authorization: Bearer` header. The implementation does not enforce HTTPS, validate the destination hostname or port, restrict requests to approved model providers, or prevent redirects to another origin. Consequently, command-line input controls the destination to which secrets and sensitive medical information are transmitted. This is particularly significant because the documented input contains patient medical records, while `--appkey` is an authentication secret. The flaw does not independently provide local code execution or elevated operating-system privileges, but it permits disclosure of both categories of sensitive data. ### Attack Path 1. An attacker influences the command invocation, a wrapper co ...[truncated 1300 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/run.py:102
Finding

Direct-Prompt Mode Bypasses Candidate-Only Output Validation

Content
View full analysis
str: raw_prompt = str(payload.get("prompt") or "").strip() if raw_prompt: return raw_prompt ``` A JSON payload containing a prompt is accepted even when candidate arrays are absent: ```python if resolved_type == "json": try: payload = _read_json(path) if payload.get("prompt") or ( any(payload.get(key) for key in ("admission", "入院情况", "treatment", "诊疗过程", "pathology", "病理")) and (payload.get("candidate_diagnoses") or payload.get("候选诊断") or candidate_diagnoses) and (payload.get("candidate_surgeries") or payload.get("候选手术") or candidate_surgeries) ): if candidate_diagnoses: payload["candidate_diagnoses"] = candidate_diagnoses if candidate_surgeries: payload["candidate_surgeries"] = candidate_surgeries return payload ``` Output membership checks are conditional and therefore skipped when the candidate lists are empty: ```python diagnoses = _string_list(source_payload.get("candidate_diagnoses") or source_payload.get("候选诊断")) surgeries = _string_list(source_payload.get("candidate_surgeries") or source_payload.get("候选手术")) if diagnoses and main_diagnosis not in diagnoses: raise ValueError(f"Model output main_diagnosis is not among candidate diagnoses: {main_diagnosis}") if surgeries and main_surgery not in surgeries: raise ValueError(f"Model output main_surgery is not among candidate surgeries: {main_surgery}") ``` ### Technical Analysis The skill states that the main diagnosis and surgery must be selected from caller-provided candidate lists. That bounda ...[truncated 1893 chars]
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Taint TrackingDirect Taint Flow, Variable-Mediated Taint Flow, Credential Exfiltration Chain
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (7)

Tainted flow: 'req' from pathlib.Path.read_text (line 201, file read) → urllib.request.urlopen (network output)

High
Category
Data Flow
Confidence
97% confidence
Finding

The skill sends extracted medical record text and candidate data to a remote LLM endpoint, and the destination base URL is user-configurable via --base. In a medical context this can expose sensitive patient data to external or attacker-controlled infrastructure, making the network egress materially dangerous even if intended for an internal API.

Content

Scanner excerpt · scripts/run.py (reported line 208)May include surrounding context.

python
headers={"Content-Type": "application/json", **{key: value for key, value in headers.items() if value}},
    )
    try:
        opener = urllib.request.urlopen(req) if not timeout else urllib.request.urlopen(req, timeout=timeout)
        with opener as resp:
            body = resp.read().decode("utf-8", errors="replace")
            return json.loads(body)

Missing User Warnings

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The core workflow transmits patient record content to a remote LLM service without any explicit warning, confirmation, or visible safeguard around PHI handling. Because this skill processes medical records, silent disclosure to a remote model endpoint raises serious confidentiality and compliance risks.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
90% confidence
Finding

The skill explicitly documents capabilities to read multiple local file formats, optionally save preprocessed medical text to disk, and send patient-derived content over the network to an internal model endpoint, yet it declares no tool scope or permission boundaries. In a medical context this is dangerous because the skill handles sensitive health information and API credentials; without explicit permission declarations, callers and enforcement layers may not be able to constrain file, write, or network behavior appropriately.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
84% confidence
Finding

The natural-language content of the skill, including description, usage, and warnings, is entirely in Chinese. Under the policy, forcing a specific language without offering the user a choice or documenting a justified locale restriction is a language/locale policy violation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The generated prompt instructs the model entirely in Chinese and requires a Chinese-structured response, but the script does not offer a language or locale option. This enforces a specific language behavior without user opt-in or a documented region-specific justification in the file.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The --save-prepared option writes preprocessed medical text to disk for debugging without a safety warning or protective controls. In this skill's healthcare setting, that creates avoidable PHI-at-rest exposure through local files, logs, backups, or shared workspaces.

Content

No source excerpt is available for this finding.

Tainted flow: 'text' from pathlib.Path.read_text (line 349, file read) → pathlib.Path.write_text (file write)

Medium
Category
Data Flow
Confidence
65% confidence
Finding

Data from a source is assigned to a variable that is later passed to a sink, creating a variable-mediated taint flow.

Content

Scanner excerpt · scripts/run.py (reported line 354)May include surrounding context.

python
if output_path:
        out_path = Path(output_path)
        out_path.parent.mkdir(parents=True, exist_ok=True)
        out_path.write_text(text, encoding="utf-8")
    print(text)
    return 0

Static analysis

No suspicious patterns detected.