T09 · Insecure Skill Coding Practices
- Location
scripts/struct_followup_record.py:61- Finding
Medical Records Are Transmitted to a Third Party Without Implemented De-identification
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This skill performs the advertised medical-record structuring task, but it sends sensitive medical text to a third-party API while promising de-identification that the code does not implement.
Review this carefully before installing in any environment that handles real patient data. Use only with records you are allowed to send to the named third-party service, assume the full record may be transmitted unless the publisher adds verifiable redaction, and avoid broad document/OCR parsing on untrusted files without sandboxing and pinned dependencies.
scripts/struct_followup_record.py:61Medical Records Are Transmitted to a Third Party Without Implemented De-identification
SKILL.md:121Python Dependencies Are Installed Without Version or Integrity Pinning
The document states that user input and intermediate results are not persisted locally and are destroyed after the call, but later defines output paths and options that save prepared text and structured JSON to disk. For sensitive medical records, this contradiction can mislead operators into handling regulated data insecurely and may cause retention of PHI on local storage or in backups.
Referenced artifact was not completely inspected
也支持通过统一入口 `scripts/run.py` 直接输入 `pdf/doc/docx/xls/xlsx/csv/txt/json`。
Referenced artifact was not completely inspected
也支持通过统一入口 `scripts/run.py` 直接输入 `pdf/doc/docx/xls/xlsx/csv/txt/json`。
Referenced artifact was not completely inspected
也支持通过统一入口 `scripts/run.py` 直接输入 `pdf/doc/docx/xls/xlsx/csv/txt/json`。
The code reads the medical record from disk and transmits it to a remote API without any privacy warning, consent flow, or sensitivity notice. In the context of outpatient follow-up records, the content is likely to include PHI/PII and clinical details, so silent exfiltration to an external processor materially increases privacy, regulatory, and data-governance risk.
The skill advertises broad capabilities in practice (network access, file I/O, shell/external tools) but does not declare an explicit tool/permission scope. That weakens enforcement and review, making it easier for a consumer to invoke the skill with more privilege than expected, especially in a medical-data workflow involving external services and local files.
The skill claims strict de-identification before sending data to any model or API, yet the documentation only describes passing prepared clinical text to a third-party backend and does not specify any concrete desensitization step within the skill. In a medical context, this creates a real risk that protected health information is transmitted off-host based on a misleading privacy assurance.
Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.
| `pdftotext`(poppler-utils) | 提取 PDF 文本 | 输入为 pdf 且未安装 pypdf 时必须 |
| `tesseract`(含 chi_sim+eng 语言包) | 图片 OCR | 输入为 png/jpg/bmp/tif 等图片时必须 |
安装(Ubuntu/Debian):`sudo apt-get install libreoffice poppler-utils tesseract-ocr tesseract-ocr-chi-sim`
> 仅使用 TXT/JSON 输入时,无需安装任何第三方包或外部工具。
The script is built around Chinese-only medical section names and aliases, and later prefers the Chinese OCR language pack (chi_sim+eng) when processing images. This imposes a specific language/locale behavior with no user opt-in or alternative language selection, which matches the policy concern for forced language handling.
The skill's stated purpose is structuring medical follow-up text, but it also executes external document-conversion and OCR binaries on user-supplied files. That broadens the trust boundary significantly and creates unnecessary exposure to file parser vulnerabilities, denial of service, and environmental abuse if the runtime is not tightly controlled.
This code invokes LibreOffice/soffice on attacker-supplied office documents for conversion. While not a shell injection bug, it expands the attack surface to complex external parsers that have a history of file-parsing vulnerabilities, so processing untrusted documents can lead to crashes, denial of service, or potentially code execution in the converter context.
if not office_bin:
raise PreprocessError("libreoffice/soffice not found for office document conversion.")
with tempfile.TemporaryDirectory(prefix="med-followup-record-") as tmp_dir:
proc = subprocess.run(
[
office_bin,
"--headless",
This code invokes pdftotext on untrusted PDF input. PDFs are a common carrier for parser bugs and resource-exhaustion attacks, so delegating parsing to an external binary increases exposure beyond the skill's core text-structuring purpose, especially if arbitrary user files are accepted.
pdf_to_text = shutil_which("pdftotext")
if pdf_to_text:
proc = subprocess.run(
[pdf_to_text, "-layout", str(path), "-"],
stdout=subprocess.PIPE,
stderr=subprocess.PIPE,
subprocess module calls execute external commands. Without careful input validation, this enables command injection.
cmd = [tesseract_bin, str(path), "stdout"]
if lang_arg:
cmd.extend(["-l", lang_arg])
proc = subprocess.run(cmd, stdout=subprocess.PIPE, stderr=subprocess.PIPE, text=True, check=False)
if proc.returncode != 0 or not proc.stdout.strip():
raise PreprocessError(f"Image OCR failed: {proc.stderr.strip() or 'no text returned'}")
return proc.stdout
subprocess module calls execute external commands. Without careful input validation, this enables command injection.
def detect_tesseract_langs(tesseract_bin: str) -> Sequence[str]:
proc = subprocess.run(
[tesseract_bin, "--list-langs"],
stdout=subprocess.PIPE,
stderr=subprocess.PIPE,
The skill sends full outpatient follow-up medical record text to an external company-controlled API, but the skill description presents the behavior as simple local JSON structuring and does not disclose third-party processing. Because the data is medical in nature, this creates a meaningful confidentiality and compliance risk: users may unknowingly transmit protected health information to a remote service outside their expected trust boundary.
The argument help text says '--timeout' is for HTTP requests and '--diag-id' is used by a backend service, yet this script contains no HTTP request code or backend service interaction. That documentation describes behavior not implemented in this file, which creates an intent/documentation mismatch for users reviewing the skill entrypoint.
The stated purpose is to structure follow-up medical record text and output JSON, but this implementation accepts any output path and writes files anywhere the process has access. That general filesystem write capability is broader than the narrow transformation purpose described in the manifest.
No suspicious patterns detected.