T09 · Insecure Skill Coding Practices
- Location
scripts/chronic_disease_review.py:109- Finding
Medical and identifying data is transmitted without the promised de-identification
- Content
View full analysis
str: blocks: List[str] = [] for item in ocr_data: file_name = item.get("fileName") or "未知文件" page = item.get("page", "") doc_type = item.get("docType") or "未分类文书" text = item.get("ocrText") or "" blocks.append(f"【{doc_type}】{file_name} 第{page}页\n{text}") return "\n\n".join(blocks) def build_review_user_prompt(*, disease_code: str, review_type: str, ocr_data: List[Dict[str, Any]]) -> str: ocr_text = format_ocr_for_prompt(ocr_data) return f"""请对以下材料进行「{review_type}」,病种:{disease_code}。 材料正文: {ocr_text} 请输出 JSON,字段如下(均为字符串): - final_decision:审核结论,取值为「通过」「不通过」「待补充」之一 - reasoning:审核原因说明,需与结论一致 示例: {{"final_decision": "通过", "reasoning": "..."}}""" ``` ```python content = llm( [ {"role": "system", "content": SYSTEM_PROMPT}, {"role": "user", "content": user_prompt}, ] ) ``` ### Technical Analysis `SKILL.md` states that identifying information is strictly de-identified before being sent to any model or interface. The implementation contains no corresponding redaction or data-minimization stage. `format_ocr_for_prompt()` directly inserts the source filename, document type, page number, and complete `ocrText` into the model prompt. The prompt is subsequently sent to the configured LLM endpoint. Names, identity numbers, phone numbers, addresses, insurance information, diagnoses, test results, and other protected medical data can therefore leave the local environment unchanged. The validation function only verifies the input shape and presence of `ocrText`; it does not detect or remove sensitive fields. ### Attack Path 1. A user supplies an OCR document cont ...[truncated 931 chars]- Remediation
View remediation
