T09 · Insecure Skill Coding Practices
- Location
scripts/chronic_disease_review.py:114- Finding
Medical Records Are Transmitted Without the Documented De-identification
- Content
View full analysis
Vulnerability Details
File Location:
scripts/chronic_disease_review.py:114-146, 178-187
Vulnerability Type: Sensitive medical-data disclosure
Risk Level: HighThe documentation states that identifiable information is strictly redacted before being sent to any model or interface. The implementation contains no redaction step: filenames and complete OCR text are inserted directly into the prompt and transmitted to the configured model service.
Complete Code Snippet
python def format_ocr_for_prompt(ocr_data: List[Dict[str, Any]]) -> str: blocks: List[str] = [] for item in ocr_data: file_name = item.get("fileName") or "未知文件" page = item.get("page", "") doc_type = item.get("docType") or "未分类文书" text = item.get("ocrText") or "" blocks.append(f"【{doc_type}】{file_name} 第{page}页\n{text}") return "\n\n".join(blocks) SYSTEM_PROMPT = """你是医疗保险门诊慢特病(慢病)理赔审核助手。 根据用户提供的 OCR 病历/检验等文书文本,判断是否符合该慢病的门诊慢特病认定或理赔审核要求。 仅依据给定文本作答,不要编造未出现的检查结果或诊断。 输出必须是合法 JSON,且只包含一个 JSON 对象,不要 markdown 代码块或额外说明。""" def build_review_user_prompt(*, disease_code: str, review_type: str, ocr_data: List[Dict[str, Any]]) -> str: ocr_text = format_ocr_for_prompt(ocr_data) return f"""请对以下材料进行「{review_type}」,病种:{disease_code}。 材料正文: {ocr_text} 请输出 JSON,字段如下(均为字符串): - final_decision:审核结论,取值为「通过」「不通过」「待补充」之一 - reasoning:审核原因说明,需与结论一致 示例: {{"final_decision": "通过", "reasoning": "..."}}"""python user_prompt = build_review_user_prompt( disease_code=disease_code, review_type=review_type or "慢病审核", ocr_data=ocr_data, ) content = llm( [ {"role": "system", "content": SYSTEM_PROMPT}, {"role": "user", "content": user_prompt}, ] )Technical Analysis
format_ocr_for_promptcopiesfileName,docType, and the completeocrTextvalue into the model prompt. Medical records can ...[truncated 1367 chars]- Remediation
View remediation
Remediation Suggestions
- Implement a dedicated de-identification stage before prompt construction.
- Detect and redact names, identity and insurance numbers, telephone numbers, email addresses, full addresses, dates of birth, and other identifying attributes.
- Exclude filenames and document metadata unless they are strictly required for the review.
- Run residual-sensitive-data detection after redaction and block transmission when unresolved identifiers remain.
- Use context-aware medical de-identification rather than relying exclusively on regular expressions.
- Add tests containing representative identifiers and verify that outbound request bodies do not contain them.
- Require explicit informed consent if any sensitive data must be transmitted.
- Update the documentation if the implementation cannot guarantee the stated level of redaction.
