T09 · Insecure Skill Coding Practices
- Location
scripts/chronic_disease_review.py:126- Finding
Medical OCR data is transmitted to an external service without de-identification
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This medical OCR review skill is purpose-aligned, but it contradicts its own privacy promises by sending raw health text to a remote service and saving results to disk by default.
Review before installing or using with real patient data. Treat inputs and outputs as sensitive medical records, confirm the backend destination and legal basis for sending data there, and do not rely on the stated de-identification or no-persistence guarantees unless the publisher fixes the implementation and documentation.
scripts/chronic_disease_review.py:126Medical OCR data is transmitted to an external service without de-identification
scripts/chronic_disease_review.py:140Sensitive medical review results are persistently stored despite a no-persistence guarantee
The documented purpose materially differs from the described or detected behavior: the skill claims to review OCR arrays for chronic-disease approval and output both raw JSON and a conclusion, while the implementation behavior appears to only format/summarize an existing review response. In a medical workflow, this can mislead operators about what data is processed, what controls apply, and whether an actual review occurred, causing unsafe reliance and possible disclosure of sensitive data to unintended components.
The documentation explicitly promises no local persistence and destruction after the call, yet later states that outputs are written to disk by default. For medical OCR and review content, this is a significant data-handling contradiction that can result in unintended retention of sensitive health information and noncompliance with privacy expectations or policy.
Referenced artifact was not completely inspected
- **发布约束**:示例输入、运行输出、自测脚本均放在 skill 包外(分别位于 `../data/`、`../runs/`、`../self_tests/`),skill 目录内仅保留可发布的核心文件(`scripts/`、`SKILL.md`、`_meta.json`)。
The skill advertises capabilities that include file read, file write, and network access, but it does not declare any explicit tool scope or permission boundaries. This increases the risk of overbroad execution in hosts that rely on manifest-level restrictions, especially because the skill handles medical OCR data and can contact a backend service and write outputs to disk.
The natural-language content of the skill, including description, instructions, and parameter explanations, is presented only in Chinese. Under the policy, forcing a specific language without user opt-in or a clearly justified locale limitation is a reportable natural-language policy issue.
The skill claims strict de-identification before any model or API call, but the documented interface only shows raw OCR content being sent to backend services and does not describe an actual sanitization step. In a healthcare context, absent or unverifiable de-identification can expose personal and medical data to internal or third-party services contrary to user expectations.
The script sends OCR medical data to a remote HTTPS endpoint, which exceeds the skill description's implied local processing model and changes the trust boundary. Even if this is functionally intended, undisclosed external transmission of potentially sensitive health information is a real security/privacy issue because users may assume the OCR JSON is processed locally.
This code path packages OCR medical content and submits it to a remote review API without any explicit user-facing warning or consent step. Since OCR input may include protected health information, silent transmission to a third-party service creates a meaningful privacy and compliance risk.
The script writes both the remote service response and a natural-language medical summary to disk without warning, which can leave sensitive health data in local artifacts. This increases exposure through backups, shared workstations, broad directory permissions, or accidental disclosure.
Data from a source is assigned to a variable that is later passed to a sink, creating a variable-mediated taint flow.
if args.output:
out_path = Path(args.output)
out_path.parent.mkdir(parents=True, exist_ok=True)
out_path.write_text(text, encoding="utf-8")
else:
print(text)
return 0
The script description and argument help require Chinese disease-code values and use Chinese defaults such as '慢病审核', while the output-building flow appears oriented to a single locale. There is no explicit opt-in, language selection, or justification that this tool is intentionally restricted to Chinese-speaking or China-region users.
The script persists raw response JSON and a natural-language summary to local files by default, despite the skill description emphasizing returned output. Because the data concerns chronic disease review, these files may contain sensitive medical information that remains on disk longer than users expect.
No suspicious patterns detected.