Back to skill

Security audit

tcm-wellness

Security checks for vulnerabilities and agentic risk

Overview

This TCM wellness skill is not malicious, but it stores sensitive health histories persistently and has unsafe user-record handling that needs careful review before installation.

Review this skill before installing, especially in shared or multi-user environments. It should only handle real health information after adding explicit opt-in storage, strong per-user identity binding, user-ID validation, restrictive file permissions, record deletion/export controls, and safer defaults that avoid persistent medical memory unless requested.

Vulnerability Patterns
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
Findings (5)

T09 · Insecure Skill Coding Practices

Error
Location
scripts/init_memory.py:119
Finding

Unsanitized User Identifier Allows Path Traversal and Arbitrary File Writes

Content
View full analysis
Remediation
View remediation
str: if not re.fullmatch(r"[A-Za-z0-9_-]{1,64}", value): raise ValueError("Invalid user identifier") return value def contained_path(base: Path, child: str) -> Path: base = base.resolve() destination = (base / child).resolve() if not destination.is_relative_to(base): raise ValueError("Path escapes the storage directory") return destination ``` 5. Reject symbolic-link destinations or operate relative to a verified directory handle. 6. Use exclusive creation or atomic replacement where overwriting existing records is not required. 7. Run the script with a dedicated, least-privileged operating-system account. ]]>

T09 · Insecure Skill Coding Practices

Error
Location
scripts/sleep_reflection.py:50
Finding

Reflection Processing Allows Out-of-Scope File Discovery and Modification

Content
View full analysis
list: """获取用户所有记忆块文件""" blocks_dir = root / "blocks" / user_id if not blocks_dir.exists(): return [] return sorted(blocks_dir.glob("*.md")) ``` ```python def archive_old_blocks(root: Path, user_id: str, config: dict, dry_run: bool = False): """归档过期的记忆块""" blocks_dir = root / "blocks" / user_id if not blocks_dir.exists(): print(" ℹ️ 无记忆块可归档") return retention = config.get("retention", {}) now = datetime.now() archived_count = 0 for block in sorted(blocks_dir.glob("*.md")): fm = parse_block_frontmatter(block) priority = fm.get("priority", "medium") ts = fm.get("timestamp", "") if not ts: continue try: block_date = datetime.fromisoformat(ts) except: continue days_map = { "high": retention.get("high_blocks_days", 180), "medium": retention.get("medium_blocks_days", 90), "low": retention.get("low_blocks_days", 30), } retention_days = days_map.get(priority, 90) if retention_days > 0: threshold = now - timedelta(days=retention_days) if block_date < threshold: if not dry_run: content = block.read_text(encoding="utf-8") if "archived: true" not in content: content = content.replace("---\n", "---\narchived: true\n", 1) block.write_text(content, encoding="utf-8") ``` ### Technical Analysis The reflection script accepts the same untrusted `user_id` and uses it directly to select a directory. No canonicalization or containment check en ...[truncated 1485 chars]
Remediation
View remediation

T05 · Unauthorized Access and Privilege Escalation

Error
Location
references/memory_system.md:263
Finding

Symptom-Based Identity Inference Can Expose Another User's Health Records

Content
View full analysis
_profile.md`(健康档案) 2. 读取 `long_term/_index.md`(长记忆索引) ``` The same behavior is required by `SKILL.md:63-67`, which directs the agent to identify returning users through nicknames, symptom matching, and follow-up language before loading their profiles. ### Technical Analysis Symptoms, constitution traits, nicknames, and phrases such as “last time” are not authentication factors. Many users can share common symptoms such as insomnia, fatigue, anxiety, or digestive discomfort. The Skill nevertheless treats those traits as sufficient evidence of identity and then directs the agent to load a stored health profile and long-term index. No authenticated platform identity, secret identifier, confirmation step, confidence threshold, or authorization check is required. This creates an insecure direct object reference at the agent workflow level: inferred identity controls access to another record. ### Attack Path 1. A victim has an existing long-term health profile. 2. Another user describes symptoms or constitution traits that overlap with the victim's records, or deliberately claims to be returning. 3. The agent infers that the requester is the victim. 4. The agent loads the corresponding profile and index. 5. Historical symptoms, medication information, known conditions, treatment preferences, or prior recommendations are incorporated into the response. 6. The requester receives information belonging to the victim or causes the victim's profile to be updated with incorrect data. ### Impact Assessment The vulnerab ...[truncated 377 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
scripts/init_memory.py:125
Finding

Sensitive Medical Records Are Persisted in Plaintext Without Consent or Access Controls

Content
View full analysis
- 年龄段:<未提供> - 生活特征:<待采集> ## 体质辨识 - **当前体质**:<待辨证> - **体质演变**: - **体质倾向**:<待判断> ## 病史脉络 ### 脾胃系统 - (暂无记录) ### 肝胆系统 - (暂无记录) ### 心肺系统 - (暂无记录) ### 肾与膀胱系统 - (暂无记录) ## 持续问题追踪 | 问题 | 首次出现 | 最近一次 | 状态 | 趋势 | |------|---------|---------|------|------| | (暂无) | - | - | - | - | ## 敏感信息 - 过敏史:<无/待补充> - 用药情况:<无/待补充> - 已知疾病:<无/待补充> ## 调理偏好 - 愿意尝试:<待了解> - 不愿尝试:<无> - 方案依从性:<待评估> """ ``` The profile is written as an ordinary UTF-8 Markdown file: ```python with open(path, "w", encoding="utf-8") as f: f.write(content) ``` `SKILL.md:70-76` further requires memory blocks, health profiles, and indexes to be written immediately after an assessment. ### Technical Analysis The design stores symptoms, medical history, allergies, medication use, known diseases, sex, age range, emotional information, constitution assessments, and treatment feedback in plaintext Markdown. No encryption at rest, platform authorization, explicit consent gate, data minimization control, restrictive file-permission setup, user deletion interface, or per-account storage isolation is implemented. File permissions depend entirely on the process's ambient umask and filesystem configuration. The use of an “anonymous ID” does not adequately de-identify the records. The specification permits nickname-derived identifiers and detailed longitudinal health histories, which can make re-identification possible. ### Attack Path 1. A user asks the Skill for health or wellness advice. 2. The Skill performs an assessment and automatically initializes or updat ...[truncated 760 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Warning
Location
scripts/sleep_reflection.py:415
Finding

Global Date-Only Reflection Files Cause Cross-User Overwrites and State Interference

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (36)

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The skill also performs retrospective analysis, archival updates, and automated reflection workflows that go beyond interactive question-answering. Hidden secondary behavior involving batch reads, file mutation, and automated report generation increases the risk of unauthorized processing and retention of sensitive user health history.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The skill also performs retrospective analysis, archival updates, and automated reflection workflows that go beyond interactive question-answering. Hidden secondary behavior involving batch reads, file mutation, and automated report generation increases the risk of unauthorized processing and retention of sensitive user health history.

Content

No source excerpt is available for this finding.

Missing User Warnings

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The skill explicitly states it will remember each user's constitution, medical history, adjustment plans, and feedback over time, but it provides no clear consent, retention, deletion, or privacy notice. Because this is health data, silent long-term retention materially raises privacy, compliance, and unauthorized disclosure risks.

Content

No source excerpt is available for this finding.

Missing User Warnings

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The skill instructs the system to create and update health-profile files and long-term indexes without warning users that sensitive data will be written to disk. Persisting structured health dossiers can expose highly sensitive personal information if logs are accessed by other skills, users, operators, or compromised tooling.

Content

No source excerpt is available for this finding.

Missing User Warnings

High
Category
Not specified by scanner
Confidence
96% confidence
Finding

The system persists sensitive health history and follow-up data but does not present a clear user warning about retention, privacy impact, or downstream reuse. In a health context, silent persistence materially increases privacy risk because users may disclose sensitive facts assuming an ephemeral consultation.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The architecture stores user-level health data in a shared directory across projects, which increases the chance of unintended correlation, over-collection, and reuse outside the original context. Cross-project sharing is especially risky for health information because it broadens exposure and weakens contextual integrity.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
97% confidence
Finding

The skill directs the agent to read and write local files for per-user memory, profiles, indexes, reflections, and evolution logs, but it declares no explicit tool/permission scope. This creates an overbroad capability surface where sensitive health data could be persisted or accessed without clear sandboxing, review, or least-privilege constraints.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

L003 将技能描述为当用户询问“中医、养生、调理、体质”等大量通用词汇时触发,这些词在普通聊天、泛健康咨询或非中医语境中都很常见,缺少足够的限定条件。虽然列出了触发词,但没有提供排除条件或负例,仍可能造成非预期调用。

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

L033-L045 规定用户只要描述头痛、失眠、胃胀、乏力等常见不适,或请求“健康总结”“回顾健康记录”就加载本技能,但没有说明必须是中医养生语境,也没有列出排除条件。这会让技能与广泛的日常健康表达、一般医疗咨询甚至普通总结请求发生重叠。

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

Natural-language instructions to persist and recall users' health history across sessions create a concrete data-retention and cross-session privacy risk. Even if intended for continuity of care, stored health context can be leaked, over-retained, or incorrectly associated with the wrong user without strong identity and storage controls.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The instructions direct ongoing reading, updating, indexing, and summarizing of per-user health profiles, which is systematic handling of sensitive personal data. This increases the attack surface for accidental disclosure, excessive profiling, and unauthorized secondary use, especially because the skill also includes automated reflection and evolution logging.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
84% confidence
Finding

All user-facing natural-language content in this skill file is presented in Chinese, and the document does not indicate that users may choose another language or that the skill is intentionally limited to a Chinese-speaking audience. Under the stated policy, forcing a specific language without opt-in is a natural-language policy concern.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The document explicitly defines persistent cross-session memory for individual users' constitution, health history, treatment plans, and follow-up data. For a wellness consultation skill, this materially expands data collection beyond a single interaction and creates a durable repository of sensitive health information without clear necessity, consent, or minimization controls.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The instructions explicitly direct the system to store and later reuse detailed health history across sessions, including diagnosis-related details and feedback. Re-surfacing prior sensitive medical information in later responses increases the risk of inappropriate disclosure, mistaken identity linkage, and overreach relative to a wellness advisory skill.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The document claims to use a hash for privacy, but the examples show readable nicknames or deterministic symptom/constitution/age composites that are easily guessable and re-identifiable. This creates a false sense of anonymization while still embedding sensitive attributes directly into identifiers and filenames.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The privacy boundary says only TCM-related health information will be stored, but later templates and profiles include broader medical data categories. This mismatch encourages collection creep and undermines user expectations about what sensitive information is retained.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
98% confidence
Finding

The long-term profile explicitly collects allergies, medications, and known diseases, which are highly sensitive medical data categories. Retaining these fields in a general wellness skill significantly raises privacy and misuse risk, especially without strong safeguards, necessity justification, or explicit consent.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The retrieval workflow directs the agent to load prior health profiles and inject historical data into current answers. This increases the risk of stale, incorrect, or misattributed medical context influencing advice, and can expose prior sensitive details unexpectedly during a new conversation.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

Broad trigger phrases like '总结一下' or '回顾一下' can be invoked during ordinary conversation and unintentionally activate a reflection workflow over stored health records. In a system with persistent sensitive memory, accidental triggering increases the chance of over-processing and unintended disclosure.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

Automatic reflection and profile updates perform additional analysis and modification of stored health records without clearly notifying the user at the time of processing. This expands processing beyond the original conversational purpose and can surprise users with derived inferences they did not knowingly authorize.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The reflection process bulk-collects all active memory blocks and indexes for a user and generates aggregated summaries and inferences. Bulk processing amplifies privacy impact because it centralizes sensitive health history, derives new conclusions, and increases the blast radius if the data is exposed or mishandled.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The self-evolution mechanism allows the skill to alter recommendation behavior and even record 'knowledge corrections' based on user feedback, which can introduce unsafe drift in a health-adjacent system. Without governance, validation, and separation between user anecdotes and curated knowledge, the system may propagate incorrect or personalized medical heuristics.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The document provides broad dietary and quasi-therapeutic recommendations for symptoms, constitutions, and seasonal health management without clear cautions about contraindications, allergies, pregnancy, chronic disease, medication interactions, or when to seek licensed medical care. In a wellness skill that may be used by symptomatic users, this can lead users to self-manage potentially serious conditions or follow unsuitable recommendations that worsen health outcomes.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

This file provides detailed syndrome differentiation, symptom matching, and treatment-reference content for health conditions without any warning that it is informational only and not a substitute for licensed medical evaluation. In the context of a wellness skill that may be used by symptomatic users, this increases the risk of self-diagnosis, delayed treatment, and unsafe reliance on TCM references for potentially serious conditions.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The module docstring warns that --force '会清除已有数据', and later console output repeats that it will clear existing data and rebuild. In reality, create_directory_structure uses exist_ok=True without deletion, and the various create_* functions only overwrite specific files if present; existing memory blocks, user profiles, and other undeclared files remain intact.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.