Back to skill

Security audit

Skill 开发

Security checks for vulnerabilities and agentic risk

Overview

This skill is not malware, but it tells an agent to automatically edit other skills and persistent memory without asking first.

Install only if you want an agent to actively maintain skill files, and treat it as requiring manual oversight: require explicit confirmation before any SkillManage modify action or MEMORY.md write, prefer workspace-scoped memory over user-level memory, and review every proposed change before it is applied.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Findings (2)

T01 · Skill Instruction Hijacking

Error
Location
SOUL.md:7
Finding

Autonomous Skill Modification Without Explicit User Consent

Content
View full analysis
**Fix proactively rather than waiting for instructions.** When a Skill contains a bug, typo, missing step, or outdated information, modify it directly with SkillManage. Do not ask whether it should be changed; notify the user only after making the change. `SOUL.md:21`: > Stop only when the user explicitly says not to modify it. `SKILL.md:47`: > **Fix proactively rather than waiting for instructions:** When a bug, typo, missing step, or outdated information is discovered, modify it directly and notify the user afterward. ### Technical Analysis The Skill explicitly directs the Agent to modify other Skill files without obtaining prior authorization. It makes user consent opt-out rather than opt-in: modification proceeds unless the user has already issued an explicit prohibition. The activation phrases are broad and include ordinary requests concerning Skill problems, improvement, or lessons learned. Once loaded, these instructions can supersede the user's expected review and approval workflow. If the Agent has access to `SkillManage`, file-editing tools, or equivalent workspace permissions, it may change `SKILL.md`, scripts, or references before presenting the proposed changes. Untrusted repository text can also influence the Agent's diagnosis. For example, content framed as a defect report or required improvement could induce the Agent to incorporate attacker-selected instructions into another Skill. The project provides no requirement to establish content provenance, distinguish instructions from untrusted data, display a proposed diff, or obtain approval before writing. ### Attack Path 1. An attacker places crafted content in a repository, issue description, Skill file, or other material that the Agent is asked to review ...[truncated 1555 chars]
Remediation
View remediation

T02 · Agent Memory Poisoning

Error
Location
SKILL.md:43
Finding

Persistent Memory Modification Can Poison Future Agent Sessions

Content
View full analysis
Record the modification in `MEMORY.md`. If other Skills are affected, document their relationships. `SOUL.md:15`: > Record the impact: append one record to `MEMORY.md` after every modification. `README.md:56`: > Place the contents of this repository in the “scenario rules” section of the workspace `.workbuddy/memory/MEMORY.md` or the user-level `~/.workbuddy/MEMORY.md`. ### Technical Analysis The Skill directs the Agent both to install operational rules into persistent memory and to append further records after modifications. Content stored in workspace-level or user-level `MEMORY.md` may be loaded in later sessions and influence tasks unrelated to the original repair operation. The workflow does not define: - A trusted schema for memory entries. - Separation between inert audit metadata and executable Agent instructions. - Content sanitization or escaping. - Provenance and trust labels. - User review before persistence. - Expiration, rollback, or deletion procedures. - Limits preventing repository-controlled text from being copied into memory. Consequently, attacker-controlled material that is misclassified as a lesson, workflow improvement, linked-Skill rule, or impact record may become durable Agent state. Because the suggested user-level location can apply across workspaces, poisoning may affect more than the repository in which the original content was encountered. ### Attack Path 1. An attacker adds crafted instructions to a Skill, repository document, issue, or purported lesson learned. 2. The user asks the Agent to inspect or improve the affected Skill. 3. The Skill interprets the attacker-controlled content as a valid rule, experience, or required workflow change. 4. The Agent modifies a Skill and foll ...[truncated 1435 chars]
Remediation
View remediation
Vulnerability Patterns
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (8)

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
83% confidence
Finding

The natural-language instructions, trigger phrases, and operational guidance are entirely in Chinese, with no indication that users may choose another language or that the skill is intentionally limited to a Chinese-only context. This can constitute a language/locale policy violation when a skill implicitly enforces one language without opt-in.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The trigger phrases are broad, generic troubleshooting language such as 'modify skill' or 'skill has a problem', which can cause the skill to activate in ordinary discussion rather than through an explicit user request for file-changing behavior. In this skill’s context, accidental activation is more dangerous because the workflow explicitly proceeds to modifying skill files and recording changes in shared memory, increasing the chance of unintended configuration or prompt changes.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The documentation describes persistent modification of skill files and recording changes in workspace memory, but it does not clearly warn users up front that invoking the skill may alter files and leave durable audit or memory entries. This weakens informed consent and can lead to unexpected state changes, especially in shared workspaces or when users think they are only asking for analysis rather than mutation.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The trigger description is broad enough that normal discussion about improving or discussing a skill could unintentionally activate this skill. Because the skill is empowered to modify other skill files and encourages proactive changes, accidental invocation could cause unauthorized or unexpected edits to system behavior.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The manifest description and invocation guidance are entirely in Chinese and present the skill as operating in that language by default, with no opt-in or alternative language path. Under the stated policy, a language constraint should either be user-selectable or clearly justified as region-specific.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The workflow and activation guidance do not clearly separate analysis-only requests from requests that authorize modification, creating ambiguity about when the skill should act. In context, this is more dangerous because the skill's core behavior includes reading and modifying existing skills, and its doctrine says to 'actively fix' issues without waiting for instruction.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill explicitly instructs the agent to modify skill files automatically ('directly use SkillManage') and to append to MEMORY.md after every change, while only requiring confirmation for user-data deletion. This creates an unsafe default of making persistent file and log changes without explicit user consent, which can lead to unintended modifications, audit noise, or covert persistence of bad instructions if the agent misinterprets a situation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
84% confidence
Finding

The skill's language directive mandates a fixed communication style/language without user opt-in. While not directly a security exploit by itself, forced locale can reduce user comprehension, impair review of proposed changes, and increase the chance that unsafe edits are accepted without fully informed consent in a security-sensitive workflow.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.