Back to skill

Security audit

鲁班.Skill

Security checks for vulnerabilities and agentic risk

Overview

This is a plausible skill optimizer, but it authorizes broad repository edits, scheduled maintenance, and persistent web/user-derived updates without consistently requiring clear user approval.

Review before installing. Use it only in a sandboxed skill repository or after disabling scheduled runs, CASCADE web updates, and automatic MUSE test persistence. Require explicit approval of exact files, sources, and diffs before it writes to Git, references, tests.yaml, or generated prompt files.

Vulnerability Patterns
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Findings (2)

T02 · Agent Memory Poisoning

Error
Location
SKILL.md:271
Finding

Unsanitized Remote Content Can Persistently Poison Agent Instructions

Content
View full analysis

Vulnerability Details

File Location: SKILL.md:271-310
Vulnerability Type: Persistent prompt injection through untrusted remote content
Risk Level: High

Vulnerable Code Snippet

text
## Four. CASCADE Module: Automatic Domain-Knowledge Updates

Trigger conditions:
- Scheduled task: automatically runs every quarter
- User command: update skill knowledge / add the latest information / refresh references
- The skill's references cite external knowledge and were last updated more than 90 days ago

Step 3: Knowledge retrieval
  - Papers: search the cited arXiv ID and check for a new version
  - APIs: fetch the latest documentation and compare the changelog
  - Standards: search for newly published versions

Step 4: Self-reflection
  - Compare new and old knowledge
  - Determine whether changes affect the validity of skill rules
  - Generate an update only when substantive changes exist

Step 5: Append-only update
  - Append new knowledge without deleting old content and include a version number
  - Format: ## [YYYY-MM-DD] Update: xxx → new content

Key design:
- Append rather than delete
- Mark each update with a date and version
- Do not automatically modify rules; only update references

The snippet above is an English rendering of the operative instructions at the cited location.

Technical Analysis

The CASCADE workflow retrieves information from external web sources and appends the resulting material to files under references/. Those files are part of the instruction context consumed by agents in later executions. Consequently, externally controlled text crosses a trust boundary and becomes persistent agent state.

The documented workflow does not require:

  • An allowlist of trusted domains or exact source URLs
  • Source authenticity, signature, or content-hash verification
  • Prompt-injection detection or removal of instruction-like text
  • Strict separation betwe ...[truncated 2401 chars]
Remediation
View remediation

Remediation Suggestions

  1. Disable unattended writes from scheduled knowledge-update jobs. Scheduled runs should produce a read-only report and proposed patch.
  2. Require explicit user approval of the exact source list and exact diff before modifying any instruction-bearing file.
  3. Allow retrieval only from explicitly approved domains and pinned canonical URLs. Do not use unrestricted search results as authoritative sources.
  4. Record provenance for every appended passage, including the canonical URL, retrieval timestamp, content hash, document version, and verification status.
  5. Verify signatures or publisher-provided checksums where available. Pin content hashes for sources that lack signed releases.
  6. Treat fetched material strictly as untrusted data. Delimit and quote it so that it cannot be interpreted as Agent instructions.
  7. Reject or quarantine content containing instruction patterns, tool-call requests, role reassignment, requests to ignore prior rules, encoded payloads, or unrelated operational directives.
  8. Use a two-stage pipeline in which one restricted process retrieves content without file-write or execution privileges, while a separate reviewer produces a sanitized summary.
  9. Store retrieved source material outside Agent instruction files. Only a reviewed, declarative summary should be eligible for inclusion in references/.
  10. Add regression tests with malicious documentation samples to verify that prompt-injection text is quarantined and cannot alter subsequent Agent behavior.

T02 · Agent Memory Poisoning

Error
Location
references/SA-DM.md:228
Finding

Reference Architecture Replicates the Unsafe Web-to-Instruction Persistence Pipeline

Content
View full analysis

Vulnerability Details

File Location: references/SA-DM.md:228-252
Vulnerability Type: Persistent prompt injection through unsanitized search and web-fetch results
Risk Level: High

Vulnerable Code Snippet

text
Step 1: Scan the skill's references/ directory
  - Identify all external references, including paper IDs, API documentation URLs,
    and standard identifiers
  - Record the last update date of each reference

Step 2: Select outdated references
  - Mark references older than the threshold as pending updates
  - Prioritize skills recently used frequently by the user

Step 3: Knowledge retrieval (web_search + web_fetch)
  - Papers: search the cited arXiv ID and check for a new version
  - APIs: fetch the latest documentation and compare the changelog
  - Standards: search for newly published versions

Step 4: Self-reflection
  - Compare new and old knowledge
  - Determine whether changes affect the validity of skill rules
  - Generate an update only when substantive changes exist

Step 5: Update references/ files
  - Append new knowledge without deleting old content and include a version number
  - Format: ## [YYYY-MM-DD] Update: xxx → new content

The snippet above is an English rendering of the operative instructions at the cited location.

Technical Analysis

This reference is described as a methodology document, but it contains concrete operational instructions that can be copied into or consumed by Skill implementations. It explicitly connects web_search and web_fetch results to writes under references/, which are subsequently loaded as Agent context.

No validation stage separates externally supplied content from persistent instruction text. The described self-reflection step assesses whether information is substantively different, but it is not a security control: it does not authenticate sources, identify prompt injection, constrain source domains, or requi ...[truncated 1427 chars]

Remediation
View remediation

Remediation Suggestions

  1. Revise the architecture so retrieval never writes directly to SKILL.md or references/.
  2. Add an explicit trust-boundary stage between retrieval and persistence.
  3. Require trusted-domain allowlisting, canonical URL pinning, provenance recording, and integrity verification.
  4. Place raw retrieved content in a quarantine area that is never loaded as Agent instructions.
  5. Generate a sanitized, declarative summary through a restricted reviewer that has no write or tool-execution permissions.
  6. Require a human to approve the source, sanitized summary, and exact destination diff before persistence.
  7. Specify mandatory prompt-injection checks for role changes, safety-constraint overrides, tool instructions, secret requests, hidden directives, and encoded payloads.
  8. Document that scheduled jobs may detect staleness and prepare proposals but must not autonomously modify instruction-bearing files.
  9. Add negative test cases demonstrating that malicious fetched content is rejected rather than appended.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Tool MisuseTool Parameter Abuse, Chaining Abuse, Unsafe Defaults
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
Findings (16)

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
65% confidence
Finding

Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).

Content

Scanner excerpt · SKILL.md (reported line 555)May include surrounding context.

md
| # | 反模式 | 替代做法 |
|---|--------|----------|
| 1 | 同 context 自评自改 | 必须 spawn 独立子 agent 评分 |
| 2 | `git reset --hard` 当回滚 | 用 `git revert HEAD` 保留追溯链 |
| 3 | 为凑分增冗余 | 触顶信号(连续 2 轮 Δ<2)→ break |
| 4 | 跳过 test-prompts 直接评分 | Phase 0.5 强制设计 2-3 prompts |
| 5 | 轮内改多个维度 | 每轮 1 个维度 |

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

This markdown file is entirely written in Chinese and all example trigger phrases are specified only in Chinese, with no indication that other languages are supported or that the language requirement is optional. Under the policy, forcing a specific language without user opt-in is a natural-language policy violation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The entire user-facing README, including installation and usage instructions, is presented only in Chinese. Under the policy for natural-language violations, forcing a specific language without user opt-in is a reportable issue unless the locale restriction is explicitly justified, which is not stated here.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The trigger list includes broad conversational phrases such as requests to review, optimize, or comment on a skill, which can overlap with ordinary discussion. This is dangerous because accidental activation of a write-capable skill can escalate a benign conversation into repository inspection, testing, or modification behavior without the user's informed intent.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The metadata presents the skill as an on-demand optimizer/reviewer, but the body defines autonomous event-driven and scheduled maintenance behavior over repositories. This capability mismatch is dangerous because users may invoke a seemingly advisory skill that is actually authorized to initiate ongoing operational workflows and queued changes beyond the immediate request.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The documented workflow includes branching, commits, reverts, backups, and persistent result file writes, which exceeds the advertised scope of scoring or optimizing a skill. This is dangerous because it normalizes source-control and filesystem mutation as part of a loosely triggered review skill, increasing the chance of unintended codebase changes and abuse through social or prompt-triggering paths.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The EvoSkill flow captures users' original requests, parameters, outputs, and feedback as failure context, which creates a clear data retention risk for potentially sensitive prompts and generated content. Persisting rich natural-language interaction traces can expose secrets, personal data, proprietary code, or confidential operational details through artifacts, logs, or later model reuse.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

Module-specific trigger phrases are short and weakly scoped, increasing the chance that unrelated user feedback or general discussion text matches them. In a skill with mutation and logging behavior, ambiguous triggers materially increase the risk of unintended analysis, persistence, or repository actions.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The stated rule of 'no new dependencies' conflicts with flows that create validators, merged references, and new files during maintenance or distillation. Such internal policy contradictions are dangerous because they make operator expectations unreliable and create room for a model to justify unbounded artifact creation under the guise of compliance.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
86% confidence
Finding

The CASCADE module expands the skill from local skill-quality review into external retrieval and documentation refresh. That scope expansion is risky because it introduces network-derived content and trust-boundary crossing without a strict necessity tied to the core optimization function, which can lead to poisoned updates or inappropriate ingestion into references.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The global rule says all modifications require user confirmation, yet the MUSE module auto-triggers after edits and appends tests to persistent files without a clear confirmation gate. This contradiction is dangerous because it creates a hidden post-edit write path that can store data or alter repository state even when the operator believes the system is in a confirmation-required mode.

Content

No source excerpt is available for this finding.

Ssd 3

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

Automatically appending interaction-derived test cases into tests.yaml can permanently embed sensitive user inputs or business data into repository files. This is especially risky because the workflow treats captured inputs as reusable regression assets, turning one conversation's private context into long-lived artifacts that may be shared or committed.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The usage table adds more broad trigger phrases that can cause the skill to activate from general conversation instead of a deliberate command. Given the document's mutation, git, and persistence behaviors, overbroad invocation substantially raises operational risk compared with a purely advisory skill.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
86% confidence
Finding

The overall activation logic maps broad user intents directly to internal modules without defining clear scope checks, target validation, or exclusion criteria. An ambiguous activation boundary increases the chance that ordinary troubleshooting or editing requests are misclassified as skill-evolution tasks, which can expose internal workflow behavior or trigger inappropriate analysis and modification suggestions.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The dispatch table uses short, generic phrases such as “检查技能”, “这个技能有问题”, and “精简技能” as activation triggers. These overlap with normal conversational requests and can cause the skill to activate on unrelated inputs, leading to prompt-routing confusion and unintended execution of maintenance or analysis behaviors on the wrong target.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The file’s user-facing instructions and trigger descriptions are entirely in Chinese, which imposes a specific language on users without stating that the skill is Chinese-only or offering an alternative. The policy scope explicitly includes language or locale constraints, and no opt-in or justification is provided here.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.