Back to skill

Security audit

Zhua Evolver

Security checks for vulnerabilities and agentic risk

Overview

This skill mostly contains harmless local scoring scripts, but its instructions authorize automatic skill installation and persistent identity changes without clear user approval or safety controls.

Review this carefully before installing. The packaged code is not overtly destructive, but the skill's instructions could lead an agent to install unreviewed skills or alter persistent identity files. Only use it with explicit approval before each install or persistent-state change, trusted/pinned skill sources, visible diffs, backups, and rollback controls.

Vulnerability Patterns
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T08 · Insecure Dependencies

Warning
Location
SKILL.md:31
Finding

Unreviewed Third-Party Skill Installation Workflow

Content
View full analysis
--limit 10 ``` ### 3. 自动安装 ```bash python3 scripts/install_skills.py --skills <技能列表> ``` ### 4. 进化记录 ```bash python3 scripts/log_evolution.py --cycle <周期号> --changes <变更描述> ``` ``` ### Technical Analysis The documented workflow instructs the agent to search a skill registry and automatically install selected third-party skills. It does not specify any of the following supply-chain controls: - An allowlist of trusted publishers or packages - Exact version pinning - Cryptographic signature or checksum verification - Source-code and instruction review before installation - Permission disclosure and least-privilege enforcement - Sandboxed execution - Explicit user approval for each installation A skill can contain executable scripts and agent-facing instructions. Installing an untrusted or compromised skill can therefore introduce malicious code, instruction hijacking, credential access, or other unauthorized behavior into the agent environment. The referenced `scripts/search_skills.py` and `scripts/install_skills.py` files are absent from the audited artifact. Consequently, the packaged Python code cannot currently perform this workflow directly. Exploitation requires either a future implementation of those scripts or an agent interpreting and carrying out the documented instructions through other available tools. ### Attack Path 1. An attacker publishes a malicious or typosquatted skill to the registry searched by the workflow. 2. The attacker selects metadata and keywords likely to match an agent capability-gap search. 3. The agent follows `SKILL.md` and searches for skills without restricting results to trusted publishers. 4. The malicious skill is selected and installed withou ...[truncated 971 chars]
Remediation
View remediation

T02 · Agent Memory Poisoning

Warning
Location
SKILL.md:55
Finding

Unrestricted Persistent Agent Identity Modification

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (16)

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

声明描述的是一个范围较大的自动化自我进化系统,包含分析、搜索、执行循环和日志记录等多个核心能力。但提供的代码只做了非常有限的能力差距计算:使用硬编码等级定义,根据当前技能数量生成建议,没有网络搜索、技能安装、循环执行、状态持久化或日志记录等功能,也没有体现更高智能水平提升的自动化机制。因此,代码实际行为与声明用途存在明显且实质性的不匹配。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

声明描述的是一个主动型自我进化系统,核心能力应包括发现不足、检索补强方案、实施进化步骤并记录过程。但提供的代码仅是一个本地评估脚本:使用硬编码的 features 和 metrics 与目标集合比较后输出分数和文案。它没有外部搜索、没有技能安装或补强、没有任何循环进化流程、没有日志持久化,也没有资源访问或自动触发逻辑。虽然输出中提到“进化系统”“持续自我进化”等概念,但这些只是静态文本与自评内容,不构成所声明功能的实现。因此声明与实际行为存在明显不匹配。

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description presents an active self-improvement system with automation around gap analysis, skill discovery, iterative evolution, and logging. The supplied code does none of that. It only defines fixed dictionaries of features/metrics, checks membership against another fixed reference set, prints a status report, and returns a computed percentage. There is no dynamic analysis, no external search, no modification of capabilities, no looped evolution process, and no log persistence. Therefore the code’s actual primary purpose is a local assessment/report generator, which materially differs from the declared self-evolution system.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The declared description promises an automated self-improvement system with gap analysis, skill discovery, iterative evolution, and logging. The supplied code only defines fixed feature/metric lists, checks which items are present in a hardcoded profile, prints a score, and returns the percentage. It does not perform dynamic analysis, external search, skill supplementation, iterative execution, or log persistence. The actual primary purpose is an assessment/reporting script, which materially differs from the declared self-evolution system.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The declared description promises an automated self-evolution system with multiple active capabilities: analyzing gaps, searching for reinforcement skills, running an evolution cycle, and recording logs. The supplied code only contains a constant definition of a target standard and a function that compares an input dictionary against predefined features/metrics, prints pass/fail markers, computes a percentage, and returns it. This is materially narrower and different in primary purpose: it is a scoring/checklist script, not an evolution engine. Therefore the description overstates what the code actually does.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

There is a clear description-behavior mismatch. The declared purpose describes a substantive autonomous self-improvement system, but the actual code chunk is a minimal example helper script with no implemented logic beyond printing a message. It does not analyze abilities, search for skills, persist logs, access resources, or perform any evolution-related behavior. This is not merely incomplete supporting code; its actual primary behavior is unrelated placeholder output rather than the claimed system functionality.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The declared description presents an automated self-evolution system with substantive functionality: analyzing gaps, searching for improvements, executing iterative enhancement, and recording logs. The supplied code does none of these actions. It only parses a --task argument, uses simple keyword checks to choose one or more predefined roles, and prints assignment/status messages. While some minion role labels loosely reference search, logging, safety, and evolution, these are descriptive text only and not actual implemented capabilities. Therefore the code's real primary purpose—CLI-based task dispatch/orchestration—materially differs from the declared self-evolution system.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The activation criteria are extremely broad: any desire for improvement, self-evolution, or higher intelligence can trigger the skill. In an agentic environment, vague triggers can cause unplanned invocation of workflows that search for and install new skills, expanding system behavior without clear operator intent.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The markdown advertises automatic skill installation and configuration but provides no warning, consent checkpoint, sandboxing requirement, or trust policy for downloaded components. In context, this is dangerous because a self-modifying agent workflow can introduce unreviewed code or capabilities into the environment, increasing the risk of supply-chain compromise, persistence, or unsafe privilege use.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

This code's user-facing natural language, including the module description, CLI description, argument help text, and output messages, is entirely in Chinese with no option to select another language or indication that the skill is intentionally limited to a Chinese-speaking context. That creates a language/locale policy concern under the rule for forced language without user opt-in.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The file’s user-facing natural language is entirely in Chinese, including the script description, role labels, status output, and CLI help text. This imposes a specific language on users without offering an opt-in, fallback, or documented locale-specific justification, which matches the language/locale policy violation criteria.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
86% confidence
Finding

The natural-language instructions and descriptions are presented entirely in Chinese, with no indication that the user can choose another language or that the skill is intentionally limited to a Chinese-language context. Under the policy, forcing a specific language without opt-in can be a locale-policy issue.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
92% confidence
Finding

The script’s natural-language docstrings and all user-facing print output are in Chinese, and there is no indication that the language is configurable or chosen by the user. This can violate a language/locale policy when a skill imposes a specific language without opt-in.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
95% confidence
Finding

This is a code file, so SQP-3 applies to natural-language strings embedded in docstrings and print statements. The file consistently forces one language for all descriptions and output, and there is no opt-in, fallback, or justification that the skill is intended only for a Chinese-speaking or region-specific context.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
95% confidence
Finding

The file’s docstrings, comments, labels, and user-facing print messages are all in Chinese, and the script provides no option to select another language. Under the stated policy, forcing a specific language without user opt-in is a natural-language policy violation unless the locale restriction is documented and justified.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
85% confidence
Finding

The module docstring and all user-facing text are written only in Chinese, with no indication that users can choose another language or locale. The policy specifically flags language or locale constraints when they are imposed without user opt-in.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.