Back to skill

Security audit

Audit Evolution

Security checks across malware telemetry and agentic risk

Overview

This is a disclosed agent self-audit tool, but it needs Review because it can persist automatic agent-routing behavior and let a very short command lead to local patching.

Install only if you want a persistent self-audit loop for your agent. Review the AGENTS.md block and .audit-evolution hooks before enabling them, avoid --force unless you are comfortable overwriting an existing audit-evolution skill directory, and require an explicit file-by-file diff approval before allowing “进化” to modify skills, configs, or memory files.

SkillSpector

By NVIDIA
Vulnerability Patterns
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (21)

Intent-Code Divergence

Medium
Confidence
89% confidence
Finding
The document claims the system will first audit and propose suggestions without directly modifying the system, but elsewhere says it may 'apply minimal local patches' during the evolution flow. That ambiguity can cause users or downstream agents to misunderstand when state-changing actions are permitted, increasing the risk of unauthorized local modifications before clear approval boundaries are enforced.

Intent-Code Divergence

Medium
Confidence
91% confidence
Finding
The document establishes a safety boundary that modifications require explicit human approval, but later says that if the user replies only “进化”, the agent may request or apply a local patch. That creates an ambiguous approval model where a broad one-word command can be interpreted as authorization for state-changing actions, increasing the risk of unintended self-modification or policy bypass.

Vague Triggers

Medium
Confidence
84% confidence
Finding
The trigger phrases for Codex are overly broad and can overlap with normal user conversation, which may cause the audit skill to activate unexpectedly. In an agent environment, ambiguous routing can be abused to force unintended skill execution, disrupt workflows, or make the agent follow a different control path than the user intended.

Vague Triggers

Medium
Confidence
86% confidence
Finding
The general activation rule is broad and ambiguous, covering many common operational states like task completion, reading more than five files, or context size thresholds. This can cause frequent automatic invocation of the skill, creating opportunities for prompt-routing manipulation, denial of normal task flow, and unintended persistence or logging behavior across routine agent actions.

Vague Triggers

Medium
Confidence
93% confidence
Finding
The activation phrase '进化' is extremely generic and can plausibly appear in ordinary conversation, especially in discussions about product improvement or agent behavior. In an agent environment, ambiguous triggers can cause unintended skill activation, leading to confusing state changes, unexpected workflows, or execution of actions the user did not clearly request.

Missing User Warnings

Medium
Confidence
89% confidence
Finding
The document explicitly promotes broad automatic triggering across many conditions ('benchmark、用户纠错、超时、上下文压力、skill 修改后都能触发') but does not pair that behavior with clear user-facing warnings, consent, or guardrails. In a skill that records evidence, memory, patches, and next-run instructions, broad trigger conditions can lead to unexpected state capture or workflow changes, making the behavior more risky than a generic design note.

Vague Triggers

Medium
Confidence
95% confidence
Finding
The trigger phrase "进化" is a single common word with broad conversational meaning, so it can be invoked accidentally in normal dialogue. In an agent skill context, ambiguous activation can cause the agent to enter audit/evolution workflows unintentionally and begin reading files or preparing changes without a clearly scoped user request.

Vague Triggers

Medium
Confidence
89% confidence
Finding
The fallback invocation text is long, natural-language, and action-oriented, but it does not define a strict trigger boundary or structured parameters. This increases the chance that similar phrasing in ordinary discussion, copied text, or contextual content could activate the skill and prompt broad file discovery across accessible data.

Vague Triggers

Medium
Confidence
88% confidence
Finding
Using the single word trigger “进化” is overly broad and highly likely to appear in ordinary conversation, especially in a workflow about iterative improvement. That ambiguity can cause unintended skill activation, leading the agent to begin evidence collection, hook-driven behavior, or self-modification proposals when the user did not intend to invoke this skill.

Vague Triggers

Low
Confidence
80% confidence
Finding
The phrase “开始调用 Audit Evolution” is presented as a trigger but without clear scope, mode, or authorization boundaries. In agent systems that parse natural language loosely, this can cause accidental activation and unexpected reading of files or generation of follow-on actions beyond the user’s precise intent.

Vague Triggers

Medium
Confidence
95% confidence
Finding
The automatic trigger rules are very broad, including common events such as user corrections, reading more than 5 files, uncertainty language, or context pressure. This creates a substantial risk of frequent unintended activation, recursive behavior, unnecessary context expansion, and agent actions being rerouted into this skill without a deliberate user request.

Vague Triggers

Medium
Confidence
92% confidence
Finding
The skill mandates automatic invocation on broad events such as user corrections, uncertainty language, context pressure, and reading more than 5 files. In an agent runtime, this can cause unprompted workflow changes, extra file access, and repeated self-audit behavior that disrupts the primary task or creates cascading prompt-control effects without clear user intent.

Natural-Language Policy Violations

Medium
Confidence
84% confidence
Finding
The skill is written to require Chinese prompts, menus, and output structures without offering language negotiation. In a mixed-language environment, this can mislead users, reduce operator comprehension of approvals and safety boundaries, and increase the chance of accidental consent or misuse of the workflow.

Vague Triggers

Medium
Confidence
92% confidence
Finding
The auto-trigger section includes broad conditions such as benchmark completion, user corrections, uncertainty language, reading more than five files, or context usage above 60%. These are common events in normal agent operation, so the skill may activate unexpectedly and start audit or patch workflows without the user intending to invoke them.

Vague Triggers

Medium
Confidence
90% confidence
Finding
The short command menu uses generic terms like '进化', '保存', '暂停', '跑分', '继续', and '详情', which can easily appear in ordinary dialogue. If the agent treats these as control commands without stronger framing, normal user language could unintentionally trigger auditing, benchmarking, or state transitions.

Vague Triggers

Medium
Confidence
89% confidence
Finding
The skill enables implicit invocation globally (`allow_implicit_invocation: true`) without any visible trigger constraints, exclusions, or user-confirmation guardrails. Because this skill performs auditing, memory generation, and evolution recommendations over prior agent runs, unintended auto-invocation could expose sensitive run history or cause the agent to process internal operational data in contexts where the user did not explicitly request it.

Natural-Language Policy Violations

Medium
Confidence
78% confidence
Finding
The interface text and default prompt are Chinese-only, which can cause the skill to activate or operate in a language the user did not choose. This is primarily a safety and usability issue: users may not understand what the skill will do, reducing informed consent and increasing the chance that auditing or memory-related actions are invoked without clear comprehension.

Missing User Warnings

Medium
Confidence
88% confidence
Finding
The page instructs users to run installation scripts with force/override flags (`-Force` / `--force`) that write into a target workspace, update routing (`AGENTS.md`), and generate hooks, but it does not present any meaningful warning, confirmation step, or scope limitation. In an agent-skill context, this is risky because it encourages unattended modification of agent behavior and persistence mechanisms, which can broaden the skill's control surface without clear user review.

Vague Triggers

Medium
Confidence
84% confidence
Finding
The invocation phrase is extremely broad ('开始调用 Audit Evolution') and paired with instructions to automatically search current context and accessible files. In an agent environment, broad triggers increase the chance of accidental activation during normal conversation, causing unintended audits, file reads, or preparation for later workspace changes.

Vague Triggers

Medium
Confidence
91% confidence
Finding
The skill advertises automatic triggering on vague conditions such as benchmark completion, task failure, or context exceeding a threshold, and also states that installation will update routing and generate hooks. In context, this creates a persistence-like mechanism where the skill may run repeatedly without a tightly bounded policy, increasing the risk of unintended execution, excessive file access, and workflow hijacking.

Missing User Warnings

Medium
Confidence
88% confidence
Finding
The script recursively deletes and overwrites paths inside the target workspace during installation, and `--force` suppresses the only safeguard. If a user points the installer at the wrong workspace or an unexpected preexisting skill directory, important files or directories under `skills/audit-evolution` can be destroyed without backup or detailed confirmation.

VirusTotal

64/64 vendors flagged this skill as clean.

View on VirusTotal

Static analysis

No suspicious patterns detected.