T09 · Insecure Skill Coding Practices
- Location
src/gep/validator/sandboxExecutor.js:243- Finding
Hub-Provided Validation Commands Execute Without OS-Level Sandbox Isolation
- Content
View full analysis
- Remediation
View remediation
Security audit
Security checks for vulnerabilities and agentic risk
This appears to be a real Evolver tool, but it needs Review because it enables persistent networked automation and default remote validation code execution that is not safely scoped or fully disclosed.
Review before installing. Only use this skill if you are comfortable with a global CLI that can install persistent agent hooks, read/write local evolution state, contact EvoMap/GitHub, and run Hub-assigned validator jobs. Set EVOLVER_VALIDATOR_ENABLED=false unless you explicitly want remote validation work, disable telemetry if unwanted, and avoid running it in a workspace or user account that has sensitive files or credentials.
src/gep/validator/sandboxExecutor.js:243Hub-Provided Validation Commands Execute Without OS-Level Sandbox Isolation
src/gep/validator/index.js:25Remote Validator Role Is Enabled by Default and Can Be Controlled Through Persisted Feature State
src/config.js:228Default-On Anti-Abuse Telemetry Is Not Declared in the Skill Configuration
src/evolve.js:1Security-Critical Hub and Evolution Modules Are Distributed as Obfuscated JavaScript
index.js:2345Direct Authenticated Hub Request Contradicts the Declared Proxy-Only Security Boundary
The 'does not do' list says Evolver does not execute arbitrary shell commands. Later, the security model explicitly states that src/gep/solidify.js executes Gene validation commands and that distributed validator mode runs proposer-declared validation commands in a sandbox. This is an active contradiction in the documentation, not mere omission.
YARA rule matched a known malware signature (reverse shell, backdoor, ransomware, C2 framework, or info stealer).
变更范围计算和固化(solidify)。在非 git 目录中运行会直接报错并退出。
npm install -g @evomap/evolver
此命令将全局安装 evolver CLI。通过 evolver --help 验证。
如在 Linux/macOS 上遇到 EACCES 错误,建议配置用户级 prefix,而不是使用 sudo:
npm config set prefix ~/.npm-global
echo 'export PATH="$HOME/.npm-global/bin:$PATH"' >> ~/.bashrc
source ~/.bashrc
Evolver 通过 setup-hooks 命令与主流 Agent 运行时集成。每个需要接入的平台执行一次即可。
evolver setup-hooks --platform=cursor
会写入 ~/.cursor/hooks.json,并将 hook 脚本安装到 ~/.cursor/hooks/。重启 Cursor(或开新会话)后生效。钩子在 sessionStart、afterFileEdit、stop 时触发。
evolver setup-hooks --platform=claude-code
通过 ~/.claude/ 向 Claude Code 的 hook
YARA rule matched a known malware signature (reverse shell, backdoor, ransomware, C2 framework, or info stealer).
变更范围计算和固化(solidify)。在非 git 目录中运行会直接报错并退出。
npm install -g @evomap/evolver
此命令将全局安装 evolver CLI。通过 evolver --help 验证。
如在 Linux/macOS 上遇到 EACCES 错误,建议配置用户级 prefix,而不是使用 sudo:
npm config set prefix ~/.npm-global
echo 'export PATH="$HOME/.npm-global/bin:$PATH"' >> ~/.bashrc
source ~/.bashrc
Evolver 通过 setup-hooks 命令与主流 Agent 运行时集成。每个需要接入的平台执行一次即可。
evolver setup-hooks --platform=cursor
会写入 ~/.cursor/hooks.json,并将 hook 脚本安装到 ~/.cursor/hooks/。重启 Cursor(或开新会话)后生效。钩子在 sessionStart、afterFileEdit、stop 时触发。
evolver setup-hooks --platform=claude-code
通过 ~/.claude/ 向 Claude Code 的 hook
YARA rule matched a known malware signature (reverse shell, backdoor, ransomware, C2 framework, or info stealer).
变更范围计算和固化(solidify)。在非 git 目录中运行会直接报错并退出。
npm install -g @evomap/evolver
此命令将全局安装 evolver CLI。通过 evolver --help 验证。
如在 Linux/macOS 上遇到 EACCES 错误,建议配置用户级 prefix,而不是使用 sudo:
npm config set prefix ~/.npm-global
echo 'export PATH="$HOME/.npm-global/bin:$PATH"' >> ~/.bashrc
source ~/.bashrc
Evolver 通过 setup-hooks 命令与主流 Agent 运行时集成。每个需要接入的平台执行一次即可。
evolver setup-hooks --platform=cursor
会写入 ~/.cursor/hooks.json,并将 hook 脚本安装到 ~/.cursor/hooks/。重启 Cursor(或开新会话)后生效。钩子在 sessionStart、afterFileEdit、stop 时触发。
evolver setup-hooks --platform=claude-code
通过 ~/.claude/ 向 Claude Code 的 hook
YARA rule matched a known malware signature (reverse shell, backdoor, ransomware, C2 framework, or info stealer).
变更范围计算和固化(solidify)。在非 git 目录中运行会直接报错并退出。
npm install -g @evomap/evolver
此命令将全局安装 evolver CLI。通过 evolver --help 验证。
如在 Linux/macOS 上遇到 EACCES 错误,建议配置用户级 prefix,而不是使用 sudo:
npm config set prefix ~/.npm-global
echo 'export PATH="$HOME/.npm-global/bin:$PATH"' >> ~/.bashrc
source ~/.bashrc
Evolver 通过 setup-hooks 命令与主流 Agent 运行时集成。每个需要接入的平台执行一次即可。
evolver setup-hooks --platform=cursor
会写入 ~/.cursor/hooks.json,并将 hook 脚本安装到 ~/.cursor/hooks/。重启 Cursor(或开新会话)后生效。钩子在 sessionStart、afterFileEdit、stop 时触发。
evolver setup-hooks --platform=claude-code
通过 ~/.claude/ 向 Claude Code 的 hook
The declared description captures only a subset of the code's behavior. While the code does include an evolution loop and hub/proxy-related functionality consistent with a self-evolution engine, this chunk is actually the main operational entrypoint for a large multi-function daemon/CLI. It includes bootstrap update recovery, singleton locking, process lifecycle management, token credential helper behavior, git review/rollback/solidify flows, remote skill download, and ATP/validator/background services. Most notably, the proxy-token path exposes a local credential helper that reads and outputs a proxy token, which is a sensitive capability not suggested by the description. The code also performs substantial shell/git and filesystem actions beyond merely analyzing runtime history and applying constrained evolution. Therefore the description materially understates and misrepresents the skill's actual capabilities and primary scope.
The declared description presents a narrow primary purpose: an AI self-evolution engine that analyzes runtime history and evolves protocols, communicating with EvoMap Hub through a local proxy mailbox. The actual code chunk instead exposes a broad multi-command CLI for marketplace, asset management, authentication, web UI, hook management, secret reset, trajectory export, recipe handling, ATP transaction flows, and experiment execution. Some pieces are adjacent to an evolution platform, but the implemented behavior is materially broader and operationally different from the declared purpose. In particular, the code performs substantial account/auth, local system configuration, file export/import, web serving, and transaction-related actions that are not conveyed by the description. The declared network and shell permissions are not themselves problematic, but the description understates the true scope and capabilities of the code.
The code does not implement a self-evolution engine. It reads preexisting genes, capsules, and events from storage, filters them for A2A export, formats them as JSON or protocol messages, and optionally sends those messages through a transport. While it does communicate outward via a transport/mailbox-like mechanism, that communication is in service of exporting/publishing assets. There is no logic for analyzing runtime history to derive improvements, selecting or applying evolutionary changes, or modifying agent protocols. Thus the declared description materially overstates and mischaracterizes the code's primary purpose and capabilities.
The code chunk is a focused ingestion utility for external A2A data. It reads text input, parses candidate objects, filters allowed assets, verifies content-hash-based asset IDs, reduces confidence for external sources, stores staged candidates, records them in a memory graph, and optionally sends reject/quarantine decisions. This is materially different from the declared purpose of a self-evolution engine that analyzes runtime history and applies constrained evolution. While some persistence and transport behavior could be supporting infrastructure, the primary behavior shown is intake and quarantine of external assets, not agent self-improvement.
The code’s primary function is a command-line asset promotion script, not an engine that analyzes agent runtime history and evolves the agent. It requires explicit --type, --id, and --validated inputs, loads recent external candidates, performs limited safety checks on Gene validation commands, marks the chosen asset as promoted, stores it locally, and may optionally send decision messages. While the description mentions protocol-constrained evolution and communication with a hub, the actual code does not analyze history, generate improvements, or perform evolution decisions beyond promoting a preexisting candidate. The optional A2A transport messaging is secondary and much narrower than the declared purpose.
The code is a local analysis/report generator. It reads evolution_history_full.md, extracts interesting entries, groups them by skill, and writes evolution_detailed_report.md. While this loosely relates to analyzing history, it does not identify improvements in any decision-making sense beyond simple keyword filtering, does not perform protocol-constrained evolution, and does not communicate with any hub, proxy, mailbox, network, or shell. The primary purpose described is materially broader and different from the actual behavior shown.
The supplied code is a packaging/build script, not an AI self-evolution engine. Its behavior is narrowly focused on local filesystem and shell-based build operations: checking tool availability, bundling JavaScript with Bun, optionally obfuscating code, compiling platform-specific executables, generating SHA256 checksum files, and smoke-testing the host binary. There is no logic for analyzing runtime history, modifying agent protocols, evolving behavior, or communicating with any EvoMap Hub or local Proxy mailbox. While the declared permissions include shell and network, the script does not actually implement the claimed networked agent-evolution functionality. This is a material description-versus-behavior mismatch.
The declared description and the code behavior are materially unrelated. The code is a release-process utility that validates frozen CHANGELOG sections against git tags using local file reads and git commands. It does not analyze runtime history, evolve agents, apply protocols for evolution, or communicate with any hub/mailbox. While declared permissions include network and shell, only local shell/git and filesystem access are used; over-declared permissions alone are not the issue. The primary purpose is clearly different, so this is a strong mismatch.
The supplied code is a standalone log-extraction utility. It reads memory/mad_dog_evolution.log, parses 'Cycle Start' timestamps and Feishu-card command lines, converts them into markdown, and writes evolution_history.md. It does not analyze runtime history for improvements in any substantive sense beyond extracting logged text, does not apply any evolution changes, and does not communicate with any hub or mailbox. Therefore the actual behavior is materially narrower and different from the declared purpose.
The declared description presents a sophisticated self-evolution engine with analysis, decision-making, and external/local mailbox communication. The supplied code does none of that. It simply invokes git log, filters commit messages by the keyword "Evolution", reformats the results, and saves them as a markdown report. While this could be tangentially related to documenting evolution history, its primary purpose and capabilities are materially narrower and different from the declared functionality. There is no evidence of runtime-history processing, applying changes/evolution, or network/proxy mailbox communication.
The declared description presents an active self-evolution component with analysis, improvement application, and hub communication. The supplied code only reads local JSON/JSONL files, computes aggregate metrics for personality states from EvolutionEvent records, and prints a summary report. There is no code to modify agent state, enact protocol-constrained evolution, send messages, access a mailbox, or use network/shell capabilities. This is a materially different primary purpose: diagnostic/report generation rather than agent self-evolution.
The declared description presents this skill as an active self-evolution engine with analysis, improvement application, and hub communication. The supplied code instead implements a report generator: it reads evolution_history_full.md, parses entries, classifies them into categories/components using keyword matching, builds a markdown summary, and writes evolution_human_summary.md. There is no evidence of agent evolution, runtime decision-making, protocol-constrained updates, shell usage, network access, or mailbox/proxy communication. This is a materially different primary purpose, so it should be flagged as a mismatch.
The code is a standalone Node.js report generator for recall_verify events. It parses CLI args, reads local memory graph events via tryReadMemoryGraphEvents, computes success/mismatch statistics and percentiles, prints Markdown/JSON, and returns exit code 0 or 2 based on thresholds. It does not analyze runtime history for improvement opportunities in the sense of evolving an agent, does not apply any changes or protocol-constrained evolution, and does not communicate with EvoMap Hub or a local Proxy mailbox. The declared description therefore materially misrepresents both the primary purpose and the capabilities of this code.
The declared description and the actual code behavior are materially unrelated. The description claims an AI self-evolution component involving runtime analysis, agent improvement logic, and mailbox-based hub communication. The supplied code instead performs a narrow release/maintenance task: it reads README markdown files, fetches GitHub stargazer counts, formats the number, and rewrites static shields.io badge text. While the declared permissions include network and shell, which the script does use through HTTPS and the gh CLI, those are over-declared/supporting details rather than evidence of purpose alignment. The primary purpose, accessed resources, and functionality do not match the declared description.
The declared description describes an autonomous self-improvement system for AI agents. The actual code does not analyze runtime history, evolve agents, or enforce any evolution protocol. Instead, it bootstraps a merchant agent with three predefined service listings and handles incoming orders with simple keyword-based canned outputs. It communicates with a hub URL through a merchant agent, not via a local Proxy mailbox as described. This is a clear material mismatch in primary purpose and capabilities.
The code is a command-line utility named skill2recipes. It parses flags like --manifest, --title, --price, and --no-publish, loads a JSON manifest or skill paths from disk, and calls composeRecipeFromSkills to build and optionally publish a recipe to EvoMap. It sets A2A_TRANSPORT='http' and defaults A2A_HUB_URL to https://evomap.ai, indicating direct network communication to the Hub. There is nothing in this code about analyzing runtime history, identifying agent improvements, or applying self-evolution constraints. The primary purpose is materially different from the declared description, so this is a clear mismatch.
The code is a release/versioning utility, not an AI self-evolution engine. Its primary function is to inspect git commit messages and current package version, determine a semver bump (major/minor/patch), and persist that suggestion locally. While it uses shell access for git commands and filesystem writes, these are in service of version suggestion only. There is no network use, no mailbox/hub communication, no runtime-history analysis of an agent, and no application of behavioral or protocol changes. This is a clear material mismatch in primary purpose and declared capabilities.
The declared description presents a sophisticated AI self-evolution component with analysis, improvement, and local hub communication capabilities. The supplied code does none of that. It is a simple developer utility script for validating JavaScript modules by requiring them and checking their exports. There is no evidence of runtime-history processing, evolution logic, network or mailbox communication, or any AI-agent behavior. The primary purpose is materially different, so this is a clear mismatch.
The supplied code does not implement an AI self-evolution engine or any of the described behaviors. It simply validates and runs JavaScript test files using node --test, manages environment variables for the test process, and reports test results. There is no network activity, no mailbox communication, no analysis of runtime history, and no application of agent evolution logic. The primary purpose is materially different from the declared description, so this is a clear mismatch.
The declared description presents the skill as an autonomous self-evolution engine with analysis and hub communication responsibilities. The supplied code chunk instead implements a platform adapter for Claude Code: it builds hook definitions, writes settings, copies scripts, injects/removes a markdown section, and uninstalls those changes. While these hooks may support a larger evolver system, this chunk’s actual behavior is installation/configuration management, not the declared core behavior. The declared triggers are also inconsistent with the concrete hook events configured by the code.
Detected: suspicious.dangerous_exec, suspicious.dynamic_code_execution, suspicious.env_credential_access (+4 more)