Back to skill

Security audit

Capability Evolver 1.40.0

Security checks for vulnerabilities and agentic risk

Overview

This is a powerful self-evolution tool with disclosed shell and network use, but it needs Review because it can spawn code-changing agents, ingest remote instructions, send local context to external services, and change installed agent behavior with weak scoping.

Install only in a sandboxed workspace with clean git state, no production secrets, and network access limited to trusted origins. Prefer review mode, disable bridge/loop/auto-publish/auto-issue/auto-update unless you explicitly need them, avoid generic npm/npx validation from remote Genes, and do not point A2A_HUB_URL or MEMORY_GRAPH_REMOTE_URL at untrusted endpoints.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (5)

T01 · Skill Instruction Hijacking

Error
Location
src/evolve.js:1978
Finding

Untrusted Hub Content Is Injected into an Executor Agent Prompt

Content
View full analysis
0) { hubLessons = fetchResult.relevant_lessons; console.log(`[LessonBank] Received ${hubLessons.length} lesson(s) from ecosystem.`); } ``` ```js const prompt = isDirectReuse ? buildReusePrompt({ capsule: hubHit.match, signals, nowIso: new Date().toISOString(), }) : buildGepPrompt({ nowIso: new Date().toISOString(), context, signals, selector, parentEventId: getLastEventId(), selectedGene, capsuleCandidates, genesPreview, capsulesPreview, capabilityCandidatesPreview, externalCandidatesPreview, hubMatchedBlock, strategyPolicy, failedCapsules: recentFailedCapsules, hubLessons, }); ``` ```js const executorTask = [ 'You are the executor (the Hand).', 'Your job is to apply a safe, minimal patch in this repo following the attached GEP protocol prompt.', artifact && artifact.promptPath ? `Prompt file: ${artifact.promptPath}` : 'Prompt file: (unavailable)', '', 'After applying changes and validations, you MUST run:', ' node index.js solidify', '', 'Loop chaining (only if you are running in loop mode): after solidify succeeds, print a sessions_spawn call to start the next loop run with a short delay.', 'Example:', 'sessions_spawn({ task: "exec: node skills/feishu-evolver-wrapper/lifecycle.js ensure", agentId: "main", cleanup: "delete", label: "gep_loop_next" })', '', 'GEP protocol prompt (may be truncated here; prefer the prompt file if provided):', clip(prompt, 24000), ].join('\n'); const spawn = renderSessionsSpaw ...[truncated 2516 chars]
Remediation
View remediation

T03 · Remote Payload Retrieval and Execution

Error
Location
src/gep/policyCheck.js:378
Finding

Validation Allowlist Permits Arbitrary npm and npx Package Execution

Content
View full analysis
c.startsWith(p))) return false; if (/`|\$\(/.test(c)) return false; const stripped = c.replace(/"[^"]*"/g, '').replace(/'[^']*'/g, ''); if (/[;&|><]/.test(stripped)) return false; if (/^node\s+(-e|--eval|--print|-p)\b/.test(c)) return false; return true; } function runValidationsOnce(gene, opts) { const repoRoot = opts.repoRoot || getRepoRoot(); const timeoutMs = Number.isFinite(Number(opts.timeoutMs)) ? Number(opts.timeoutMs) : 180000; const validation = Array.isArray(gene && gene.validation) ? gene.validation : []; const results = []; const startedAt = Date.now(); for (const cmd of validation) { const c = String(cmd || '').trim(); if (!c) continue; if (!isValidationCommandAllowed(c)) { results.push({ cmd: c, ok: false, out: '', err: 'BLOCKED: validation command rejected by safety check (allowed prefixes: node/npm/npx; shell operators prohibited)' }); return { ok: false, results, startedAt, finishedAt: Date.now() }; } const r = tryRunCmd(c, { cwd: repoRoot, timeoutMs }); results.push({ cmd: c, ok: r.ok, out: String(r.out || ''), err: String(r.err || '') }); if (!r.ok) return { ok: false, results, startedAt, finishedAt: Date.now() }; } return { ok: true, results, startedAt, finishedAt: Date.now() }; } ``` The accepted command is passed to a shell: ```js function runCmd(cmd, opts = {}) { const cwd = opts.cwd || getRepoRoot(); const timeoutMs = Number.isFinite(Number(opts.timeout ...[truncated 2399 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
src/gep/questionGenerator.js:91
Finding

User Conversation Excerpts Can Be Sent to the Hub Without Redaction

Content
View full analysis
0) { var featureContext = featureLines[0] .replace(/\s+/g, ' ') .trim() .slice(0, 150); candidates.push({ question: 'User requested a feature that may benefit from community solutions: ' + featureContext + ' -- Are there existing implementations or best practices for this?', amount: 0, signals: ['user_feature_request', 'community_solution_sought'], priority: 1, }); } } ``` The ge ...[truncated 1955 chars]
Remediation
View remediation

other

Warning
Location
src/gep/envFingerprint.js:46
Finding

Stable Hardware-Derived Device Fingerprint Is Transmitted to the Hub

Content
View full analysis
0) { const raw = os.hostname() + '|' + macs.join(','); return crypto.createHash('sha256') .update('evomap:' + raw) .digest('hex') .slice(0, 32); } return crypto.randomBytes(16).toString('hex'); } ``` The value is included directly in the environment fingerprint: ```js return { device_id: getDeviceId(), node_version: process.version, platform: process.platform, arch: process.arch, os_release: os.release(), hostname: crypto.createHash('sha256') .update(os.hostname()) .digest('hex') .slice(0, 12), evolver_version: pkgVersion, client: pkgName || 'evolver', client_version: pkgVersion, region: region, cwd: crypto.createHash('sha256') .update(process.cwd()) .digest('hex') .slice(0, 12), container: isContainer(), captured_at: new Date().toISOString(), }; ``` Hello messages include the fingerprint: ```js function buildHello(opts) { const o = opts || {}; return buildMessage({ messageType: 'hello', senderId: o.nodeId, payload: { capabilities: o.capabilities || {}, gene_count: typeof o.geneCount === 'number' ? o.geneCount : n ...[truncated 1802 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
src/gep/a2aProtocol.js:422
Finding

Configurable Hub URL Receives Persisted Bearer Credentials Without Transport or Origin Enforcement

Content
View full analysis
Remediation
View remediation
Vulnerability Patterns
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (204)

Context-Inappropriate Capability

Critical
Category
Not specified by scanner
Confidence
99% confidence
Finding

The code passes the contents of the environment variable INTEGRATION_STATUS_CMD directly to execSync, which executes through a shell. Any attacker who can influence environment variables, configuration, or process startup can achieve arbitrary command execution under the agent's privileges.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

The README asserts that the tool does not execute shell commands, but the same document later describes lifecycle operations and solidify.js validation commands that do execute commands under constraints. This contradiction can cause operators or integrating platforms to under-assess execution risk and deploy the skill with elevated trust or insufficient sandboxing.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The skill advertises a narrow self-evolution purpose, but the documented behavior spans remote hub communication, asset import/export, Git mutation, process management, packaging, issue reporting, and other operational capabilities. This mismatch is dangerous because operators may grant powerful network and shell permissions under a misleading mental model, increasing the chance of unsafe deployment and unexpected code/data changes.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The skill advertises a narrow self-evolution purpose, but the documented behavior spans remote hub communication, asset import/export, Git mutation, process management, packaging, issue reporting, and other operational capabilities. This mismatch is dangerous because operators may grant powerful network and shell permissions under a misleading mental model, increasing the chance of unsafe deployment and unexpected code/data changes.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The skill advertises a narrow self-evolution purpose, but the documented behavior spans remote hub communication, asset import/export, Git mutation, process management, packaging, issue reporting, and other operational capabilities. This mismatch is dangerous because operators may grant powerful network and shell permissions under a misleading mental model, increasing the chance of unsafe deployment and unexpected code/data changes.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
93% confidence
Finding

The skill advertises a narrow self-evolution purpose, but the documented behavior spans remote hub communication, asset import/export, Git mutation, process management, packaging, issue reporting, and other operational capabilities. This mismatch is dangerous because operators may grant powerful network and shell permissions under a misleading mental model, increasing the chance of unsafe deployment and unexpected code/data changes.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The skill advertises a narrow self-evolution purpose, but the documented behavior spans remote hub communication, asset import/export, Git mutation, process management, packaging, issue reporting, and other operational capabilities. This mismatch is dangerous because operators may grant powerful network and shell permissions under a misleading mental model, increasing the chance of unsafe deployment and unexpected code/data changes.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The skill advertises a narrow self-evolution purpose, but the documented behavior spans remote hub communication, asset import/export, Git mutation, process management, packaging, issue reporting, and other operational capabilities. This mismatch is dangerous because operators may grant powerful network and shell permissions under a misleading mental model, increasing the chance of unsafe deployment and unexpected code/data changes.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The skill advertises a narrow self-evolution purpose, but the documented behavior spans remote hub communication, asset import/export, Git mutation, process management, packaging, issue reporting, and other operational capabilities. This mismatch is dangerous because operators may grant powerful network and shell permissions under a misleading mental model, increasing the chance of unsafe deployment and unexpected code/data changes.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The skill advertises a narrow self-evolution purpose, but the documented behavior spans remote hub communication, asset import/export, Git mutation, process management, packaging, issue reporting, and other operational capabilities. This mismatch is dangerous because operators may grant powerful network and shell permissions under a misleading mental model, increasing the chance of unsafe deployment and unexpected code/data changes.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The skill advertises a narrow self-evolution purpose, but the documented behavior spans remote hub communication, asset import/export, Git mutation, process management, packaging, issue reporting, and other operational capabilities. This mismatch is dangerous because operators may grant powerful network and shell permissions under a misleading mental model, increasing the chance of unsafe deployment and unexpected code/data changes.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The skill advertises a narrow self-evolution purpose, but the documented behavior spans remote hub communication, asset import/export, Git mutation, process management, packaging, issue reporting, and other operational capabilities. This mismatch is dangerous because operators may grant powerful network and shell permissions under a misleading mental model, increasing the chance of unsafe deployment and unexpected code/data changes.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
94% confidence
Finding

The skill advertises a narrow self-evolution purpose, but the documented behavior spans remote hub communication, asset import/export, Git mutation, process management, packaging, issue reporting, and other operational capabilities. This mismatch is dangerous because operators may grant powerful network and shell permissions under a misleading mental model, increasing the chance of unsafe deployment and unexpected code/data changes.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The skill advertises a narrow self-evolution purpose, but the documented behavior spans remote hub communication, asset import/export, Git mutation, process management, packaging, issue reporting, and other operational capabilities. This mismatch is dangerous because operators may grant powerful network and shell permissions under a misleading mental model, increasing the chance of unsafe deployment and unexpected code/data changes.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The skill advertises a narrow self-evolution purpose, but the documented behavior spans remote hub communication, asset import/export, Git mutation, process management, packaging, issue reporting, and other operational capabilities. This mismatch is dangerous because operators may grant powerful network and shell permissions under a misleading mental model, increasing the chance of unsafe deployment and unexpected code/data changes.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The skill advertises a narrow self-evolution purpose, but the documented behavior spans remote hub communication, asset import/export, Git mutation, process management, packaging, issue reporting, and other operational capabilities. This mismatch is dangerous because operators may grant powerful network and shell permissions under a misleading mental model, increasing the chance of unsafe deployment and unexpected code/data changes.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The skill advertises a narrow self-evolution purpose, but the documented behavior spans remote hub communication, asset import/export, Git mutation, process management, packaging, issue reporting, and other operational capabilities. This mismatch is dangerous because operators may grant powerful network and shell permissions under a misleading mental model, increasing the chance of unsafe deployment and unexpected code/data changes.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The skill advertises a narrow self-evolution purpose, but the documented behavior spans remote hub communication, asset import/export, Git mutation, process management, packaging, issue reporting, and other operational capabilities. This mismatch is dangerous because operators may grant powerful network and shell permissions under a misleading mental model, increasing the chance of unsafe deployment and unexpected code/data changes.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The skill advertises a narrow self-evolution purpose, but the documented behavior spans remote hub communication, asset import/export, Git mutation, process management, packaging, issue reporting, and other operational capabilities. This mismatch is dangerous because operators may grant powerful network and shell permissions under a misleading mental model, increasing the chance of unsafe deployment and unexpected code/data changes.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The skill advertises a narrow self-evolution purpose, but the documented behavior spans remote hub communication, asset import/export, Git mutation, process management, packaging, issue reporting, and other operational capabilities. This mismatch is dangerous because operators may grant powerful network and shell permissions under a misleading mental model, increasing the chance of unsafe deployment and unexpected code/data changes.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The skill advertises a narrow self-evolution purpose, but the documented behavior spans remote hub communication, asset import/export, Git mutation, process management, packaging, issue reporting, and other operational capabilities. This mismatch is dangerous because operators may grant powerful network and shell permissions under a misleading mental model, increasing the chance of unsafe deployment and unexpected code/data changes.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The skill advertises a narrow self-evolution purpose, but the documented behavior spans remote hub communication, asset import/export, Git mutation, process management, packaging, issue reporting, and other operational capabilities. This mismatch is dangerous because operators may grant powerful network and shell permissions under a misleading mental model, increasing the chance of unsafe deployment and unexpected code/data changes.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

The skill advertises a narrow self-evolution purpose, but the documented behavior spans remote hub communication, asset import/export, Git mutation, process management, packaging, issue reporting, and other operational capabilities. This mismatch is dangerous because operators may grant powerful network and shell permissions under a misleading mental model, increasing the chance of unsafe deployment and unexpected code/data changes.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

The skill advertises a narrow self-evolution purpose, but the documented behavior spans remote hub communication, asset import/export, Git mutation, process management, packaging, issue reporting, and other operational capabilities. This mismatch is dangerous because operators may grant powerful network and shell permissions under a misleading mental model, increasing the chance of unsafe deployment and unexpected code/data changes.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The skill advertises a narrow self-evolution purpose, but the documented behavior spans remote hub communication, asset import/export, Git mutation, process management, packaging, issue reporting, and other operational capabilities. This mismatch is dangerous because operators may grant powerful network and shell permissions under a misleading mental model, increasing the chance of unsafe deployment and unexpected code/data changes.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.dangerous_exec, suspicious.env_credential_access, suspicious.exposed_secret_literal (+1 more)

Shell command execution detected (child_process).

Critical
Code
suspicious.dangerous_exec
Location
index.js:242

Shell command execution detected (child_process).

Critical
Code
suspicious.dangerous_exec
Location
scripts/build_public.js:170

Shell command execution detected (child_process).

Critical
Code
suspicious.dangerous_exec
Location
scripts/generate_history.js:17

Shell command execution detected (child_process).

Critical
Code
suspicious.dangerous_exec
Location
scripts/publish_public.js:13

Shell command execution detected (child_process).

Critical
Code
suspicious.dangerous_exec
Location
scripts/recover_loop.js:19

Shell command execution detected (child_process).

Critical
Code
suspicious.dangerous_exec
Location
scripts/suggest_version.js:27

Shell command execution detected (child_process).

Critical
Code
suspicious.dangerous_exec
Location
scripts/validate-suite.js:19

Shell command execution detected (child_process).

Critical
Code
suspicious.dangerous_exec
Location
src/evolve.js:485

Shell command execution detected (child_process).

Critical
Code
suspicious.dangerous_exec
Location
src/gep/gitOps.js:12

Shell command execution detected (child_process).

Critical
Code
suspicious.dangerous_exec
Location
src/gep/idleScheduler.js:39

Shell command execution detected (child_process).

Critical
Code
suspicious.dangerous_exec
Location
src/gep/llmReview.js:70

Shell command execution detected (child_process).

Critical
Code
suspicious.dangerous_exec
Location
src/ops/health_check.js:20

Shell command execution detected (child_process).

Critical
Code
suspicious.dangerous_exec
Location
src/ops/lifecycle.js:27

Shell command execution detected (child_process).

Critical
Code
suspicious.dangerous_exec
Location
src/ops/self_repair.js:17

Shell command execution detected (child_process).

Critical
Code
suspicious.dangerous_exec
Location
src/ops/skills_monitor.js:96

Shell command execution detected (child_process).

Critical
Code
suspicious.dangerous_exec
Location
test/bridge.test.js:98

Shell command execution detected (child_process).

Critical
Code
suspicious.dangerous_exec
Location
test/loopMode.test.js:129

Environment variable access combined with network send.

Critical
Code
suspicious.env_credential_access
Location
index.js:109

Environment variable access combined with network send.

Critical
Code
suspicious.env_credential_access
Location
src/gep/a2aProtocol.js:75

Environment variable access combined with network send.

Critical
Code
suspicious.env_credential_access
Location
src/gep/hubReview.js:104

Environment variable access combined with network send.

Critical
Code
suspicious.env_credential_access
Location
src/gep/hubSearch.js:75

Environment variable access combined with network send.

Critical
Code
suspicious.env_credential_access
Location
src/gep/issueReporter.js:21

Environment variable access combined with network send.

Critical
Code
suspicious.env_credential_access
Location
src/gep/memoryGraphAdapter.js:77

Environment variable access combined with network send.

Critical
Code
suspicious.env_credential_access
Location
src/gep/taskReceiver.js:11

File appears to expose a hardcoded API secret or token.

Critical
Code
suspicious.exposed_secret_literal
Location
test/sanitize.test.js:12

Sensitive-looking file read is paired with a network send.

Warn
Code
suspicious.potential_exfiltration
Location
src/gep/a2aProtocol.js:431