Back to skill

Security audit

Ralph Evolver

Security checks for vulnerabilities and agentic risk

Overview

This skill is a repository self-improvement prompt generator, but it also runs project scripts and turns repository text into agent instructions without enough warning or control.

Install only if you intend to use it on trusted repositories or inside a sandbox. Expect it to read repository metadata and files, write .ralph state, print project-derived prompt content, and run project-controlled build/test/lint commands under your account.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T09 · Insecure Skill Coding Practices

Error
Location
evolve.js:585
Finding

Automatic Execution of Untrusted Project Scripts During Health Checks

Content
View full analysis
0) { health.errors.push({ type: 'test_failure', message: `Tests failed: ${health.testResults.failed} failures` }); } } } else { health.testResults = { passed: 0, failed: 0, total: 0, skipped: true }; } // 3. Lint check (if exists) if (pkg.scripts && pkg.scripts.lint) { try { execFileSync('npm', ['run', 'lint'], { cwd: this.projectPath, encoding: 'utf-8', timeout: 30000, stdio: ['pipe', 'pipe', 'pipe'] }); } catch (e) { const output = (e.stderr || e.stdout || e.message || ...[truncated 3085 chars]
Remediation
View remediation

T01 · Skill Instruction Hijacking

Error
Location
evolve.js:853
Finding

Repository-Controlled Content Is Embedded into Actionable Agent Instructions Without a Trust Boundary

Content
View full analysis
🔍 **Hypothesis prompt**: If most commits are "fix" or "patch", what root cause keeps producing bugs?\n` : ''} ${signals.todos ? `### TODO/FIXME (distress signals in the code)\n\`\`\`\n${signals.todos}\n\`\`\`\n> 🔍 **Hypothesis prompt**: Are these TODOs procrastination, or signs of incomplete design?\n` : ''} ${signals.errorPatterns ? `### Error Handling Patterns (where are the fragile points)\n\`\`\`\n${signals.errorPatterns}\n\`\`\`\n> 🔍 **Hypothesis prompt**: Is error handling consistent (good) or sca ...[truncated 4227 chars]
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Output HandlingUnvalidated Output Injection, Cross-Context Output, Unbounded Output
  • System Prompt LeakageDirect Leakage, Indirect Extraction, Tool-Based Exfiltration
  • Memory PoisoningPersistent Context Injection, Context Window Stuffing, Memory Manipulation
Findings (19)

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

The skill appears to expose CLI flags, filesystem read/write access, and persistent state management while being described only as a conceptual self-improvement tool. This is risky because hidden operational controls and persistence can be abused or accidentally triggered in sensitive repositories, and the mismatch reduces the chance that users will sandbox or restrict the skill appropriately.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
96% confidence
Finding

The skill appears to expose CLI flags, filesystem read/write access, and persistent state management while being described only as a conceptual self-improvement tool. This is risky because hidden operational controls and persistence can be abused or accidentally triggered in sensitive repositories, and the mismatch reduces the chance that users will sandbox or restrict the skill appropriately.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

The skill appears to expose CLI flags, filesystem read/write access, and persistent state management while being described only as a conceptual self-improvement tool. This is risky because hidden operational controls and persistence can be abused or accidentally triggered in sensitive repositories, and the mismatch reduces the chance that users will sandbox or restrict the skill appropriately.

Content

No source excerpt is available for this finding.

Unvalidated Output Injection

High
Category
Output Handling
Confidence
100% confidence
Finding

Model output is used without validation or sanitization. Unvalidated output injected into downstream contexts (SQL, shell, HTML) enables injection attacks and arbitrary code execution.

Content

Scanner excerpt · evolve.js (reported line 725)May include surrounding context.

js
const errors = [];
    const errorRegex = /(.+\.ts)\((\d+),(\d+)\): error TS\d+: (.+)/g;
    let match;
    while ((match = errorRegex.exec(output)) !== null) {
      errors.push({
        type: 'typescript',
        file: match[1],

Direct Prompt Extraction

High
Category
System Prompt Leakage
Confidence
85% confidence
Finding

Skill contains instructions that could directly expose system prompts, internal rules, or hidden instructions to users or external parties.

Content

Scanner excerpt · evolve.js (reported line 1226)May include surrounding context.

js
}

  // 4. Generate Prompt
  return promptGenerator.generate(projectPath, health, mutation, state, task);
}

module.exports = {

Direct Prompt Extraction

High
Category
System Prompt Leakage
Confidence
85% confidence
Finding

Skill contains instructions that could directly expose system prompts, internal rules, or hidden instructions to users or external parties.

Content

Scanner excerpt · index.js (reported line 8)May include surrounding context.

js
* Usage:
 *   node index.js --project <path>           # Single iteration
 *   node index.js --project <path> --loop 5  # Run 5 cycles
 *   node index.js --project <path> --spawn   # Output prompt for sessions_spawn
 */

const path = require('path');

Direct Prompt Extraction

High
Category
System Prompt Leakage
Confidence
85% confidence
Finding

Skill contains instructions that could directly expose system prompts, internal rules, or hidden instructions to users or external parties.

Content

Scanner excerpt · index.js (reported line 143)May include surrounding context.

js
* Usage:
 *   node index.js --project <path>           # Single iteration
 *   node index.js --project <path> --loop 5  # Run 5 cycles
 *   node index.js --project <path> --spawn   # Output prompt for sessions_spawn
 */

const path = require('path');

Memory Manipulation

High
Category
Memory Poisoning
Confidence
90% confidence
Finding

The --reset path recursively deletes projectPath/.ralph using fs.rmSync(..., { recursive: true }) with a user-controlled project path and no validation that the target is the expected application state directory. If projectPath is pointed at an unexpected location or a symlinked path, this can destroy arbitrary local data under that .ralph directory and is a real destructive file-operation risk.

Content

Scanner excerpt · index.js (reported line 158)May include surrounding context.

js
const { projectPath, loopCount, isSpawn, isReset, task } = parseArgs(args);

  // Reset state
  if (isReset) {
    const stateDir = path.join(projectPath, '.ralph');
    if (fs.existsSync(stateDir)) {

Direct Prompt Extraction

High
Category
System Prompt Leakage
Confidence
85% confidence
Finding

In spawn mode, the program prints the full generated evolution prompt between clear markers, making internal prompt content directly extractable to stdout for downstream consumption. If generateEvolutionPrompt incorporates sensitive repository content, hidden instructions, or secrets from the target project, this becomes a straightforward exfiltration channel into other agent sessions or logs.

Content

Scanner excerpt · index.js (reported line 190)May include surrounding context.

js
// Generate evolution prompt
    const evolutionPrompt = evolve.generateEvolutionPrompt(projectPath, state, task);

    // Spawn mode: output prompt for sessions_spawn
    if (isSpawn) {
      console.log('\n=== EVOLUTION_PROMPT_START ===\n');
      console.log(evolutionPrompt);

Direct Prompt Extraction

High
Category
System Prompt Leakage
Confidence
85% confidence
Finding

Normal mode also prints the generated evolution prompt directly to stdout, which can leak project-derived sensitive information to terminals, logs, CI systems, or calling tools. In a self-improvement/agent skill context, directly emitting generated prompts is more dangerous because prompts may contain instructions or data intended only for internal agent use.

Content

Scanner excerpt · index.js (reported line 199)May include surrounding context.

js
return;
    }

    // Normal mode: output prompt directly
    console.log('\n' + evolutionPrompt);
    const completed = updateLoopProgress(projectPath, state);
    if (!completed) {

Self-Modification

High
Category
Rogue Agent
Confidence
91% confidence
Finding

The manifest explicitly advertises a 'self-evolve' script that runs the package against the current project in a loop, indicating self-modification or recursive code rewriting behavior. In an agent skill context, that substantially raises the risk of uncontrolled file changes, persistence of unsafe logic, and amplification of harmful behavior if the underlying tool lacks strict path, review, and execution constraints.

Content

Scanner excerpt · package.json (reported line 12)May include surrounding context.

json
"scripts": {
    "start": "node index.js",
    "evolve": "node index.js --project .",
    "self-evolve": "node index.js --project . --loop 3",
    "test": "vitest run"
  },
  "keywords": [

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The description is broad and aspirational, with no clear trigger boundaries or invocation constraints, making it easier for an orchestrator to call the skill in contexts where repository mutation or deep analysis is inappropriate. Ambiguous invocation language increases the attack surface because a powerful skill may activate outside the narrow scenarios its operator intended.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The documentation describes improvement tracking and repository signal collection but omits a clear warning that repository history is recorded/analyzed and that the tool may modify its own state. This omission undermines informed consent and may expose sensitive development metadata or create unexpected persistent artifacts in the repository.

Content

No source excerpt is available for this finding.

Ae4

Medium
Category
analysis-evasion
Confidence
80% confidence
Finding

Suspicious Unicode normalization or mixed-script content

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The skill executes multiple local subprocesses for git/grep/npm/python/pytest analysis without any user-facing confirmation or safety gate. Although commands are invoked via execFileSync rather than a shell, they can still trigger untrusted project hooks, build scripts, tests, or tooling with side effects when run against an attacker-controlled repository.

Content

No source excerpt is available for this finding.

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
89% confidence
Finding

The code persists history into .ralph/improvements.json inside the target project and the generated prompt instructs further writes, but does not clearly warn users that project files will be modified. Silent workspace modification is risky in automation contexts because it can taint repositories, affect CI, or overwrite user expectations about read-only analysis.

Content

No source excerpt is available for this finding.

Vague Triggers

Low
Category
Not specified by scanner
Confidence
83% confidence
Finding

This is a manifest file, so vague-trigger analysis applies. The description 'Recursive self-improvement engine - Think from first principles, let insights emerge' is highly general and lacks specific activation scope, constraints, or exclusion conditions, which could cause overly broad matching in systems that use manifest text to decide when to invoke a skill.

Content

No source excerpt is available for this finding.

Unpinned Dependencies

Low
Category
Supply Chain
Confidence
40% confidence
Finding

Dependencies lack version pinning, allowing potential malicious package updates. Consider pinning versions.

Content

Scanner excerpt · package.json (reported line 25)May include surrounding context.

json
"author": "haishushan",
  "license": "MIT",
  "devDependencies": {
    "vitest": "^1.6.1"
  },
  "engines": {
    "node": ">=18.0.0"

Unverifiable Dependency: vitest has 3 known advisory(ies) (CVE-2026-47429 (When Vitest UI server is listening, arbitrary file can be read and executed); CVE-2026-84373 (Vitest: Path Traversal / Arbitrary File Read via @vitest/mocker Redirect Mock); CVE-2025-24964 (Vitest allows Remote Code Execution when accessing a malicious website while Vit)), but the manifest does not pin a version, so it is unknown whether the installed release is affected

Low
Category
Supply Chain
Confidence
40% confidence
Finding

Dependency has known vulnerabilities (CVEs). Using packages with unpatched security flaws exposes the environment to known exploits.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.dangerous_exec

Shell command execution detected (child_process).

Critical
Code
suspicious.dangerous_exec
Location
evolve.js:725