Back to skill

Security audit

ralph

Security checks for vulnerabilities and agentic risk

Overview

The skill is a mostly transparent Ralph setup guide, but it recommends a broad permission-bypass wrapper for autonomous Claude Code loops without enforcing replacement safeguards.

Review before installing or following the setup. The skill does not contain malicious code, but do not run Ralph through a bypassPermissions wrapper unless you have explicit, tested guard hooks or another approval mechanism for destructive commands, credential-sensitive actions, and public GitHub posting.

Vulnerability Patterns
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (1)

T05 · Unauthorized Access and Privilege Escalation

Error
Location
SKILL.md:29
Finding

Autonomous execution configured to bypass native permission safeguards

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, lines 29–36
Vulnerability Type: Permission safeguards bypass
Risk Level: High

Vulnerable snippet:

markdown
Ralph's circuit breaker counts permission denials. Claude Code's permission system splits commands at `|`, `&&`, etc. without honoring shell quotes, so commands like `gh ... --jq '.[] | select(...)'` get denied even though they are safe. A wrapper that forces `--permission-mode bypassPermissions` works around this while your own destructive-command guard hooks (if any) stay active.

```bash
# ~/.ralphrc (per workspace)
CLAUDE_CODE_CMD="claude-wrapper-ralph.sh"
text

### Technical Analysis

The Skill recommends using a wrapper that unconditionally starts Claude Code with `--permission-mode bypassPermissions`. This is a broad bypass rather than a narrowly scoped exception for the command-parsing false positives described in the document.

The suggested compensating control—destructive-command guard hooks—is optional (“if any”), is not shipped by this project, and is not technically enforced by the documented setup. Consequently, an operator can follow the recommended configuration while having no replacement authorization gate.

The trigger is installation of the described wrapper and configuration of `CLAUDE_CODE_CMD` to invoke it. Once an autonomous Ralph loop runs through that wrapper, model-generated tool actions cross from an approval-gated environment into execution under the operator's account without normal per-operation permission checks.

### Attack Path

1. An operator follows the recommended wrapper setup in `SKILL.md`.
2. The wrapper launches Claude Code with `--permission-mode bypassPermissions`.
3. The operator configures `.ralphrc` so the autonomous loop uses that wrapper.
4. No destructive-command guard is installed, or an existing guard does not cover the generated action.
5. The autonomous loop produces a sensitive or
...[truncated 894 chars]
Remediation
View remediation

Remediation Suggestions

  • Do not recommend or automatically use --permission-mode bypassPermissions.
  • Address the circuit-breaker false positives with narrowly scoped command allowlists or validated wrapper rules that permit only the specific safe command forms required by Ralph.
  • Retain explicit approval for destructive, credential-sensitive, network-publishing, and privilege-affecting operations.
  • If a wrapper is necessary, make it fail closed and reject commands outside a documented allowlist.
  • Ship or require an enforceable guard rather than relying on optional operator-specific hooks.
  • Validate guard availability before starting the autonomous loop and abort if the required authorization controls are missing.
  • Document the exact permissions granted to the loop and require the operator to opt in explicitly after reviewing that scope.
Vulnerability Patterns
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Tool MisuseTool Parameter Abuse, Chaining Abuse, Unsafe Defaults
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
Findings (3)

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
75% confidence
Finding

Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).

Content

Scanner excerpt · SKILL.md (reported line 37)May include surrounding context.

md
## PUBLIC repo language guard (autonomous loops)

Ralph autonomous loops MUST translate non-English content to English before posting to any GitHub PUBLIC repo (issues, PRs, comments). A `--no-confirm` flag does not exempt this rule. Per-call procedure:

1. `gh repo view --json isPrivate -q '.isPrivate'` — `false` means PUBLIC
2. Scan the body for non-ASCII script characters; if any are found, translate to English first

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 67)May include surrounding context.

md
Either path provides this `SKILL.md` for reference only. It intentionally has no SessionStart guards and no PreToolUse hooks — those are enforcement concerns yo

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
65% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.md (reported line 37)May include surrounding context.

md
## PUBLIC repo language guard (autonomous loops)

Ralph autonomous loops MUST translate non-English content to English before posting to any GitHub PUBLIC repo (issues, PRs, comments). A `--no-confirm` flag does not exempt this rule. Per-call procedure:

1. `gh repo view --json isPrivate -q '.isPrivate'` — `false` means PUBLIC
2. Scan the body for non-ASCII script characters; if any are found, translate to English first

Static analysis

No suspicious patterns detected.