Back to skill

Security audit

Epic AI Swarm Orchestration

Security checks for vulnerabilities and agentic risk

Overview

This is a real multi-agent coding swarm tool, but it gives autonomous agents broad execution, network, git, and persistence authority with several unsafe defaults.

Install only in an isolated workspace or disposable account with provider credentials scoped for this swarm. Review and harden spawn-agent.sh before use, especially sandbox bypass, automatic dependency installs, task input validation, /tmp file handling, and fail-open review logic. Avoid running it on private or sensitive repositories unless you are comfortable with external model providers, git pushes, background watchers, and configured notification targets receiving project metadata.

Vulnerability Patterns
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • Insecure DependenciesIntroduces malicious components through unsafe dependency sources
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
Findings (5)

T05 · Unauthorized Access and Privilege Escalation

Error
Location
scripts/spawn-agent.sh:225
Finding

AI Coding Agent Runs with Approval and Sandbox Protections Disabled

Content
View full analysis
"` # fails in tmux with "stdout is not a terminal". exec is the designed # one-shot entry point. # Use --dangerously-bypass-approvals-and-sandbox so codex can write to # the parent repo's .git/worktrees/ metadata (required for `git commit` # inside a worktree). Default workspace-write sandbox blocks ORIG_HEAD.lock. codex exec --model "$model" -c "model_reasoning_effort=$REASONING" --dangerously-bypass-approvals-and-sandbox "$PROMPT" ;; ``` ### Technical Analysis The `--dangerously-bypass-approvals-and-sandbox` option explicitly disables both command approval and filesystem sandbox protections. The resulting model-driven process inherits the invoking user's effective permissions rather than receiving access only to the selected project worktree and necessary Git metadata. The prompt can include operator-supplied task text, repository-derived content, and later instructions derived from work logs. Repository files may contain prompt-injection content. Because the agent can autonomously interpret that content and invoke tools without approval, untrusted repository instructions can cause access to files and commands unrelated to the coding task. The stated need to update Git worktree metadata does not justify unrestricted access to the user's home directory, provider credentials, SSH configuration, unrelated repositories, notification configuration, or arbitrary network-capable commands. ### Attack Path 1. An operator starts a task against a repository containing malicious instructions in source code, documentation, issue text, or generated files. 2. `spawn-agent.sh` incorporates the task prompt and directs Codex to inspect and modify ...[truncated 1259 chars]
Remediation
View remediation
` metadata exposed as writable. - The rest of the repository and host mounted read-only or not mounted. - No access to SSH agents, cloud credentials, browser profiles, or unrelated home-directory files. 3. Require explicit operator approval for: - Access outside the project. - Network clients other than the selected model provider. - Package installation and lifecycle scripts. - Git pushes, PR creation, and writes to protected branches. 4. Use a dedicated low-privilege service account with isolated provider credentials. 5. Apply outbound network allowlisting and prevent arbitrary destinations. 6. Treat repository text as untrusted data and tell agents not to follow instructions found in repository content unless they are relevant to the endorsed task. 7. If worktree metadata remains incompatible with the default sandbox, create a narrow helper for required Git operations rather than disabling the entire sandbox. ]]>

T09 · Insecure Skill Coding Practices

Error
Location
scripts/spawn-agent.sh:325
Finding

Shell and Python Code Injection Through Unsafely Interpolated Task Parameters

Content
View full analysis
/dev/null || { echo "⚠️ tmux session $TMUX_SESSION already exists" exit 1 } # ... python3 -c " import json task = { 'id': '$TASK_ID', 'tmuxSession': '$TMUX_SESSION', 'role': '$ROLE', 'agent': '$AGENT', 'model': '$MODEL', 'description': '''$DESCRIPTION'''[:200], 'project': '$PROJECT_NAME', 'projectDir': '$PROJECT_DIR', 'worktree': '$WORK_DIR', 'branch': '$BRANCH', 'startedAt': $TIMESTAMP, 'status': 'running', 'retries': 0, 'notifyOnComplete': True } with open('$TASKS_FILE') as f: data = json.load(f) data['tasks'].append(task) with open('$TASKS_FILE', 'w') as f: json.dump(data, f, indent=2) " 2>/dev/null && echo "📋 Task registered" || echo "⚠️ Task registration failed (non-critical)" [[ -f "$USAGE_LOG" ]] || echo '[]' > "$USAGE_LOG" python3 -c " import json from datetime import datetime entry = { 'timestamp': datetime.now().isoformat(), 'taskId': '$TASK_ID', 'role': '$ROLE', 'agent': '$AGENT', 'model': '$MODEL', 'project': '$PROJECT_NAME' } with open('$USAGE_LOG') as f: data = json.load(f) data.append(entry) with open('$USAGE_LOG', 'w') as f: json.dump(data, f, indent=2) " 2>/dev/null && echo "📊 Usage logged" || true ``` ### Technical Analysis Task IDs, d ...[truncated 2283 chars]
Remediation
View remediation

T08 · Insecure Dependencies

Error
Location
scripts/spawn-agent.sh:187
Finding

Automatic Dependency Installation Executes Untrusted Package Lifecycle Scripts

Content
View full analysis
/dev/null || true elif [[ -f "yarn.lock" ]]; then yarn install 2>/dev/null || true elif [[ -f "package-lock.json" ]]; then npm install 2>/dev/null || true fi fi ``` ### Technical Analysis Starting an agent automatically runs the target repository's package manager. Standard package-manager installation can execute package lifecycle hooks such as `preinstall`, `install`, `postinstall`, and `prepare`. These hooks are executable code supplied by the repository or its transitive dependencies. There is no explicit installation consent, lifecycle-script suppression, isolated execution environment, source allowlist, dependency integrity policy beyond normal package-manager behavior, or separation from host credentials. Errors are also suppressed with `2>/dev/null || true`, reducing visibility into malicious or failed installation activity. Because installation occurs before model review and because the main agent can be launched without sandbox protections, a malicious repository or compromised dependency can execute with the same user privileges as the Skill. ### Attack Path 1. The operator targets an untrusted or compromised repository containing `package.json` and a recognized lockfile. 2. The repository declares a malicious lifecycle hook, or its lockfile references a compromised package with such a hook. 3. `spawn-agent.sh` automatically runs `npm install`, `pnpm install`, or `yarn install`. 4. The package manager executes the lifecycle hook without separate operator approval. 5. The hook reads credentials, alters files, contacts external systems, or installs additional user-level code. 6. Installation failures or suspicious outp ...[truncated 637 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
scripts/notify-on-complete.sh:208
Finding

Predictable Shared Temporary Files Enable Symlink Attacks and Review Manipulation

Content
View full analysis
"$REVIEW_PROMPT_FILE" << REVIEW_PROMPT_EOF You are a code REVIEWER+FIXER at $REVIEW_PROJECT. Round $LOOP of $MAX_LOOPS. # ... REVIEW_PROMPT_EOF # ... REVIEW_WRAPPER="/tmp/review-wrapper-${REVIEW_SESSION}.sh" if echo "$REVIEW_CMD" | grep -q "^gemini"; then GEMINI_BASE=$(echo "$REVIEW_CMD" | sed -E 's/[[:space:]]+-p[[:space:]]*$//; s/[[:space:]]*$//') cat > "$REVIEW_WRAPPER" << WRAPPER_EOF #!/usr/bin/env bash cd "$REVIEW_PROJECT" $GEMINI_BASE -y -p "\$(cat $REVIEW_PROMPT_FILE)" WRAPPER_EOF else cat > "$REVIEW_WRAPPER" << WRAPPER_EOF #!/usr/bin/env bash cd "$REVIEW_PROJECT" $REVIEW_CMD "\$(cat $REVIEW_PROMPT_FILE)" WRAPPER_EOF fi chmod +x "$REVIEW_WRAPPER" tmux new-session -d -s "$REVIEW_SESSION" -c "$REVIEW_PROJECT" "bash $REVIEW_WRAPPER" ``` Other predictable temporary paths include: ```bash WORKLOG="/tmp/worklog-${TMUX_SESSION}.md" ``` ```bash VERDICT_FILE="/tmp/review-verdict-${INTEG_SESSION}.json" INTEG_PROMPT_FILE="/tmp/integration-prompt-${INTEG_SESSION}.md" INTEG_WRAPPER="/tmp/integ-wrapper-${INTEG_SESSION}.sh" ``` ```python tmp = tempfile.mktemp(prefix='/tmp/standup_list_') ``` ### Technical Analysis The Skill creates prompts, work logs, verdict files, and executable wrapper scripts directly in the shared `/tmp` directory using names derived from predictable session identifiers. It does not create a private per-run directory, use exclusive file creation, verify file ownership, or reject symbolic links. For wrapper scripts, the code writes to a predic ...[truncated 1961 chars]
Remediation
View remediation

T09 · Insecure Skill Coding Practices

Error
Location
scripts/notify-on-complete.sh:304
Finding

Missing or Malformed Review Verdicts Fail Open as Successful Reviews

Content
View full analysis
/dev/null | head -1) WORKLOG_UPDATED=false if [[ -f "$WORKLOG" ]] && grep -q "Review.*Round $LOOP\|review.*round $LOOP\|Round $LOOP" "$WORKLOG" 2>/dev/null; then WORKLOG_UPDATED=true fi if echo "$LATEST_MSG" | grep -qi "review\|fix\|pass\|clean"; then SUMMARY="Review passed — reviewer fixed issues (commit: $LATEST_MSG)" PASS="True" send_telegram "✅ $BASE_SESSION review passed (round $LOOP): fixed issues → $LATEST_MSG" elif [[ "$WORKLOG_UPDATED" == "true" ]]; then SUMMARY="Review passed — reviewer found no issues (work log updated, no fixes needed)" PASS="True" send_telegram "✅ $BASE_SESSION review passed (round $LOOP): no issues found" else SUMMARY="Review passed — reviewer exited cleanly (auto-pass: clean exit, no issues indicated)" PASS="True" send_telegram "✅ $BASE_SESSION review auto-passed (round $LOOP): clean exit, no issues indicated" fi if [[ -x "$SWARM_DIR/update-task-status.sh" ]]; then "$SWARM_DIR/update-task-status.sh" --session "$BASE_SESSION" "done" 2>/dev/null || true fi ``` The integration watcher applies equivalent fail-open behavior: ```bash if [[ ! -f "$VERDICT_FILE" ]]; then LATEST_MSG=$(cd "$PROJECT_DIR" && git log --oneline -1 2>/dev/null | head -1) INTEG_LOG_UPDATED=false # ... if echo "$LATEST_MSG" | grep -qi "integration\|review\|fix\|pass\|c ...[truncated 2721 chars]
Remediation
View remediation
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Tool MisuseTool Parameter Abuse, Chaining Abuse, Unsafe Defaults
  • Rogue AgentSelf-Modification, Session Persistence
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
Findings (57)

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

Uninstall/removal and role deactivation are administrative/destructive capabilities that go beyond the visible install/run/diagnose framing. On shared OpenClaw hosts, hidden removal behavior can disable roles, delete runtime assets, and disrupt other users or automations if invoked unintentionally.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

Uninstall/removal and role deactivation are administrative/destructive capabilities that go beyond the visible install/run/diagnose framing. On shared OpenClaw hosts, hidden removal behavior can disable roles, delete runtime assets, and disrupt other users or automations if invoked unintentionally.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
99% confidence
Finding

Uninstall/removal and role deactivation are administrative/destructive capabilities that go beyond the visible install/run/diagnose framing. On shared OpenClaw hosts, hidden removal behavior can disable roles, delete runtime assets, and disrupt other users or automations if invoked unintentionally.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

Uninstall/removal and role deactivation are administrative/destructive capabilities that go beyond the visible install/run/diagnose framing. On shared OpenClaw hosts, hidden removal behavior can disable roles, delete runtime assets, and disrupt other users or automations if invoked unintentionally.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
98% confidence
Finding

Uninstall/removal and role deactivation are administrative/destructive capabilities that go beyond the visible install/run/diagnose framing. On shared OpenClaw hosts, hidden removal behavior can disable roles, delete runtime assets, and disrupt other users or automations if invoked unintentionally.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
95% confidence
Finding

Uninstall/removal and role deactivation are administrative/destructive capabilities that go beyond the visible install/run/diagnose framing. On shared OpenClaw hosts, hidden removal behavior can disable roles, delete runtime assets, and disrupt other users or automations if invoked unintentionally.

Content

No source excerpt is available for this finding.

Tp4

High
Category
MCP Tool Poisoning
Confidence
97% confidence
Finding

Uninstall/removal and role deactivation are administrative/destructive capabilities that go beyond the visible install/run/diagnose framing. On shared OpenClaw hosts, hidden removal behavior can disable roles, delete runtime assets, and disrupt other users or automations if invoked unintentionally.

Content

No source excerpt is available for this finding.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 149)May include surrounding context.

md
`spawn-batch.sh` starts `integration-watcher.sh`, which waits for all sessions, merges branches, runs verification, persists work logs/ESR, and writes notificat

External Model or Provider Selection

High
Category
Excessive Agency
Confidence
90% confidence
Finding

Skill selects an external model or provider that may use a different account or billing plan than the operator expects. Undisclosed model switches can cause unexpected cost or quota consumption.

Content

Scanner excerpt · references/tools.md (reported line 100)May include surrounding context.

bash
bash model-fallback.sh builder deepseek deepseek-v4-pro-max
# → codex|gpt-5.5|codex exec --model gpt-5.5 --full-auto

Monitoring

Chaining Abuse

High
Category
Tool Misuse
Confidence
75% confidence
Finding

Tool calls are chained to bypass individual safety checks or escalate capabilities beyond what any single tool call would allow.

Content

Scanner excerpt · scripts/assess-models.sh (reported line 69)May include surrounding context.

sh
fi

  rm -f "$result_file" 2>/dev/null
  [[ -n "$tmpdir" ]] && rm -rf "$tmpdir" 2>/dev/null || true
}

PROBE="Reply with ONLY the word HELLO, nothing else."

Missing User Warnings

High
Category
Not specified by scanner
Confidence
96% confidence
Finding

The script sends full prompt contents to external agent CLIs and, for Codex, does so with sandbox protections disabled, yet it does not clearly disclose the data-transmission and elevated-execution implications. Because this skill is specifically designed to orchestrate external coding agents across providers, sensitive repository data and operator instructions may be exposed to third parties while the agent runs with broad authority.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

High
Category
Not specified by scanner
Confidence
99% confidence
Finding

The runner invokes Codex with --dangerously-bypass-approvals-and-sandbox, explicitly disabling key safety controls and allowing the model-directed process to operate with broader filesystem and git metadata access. In this skill's context, prompts are built from task descriptions and files, so prompt injection or unsafe instructions could directly lead to unauthorized code changes, command execution, or data exposure.

Content

No source excerpt is available for this finding.

External Model or Provider Selection

High
Category
Excessive Agency
Confidence
94% confidence
Finding

This command delegates task content to an external model/provider selected at runtime, creating a data egress path for repository context, prompts, and potentially sensitive content. In this orchestration skill, external provider use is core functionality, which makes the behavior expected but still security-significant when not tightly governed or disclosed.

Content

Scanner excerpt · scripts/spawn-agent.sh (reported line 231)May include surrounding context.

sh
# Use --dangerously-bypass-approvals-and-sandbox so codex can write to
      # the parent repo's .git/worktrees/ metadata (required for `git commit`
      # inside a worktree). Default workspace-write sandbox blocks ORIG_HEAD.lock.
      codex exec --model "$model" -c "model_reasoning_effort=$REASONING" --dangerously-bypass-approvals-and-sandbox "$PROMPT"
      ;;
    gemini)
      gemini --model "$model" -p "$PROMPT"

Intent-Code Divergence

High
Category
Not specified by scanner
Confidence
97% confidence
Finding

The script automatically creates per-task endorsement files solely so that spawn-agent.sh approval checks pass, but it does not verify any persisted human approval state before doing so. In a swarm orchestration context, this weakens a safety gate intended to require explicit authorization for each task and can allow unreviewed agent actions to proceed under the false appearance of approval.

Content

No source excerpt is available for this finding.

Tool Parameter Abuse

High
Category
Tool Misuse
Confidence
95% confidence
Finding

Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).

Content

Scanner excerpt · uninstall.sh (reported line 74)May include surrounding context.

sh
if [[ "$KEEP_ROLE" != "1" ]]; then
  if [[ -d "$WORKSPACE_DIR/roles/swarm-lead" ]]; then
    run rm -rf "$WORKSPACE_DIR/roles/swarm-lead"
  fi
  if [[ -f "$WORKSPACE_DIR/roles/active.json" ]]; then
    if [[ "$YES" != "1" ]]; then

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
88% confidence
Finding

The skill describes operational shell scripts, file mutation, environment use, and outbound-capable integrations, but it does not declare any explicit tool scope or permission boundaries. In a portable orchestration/install skill, missing scope increases the chance that an agent or host will grant broader-than-necessary shell, filesystem, or network access, which can enable unintended installs, state changes, or notifications.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

Broad trigger phrases such as 'spawn agents', 'run the swarm', or generic project-management language increase the risk of accidental invocation. Because this skill includes installation, file mutation, state management, and possible notification behaviors, unintended activation can lead to disruptive or destructive actions in the wrong context.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · references/duty-table.md (reported line 110)May include surrounding context.

"manualOverride": { "enabled": true, "setBy": "operator", "note": "Pinned during provider outage; do not auto-reset without approval." } }

text

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

This markdown instructs the agent to clear ~/workspace/swarm/pending-notifications.txt using shell redirection after relaying its contents, but it does not explicitly warn that the operation deletes the file contents. For markdown files, SQP-2 applies when descriptions omit warnings about behaviors that can affect user data or system state.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The script actively invokes external provider CLIs (Codex, Gemini, DeepSeek) to probe model availability and may also send a notification afterward, but it provides no explicit warning, confirmation step, or offline mode by default. In a portable swarm/orchestration skill, this can unexpectedly transmit prompts, metadata, and possibly use authenticated accounts or billable APIs when a user runs what appears to be a local assessment task.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
86% confidence
Finding

The manifest describes orchestration, packaging/install, model rotation, and runtime health for parallel AI coding swarms. This script implements a daily standup broadcaster via openclaw message send to Telegram, which is a team-notification capability rather than an obvious or explicitly declared part of swarm orchestration itself.

Content

No source excerpt is available for this finding.

Context-Inappropriate Capability

Medium
Category
Not specified by scanner
Confidence
84% confidence
Finding

The script can send messages to any externally configured channel and target from values loaded out of swarm.conf, with no validation, allowlist, or interactive confirmation. In a portable multi-host swarm package, that creates a real risk of unintended data disclosure or covert outbound signaling if configuration is altered, inherited, or maliciously planted.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The notification includes project name, repository identity, commit prefix, and build status, and transmits that information to an external destination without any user-facing notice at runtime. While the leaked data is not highly sensitive in all environments, in private repositories or internal projects it can expose development activity, repository identifiers, and failure timing to third parties.

Content

No source excerpt is available for this finding.

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
88% confidence
Finding

The manifest describes orchestration of parallel AI coding swarms, packaging/installing the swarm, model rotation, and runtime health. This script instead monitors GitHub Actions for a commit and sends out Telegram/OpenClaw notifications, which is CI/deployment notification behavior not clearly covered by the stated swarm orchestration purpose.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
80% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · scripts/inbox-add.sh (reported line 2)May include surrounding context.

sh
#!/usr/bin/env bash
# inbox-add.sh — Add a task to the swarm inbox for later batching
#
# Usage: inbox-add.sh <project-dir> <task-id> <description> [priority]
#   priority: high | medium | low (default: medium)

Static analysis

No suspicious patterns detected.