Back to skill

Security audit

Agent Orchestrator

Security checks for vulnerabilities and agentic risk

Overview

This skill matches its agent-orchestration purpose, but it asks for broad persistent tool access and can route ordinary coding or status messages into high-impact agent workflows.

Install only if you want AO to become the default path for coding work in OpenClaw and you are comfortable with persistent sessions, worktrees, PR activity, GitHub CLI credentials, and Anthropic API usage. Prefer dedicated, least-privilege GitHub and LLM credentials, do not symlink .env or secret files into worktrees, and avoid enabling broad plugin access unless you have reviewed the other installed plugins.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T01 · Skill Instruction Hijacking

Error
Location
SKILL.md:24
Finding

Mandatory AO Routing Hijacks the Agent's Normal Tool Selection

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, lines 24-30, 45-47, and 84-87
Vulnerability Type: T01: Skill Instruction Hijacking
Risk Level: High

Vulnerable Code

text
## How You Think

Every user message is either:
1. **About work or code** → use AO tools
2. **About something else** → respond normally

When the user explicitly asks about work, issues, or status — use the tools for live data instead of answering from memory.
text
You don't wait for the user to say "spawn" or "use AO." You detect intent and act.
text
### Rule 1: Tools first, always
When the user asks anything about work, tasks, issues, status, or projects:
- FIRST call tools to get live data
- THEN present the results
- NEVER answer work questions from memory

Technical Analysis

The Skill defines mandatory, session-level routing instructions that intercept nearly every coding, project, issue, and status request. It explicitly directs the Agent to invoke AO without requiring the user to request AO and prohibits the Agent from selecting an ordinary workflow.

This exceeds the minimum behavior necessary for an optional orchestration Skill. A safer design would expose AO as an opt-in capability while allowing the host Agent and user to decide whether orchestration is appropriate. Instead, the phrases “Tools first, always,” “You don't wait,” and “NEVER” attempt to supersede normal tool-selection behavior.

Although the Skill separately requires confirmation before spawning an agent, that safeguard only covers spawning and batch spawning. It does not prevent automatic calls to status, issue, session, review, verification, or other AO tools, nor does it address the broader redirection of the Agent's behavior.

Attack Path

  1. The Skill is loaded into an Agent session.
  2. A user submits an ordinary coding, issue, project, or status request without explicitly selecting AO.
  3. The mandatory in ...[truncated 1062 chars]
Remediation
View remediation

Remediation Suggestions

  1. Make AO explicitly opt-in. Invoke its tools only when the user directly asks to use AO or clearly approves its use for a specific task.
  2. Remove mandatory language such as “Tools first, always,” “You don't wait,” and “NEVER answer.”
  3. Preserve the host Agent's normal policies, safety constraints, and tool-selection authority.
  4. Require confirmation before every operation that can access repositories, contact external services, alter branches, create pull requests, or launch coding agents.
  5. Before confirmation, disclose which repository, external provider, credentials, and operations will be involved.
  6. Limit AO routing to narrowly defined requests rather than treating every coding or project request as an AO request.
  7. Allow the user to choose a non-AO workflow without disabling the Skill.

T05 · Unauthorized Access and Privilege Escalation

Error
Location
SKILL.md:119
Finding

Setup Enables an Overbroad Global Plugin Tool Profile

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, lines 119-128
Vulnerability Type: T05: Unauthorized Access and Privilege Escalation
Risk Level: High

Vulnerable Code

bash
# Required: allow plugin tools to be visible to the AI
# (plugin tools are optional by default in OpenClaw — this enables them)
openclaw config set tools.profile "full"
openclaw config set tools.allow '["group:plugins"]'

# Required: trust this plugin
openclaw config set plugins.allow '["agent-orchestrator"]'

# Optional: increase message context for group chats
openclaw config set messages.groupChat.historyLimit 100

The accompanying explanation states:

text
**Why `tools.profile: "full"`?** OpenClaw's default `coding` profile only includes built-in tools. Plugin-provided tools (like `ao_spawn`, `ao_issues`) require the `full` profile to be visible to the AI. This does not grant additional system permissions — it only makes plugin tools discoverable.

Technical Analysis

The setup procedure changes the global tool profile to full and allows the entire group:plugins namespace. This is broader than allowing only the Agent Orchestrator tools needed for the declared functionality.

Tool visibility is an effective capability boundary even when it does not directly modify operating-system permissions. Making additional plugin tools callable expands the actions the Agent can request and increases exposure to vulnerable, malicious, misconfigured, or subsequently installed plugins. Therefore, the claim that the change does not grant additional system permissions understates the practical security impact.

The risk is amplified by the Skill's mandatory routing instructions. Once the global plugin group is available, broad work-related prompts are automatically directed toward plugin-based workflows. The project documentation identifies legitimate external access through GitHub credentials and ANTHROPIC_API_KEY; su ...[truncated 1499 chars]

Remediation
View remediation

Remediation Suggestions

  1. Do not require the global full profile when a narrower profile can support the Skill.
  2. Replace group:plugins with an explicit allowlist containing only the AO tools required for approved operations.
  3. Separate read-only tools, such as status inspection, from state-changing tools, such as spawning, killing, cleanup, repository modification, and pull-request creation.
  4. Require per-operation user approval for state-changing or network-enabled tools.
  5. Apply configuration at the narrowest available scope, such as the specific Skill, project, or session, rather than globally.
  6. Document the exact capabilities enabled by each AO tool and identify which external services and credentials it can use.
  7. Use a fine-grained GitHub token restricted to necessary repositories and permissions, and a dedicated LLM API key with spending limits.
  8. Add validation that rejects secret-file symlinks and prevents .env, credentials, tokens, and other sensitive files from entering agent worktrees.
  9. Revise the setup explanation to acknowledge that expanded tool visibility increases the Agent's effective capability and attack surface even if operating-system permissions are unchanged.
Vulnerability Patterns
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
Findings (4)

Vague Triggers

High
Category
Not specified by scanner
Confidence
98% confidence
Finding

Including vague confirmations like 'sure', 'yes', 'go ahead', and 'do it' as direct coding-action triggers is dangerous because those phrases are common, context-dependent, and can be elicited after an attacker or prior prompt frames a spawn action. Given this skill can create git worktrees, start autonomous coding agents, and potentially open PRs, an accidental match can trigger real external actions with code and GitHub side effects.

Content

No source excerpt is available for this finding.

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · references/config.md (reported line 61)May include surrounding context.

md
# Options: reuse | delete | ignore | delete-new | ignore-new | kill-previous

    # Workspace setup (optional)
    symlinks:                  # Symlink into worktrees (avoid .env or secret files)
      - node_modules
    postCreate:                # Run after workspace creation
      - pnpm install

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The work/issue trigger examples are overly broad and include casual phrases like 'let's go', 'morning', and 'ready to work', which may be interpreted as instructions to query repositories and sessions. This increases the chance of unintended tool use and repository/task disclosure from ordinary conversation rather than deliberate operator intent.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The skill contains conflicting control logic: it says to always confirm before spawning, but the earlier intent mapping also treats broad confirmations like 'yes', 'go ahead', and 'do it' as triggers for ao_spawn. In practice, this can cause ambiguous or context-free confirmations to launch coding agents and create worktrees/PR activity without a clearly scoped, freshly confirmed action.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.