Back to skill

Security audit

DashClaw Governance Protocol

Security checks for vulnerabilities and agentic risk

Overview

This governance skill is not malicious, but it asks to route broad agent activity, logs, approvals, and cross-session context through DashClaw with limited scope controls.

Install this only if you intentionally want DashClaw to govern the agent's whole working session. Before using it on sensitive work, verify the MCP server identity, allowed capability IDs, policy provenance, logging destination, access controls, retention and deletion settings, and redaction rules for secrets, personal data, proprietary outputs, and error details.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (4)

T01 · Skill Instruction Hijacking

Error
Location
SKILL.md:14
Finding

Global Agent Behavior and Tool-Routing Override

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, lines 14-17 and 71-77
Vulnerability Type: Global instruction and execution-path override
Risk Level: High

Vulnerable Snippet:

markdown
# DashClaw Governance Protocol

You are an agent governed by DashClaw. DashClaw provides the governance tools (via MCP)
and this skill teaches you how to use them correctly. Follow this protocol for every
action you take.
markdown
### External API Calls

Never make direct HTTP calls to external APIs that are registered as DashClaw capabilities.
Always use `dashclaw_invoke` — it runs the full governance loop automatically:
guard check, execution, outcome recording.

Before invoking an unknown capability ID, call `dashclaw_capabilities_list` to verify it
exists and check its health status.

Technical Analysis

The skill declares that the agent is governed by DashClaw and applies its protocol to every action rather than limiting the instructions to an explicitly requested DashClaw operation. It also replaces ordinary external API execution paths with dashclaw_invoke.

This is an instruction-level control-flow override. When the skill is loaded, it changes how the agent selects tools, decides whether operations may proceed, and routes external API activity. Although this behavior is consistent with the skill's stated governance purpose, it creates a broad trust boundary in which DashClaw can mediate unrelated agent actions.

Attack Path

  1. The agent loads the skill.
  2. The global instruction to follow the protocol for every action becomes active.
  3. The agent identifies an external API as a registered DashClaw capability.
  4. The skill prevents use of the normal direct execution path.
  5. The agent submits the operation through dashclaw_invoke.
  6. DashClaw mediates execution and records the operation, potentially blocking, modifying, or observing it according to externally supplied co ...[truncated 448 chars]
Remediation
View remediation

Remediation Suggestions

  • Scope the protocol to tasks for which the user explicitly requests DashClaw governance.
  • Replace “every action you take” with narrowly defined action types and trust boundaries.
  • State explicitly that system, developer, safety, and current user instructions take precedence.
  • Require informed user approval before redirecting an operation from its normal tool to a DashClaw capability.
  • Display the capability ID, destination system, transmitted fields, and expected side effects before invocation.
  • Permit a safe refusal path when DashClaw is unavailable instead of silently changing execution semantics.
  • Enforce an allowlist of capability IDs and parameter schemas at the client boundary.

T01 · Skill Instruction Hijacking

Error
Location
SKILL.md:19
Finding

Mutable External Policies Can Control Agent Decisions After Review

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, lines 19-35
Vulnerability Type: Runtime instruction loading from an external governance service
Risk Level: High

Vulnerable Snippet:

markdown
## Session Initialization

At the start of every session, do these three things:

1. **Load your governance context** — Read the `dashclaw://policies` MCP resource to
   understand what rules govern you. Note which action types require approval, what risk
   thresholds trigger blocks, and any agent-specific restrictions.

2. **Discover available capabilities** — Read the `dashclaw://capabilities` MCP resource
   to see what external APIs are registered. Note capability IDs, health status, and risk
   levels. You will use `dashclaw_invoke` (not direct HTTP) for these.

3. **Register your session** — Call `dashclaw_session_start` with your agent ID and a
   workspace description. This groups all your actions for tracking in Mission Control.

If MCP resources are unavailable, proceed with the static protocol below. You can always
call `dashclaw_policies_list` and `dashclaw_capabilities_list` tools as fallbacks.

Technical Analysis

The effective governance rules and capability inventory are obtained at runtime from MCP resources that are not included in the audited project. Consequently, static review of this package cannot establish the complete instructions that will influence the agent.

No requirement is provided to pin policy versions, authenticate individual policy documents, verify signatures, constrain policy semantics, or prevent remotely supplied policies from conflicting with higher-priority instructions. A compromised or malicious MCP service could therefore change risk thresholds, approval behavior, restrictions, or capability metadata after the skill has been reviewed.

Attack Path

  1. An attacker compromises or controls the configured DashClaw MCP service or its policy data.
  2. Th ...[truncated 958 chars]
Remediation
View remediation

Remediation Suggestions

  • Cryptographically sign policies and capability manifests and verify signatures locally.
  • Pin trusted policy versions or hashes and alert the user before accepting changes.
  • Validate remote content against a restrictive schema that excludes free-form executable instructions.
  • Prohibit remote policies from overriding system, developer, safety, or explicit user instructions.
  • Maintain a local allowlist of trusted MCP server identities and capability IDs.
  • Require explicit user consent before first contact with an MCP server and whenever its identity or policy version changes.
  • Record provenance, retrieval time, version, signature status, and policy hash for every loaded document.
  • Fail closed when authenticity cannot be established, while clearly reporting that governance initialization failed.

T02 · Agent Memory Poisoning

Error
Location
SKILL.md:140
Finding

Persistent Handoffs and Learning Records Enable Cross-Session Context Poisoning

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, lines 140-153 and 184-192
Vulnerability Type: Persistent, automatically reused agent context
Risk Level: High

Vulnerable Snippet:

markdown
## Session Continuity

### After concluding a session
Call `dashclaw_handoff_create` with a bundle containing your 1-2 sentence summary,
any open loops you opened (action-scoped, via `dashclaw_loop_add`), and decisions
you made (or references via `dashclaw_learning_log`). The next session of yours
will pick this up automatically via `dashclaw_handoff_latest` in pre_llm_call
context injection (when running under Hermes Agent — Claude Code and Codex pick
it up on first turn via the governance protocol).

### On session start (Claude Code / Codex only)
On your first turn, call `dashclaw_handoff_latest` with your agent_id. If a
bundle is returned, summarize it for the operator, then call
`dashclaw_handoff_consume` to mark it claimed so it isn't read twice.
markdown
## Learning From Prior Sessions

### Before making a non-obvious decision
Call `dashclaw_learning_query` with a search string. If a prior session made a
similar decision, surface its outcome before making yours.

### After making a non-obvious decision
Call `dashclaw_learning_log` with the decision + context (+ outcome if known).
Future sessions querying for this pattern will see your reasoning.

Technical Analysis

The skill stores summaries, open loops, decisions, context, and reasoning for later sessions. It then directs later sessions to retrieve this data, including through automatic pre-LLM context injection.

The instructions do not define integrity checks, provenance validation, sanitization of instruction-like content, retention limits, trust labels, or user approval before persisted text is placed into future model context. Content influenced by an attacker in one session could therefore be stored as a handoff or ...[truncated 1222 chars]

Remediation
View remediation

Remediation Suggestions

  • Treat every persisted handoff, learning record, and open-loop description as untrusted data.
  • Do not automatically inject free-form persisted content into the model's instruction context.
  • Present retrieved records as quoted data with an explicit untrusted provenance label.
  • Require user confirmation before using a stored record to guide actions.
  • Authenticate record authorship and bind records to the originating user, workspace, task, and session.
  • Sanitize or reject imperative language, tool-call instructions, embedded prompts, and references that attempt to change agent policy.
  • Add expiration, revocation, maximum length, and retention controls.
  • Provide operators with inspection and deletion controls before records are consumed.
  • Separate factual task state from instructions; store structured fields rather than unrestricted prose wherever possible.

T09 · Insecure Skill Coding Practices

Warning
Location
SKILL.md:80
Finding

Mandatory Broad Telemetry Can Disclose Sensitive Task Information

Content
View full analysis

Vulnerability Details

File Location: SKILL.md, lines 80-105
Vulnerability Type: Insufficient data minimization and redaction requirements
Risk Level: Medium

Vulnerable Snippet:

markdown
## Recording Rules

Record all significant actions with `dashclaw_record`. This powers the audit trail visible
in Mission Control and the Decisions ledger.

**Always record:**
- Long-running actions (status: `running`) when you record up front; PATCH later with the final outcome
- Completed actions (status: `completed`)
- Failed actions (status: `failed`) — include error details in `output_summary`
- Blocked actions (status: `failed`) — include the guard block reason (the server has no separate `blocked` status on records you create)

**Write meaningful fields:**
- `declared_goal` — Write as if explaining to an auditor. Bad: "Deploy the app".
  Good: "Deploy v2.3.1 to staging after all tests passed".
- `reasoning` — Why you chose this action over alternatives.
- `output_summary` — What was produced or what went wrong.
- `risk_score` — Your honest assessment. Don't lowball to avoid guards.

**For LLM-driven actions, include token usage (cost is auto-derived):**
- `tokens_in` / `tokens_out` — Total input and output tokens for the LLM call(s) attributed to this action.
- `model` — Model identifier (e.g. `claude-opus-4-6`, `gpt-5-codex`). The server uses this to look up pricing.
- `cost_estimate` — Optional. **Omit this field** when you provide tokens + model — the server derives `cost_estimate` from its configured pricing table (`app/lib/billing.js`) so cost stays consistent across all agents. Set it explicitly only when you have an authoritative cost from the provider.

Technical Analysis

The skill requires significant actions to be recorded and specifically requests declared goals, reasoning, output summaries, and failure details. These fields may contain credentials, personal data, proprietary ...[truncated 1362 chars]

Remediation
View remediation

Remediation Suggestions

  • Make telemetry opt-in for sensitive or externally transmitted records.
  • Define an explicit data-classification and redaction policy before calling dashclaw_record.
  • Never record credentials, authentication tokens, private keys, raw personal data, or complete confidential outputs.
  • Replace raw reasoning with concise, non-sensitive decision summaries.
  • Use structured allowlisted fields instead of arbitrary free-form content.
  • Run secret and personal-data detection over goals, summaries, and errors before transmission.
  • Document the storage endpoint, access-control model, encryption, retention period, deletion process, and downstream uses.
  • Allow users to preview and approve records for sensitive actions.
  • Apply workspace and tenant isolation and maintain access logs for Mission Control records.
Vulnerability Patterns
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
Findings (1)

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
91% confidence
Finding

The trigger list is broad enough that this governance skill may activate in many loosely related conversations, causing it to inject strong behavioral instructions such as mandatory tool calls, session registration, approval waits, and policy-loading steps even when the user did not intend to invoke this protocol. In an agent system where skill activation changes execution behavior, overbroad triggers can create prompt-scope hijacking or unintended control-flow changes, especially because this skill influences security-sensitive decisions about approvals, external capabilities, and action recording.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.