Back to skill

Security audit

codebuddy-cli-delegation

Security checks for vulnerabilities and agentic risk

Overview

This skill is a disclosed workflow for delegating large coding tasks to CodeBuddy CLI, with expected risks around unattended code changes and stored logs.

Install only if you intend to delegate repository changes to CodeBuddy CLI. Use it with a clean worktree or separate worktree, give narrow file and tool scope, review diffs and rerun tests yourself, and keep secrets out of prompts and .cli-runs logs. Treat bypassPermissions as local automatic approval for the child CLI, not permission to push, force-reset, delete data, or touch production without explicit approval.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Privilege EscalationExcessive Permissions, Sudo/Root Execution, Credential Access
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Rogue AgentSelf-Modification, Session Persistence
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
Findings (12)

Credential Access

High
Category
Privilege Escalation
Confidence
60% confidence
Finding

Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.

Content

Scanner excerpt · .gitignore (reported line 25)May include surrounding context.

text
.tmp/

# 本地环境与凭据(绝不入库)
.env
.env.*
*.local

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · README.md (reported line 9)May include surrounding context.

md
**English summary** — A reusable Agent Skill for delegating large batches of coding work to a
standalone **headless CodeBuddy CLI** agent. It covers the clean-environment launcher (which
strips host-injected environment variables that otherwise cause silent zero-byte hangs),
instruction budgeting, parallelism admission rules (4-way isolation + disjoint write sets),
per-stage parameter allocation, and independent audit & acceptance through adversarial probes
and mutation controls — never trusting the agent's own report. All 21 supported models and every
other claim are empirically verified against **CodeBuddy CLI 2.156.0**.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

The README says the skill can be triggered by directly saying 「用 CLI 派活」, but it does not define the exact trigger scope, exclusions, or negative examples. For a markdown skill description, this phrase is broad enough to overlap with ordinary requests about delegating work via a CLI, which could cause unintended activation.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
70% confidence
Finding

Without declared permissions the skill's intent is opaque and cannot be validated.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
97% confidence
Finding

The trigger list is broad and includes common phrases like '交给 CLI 做' and '前台会超时', which increases the chance that the skill activates in situations the user did not intend. Because this skill advocates unattended execution, permission bypass, and delegated code changes, accidental invocation can expand the blast radius from a simple misunderstanding into unsupervised actions or risky recommendations.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
92% confidence
Finding

The skill explicitly recommends persistent local logging of prompts and full CLI output into .cli-runs/, and elsewhere references session logs under ~/.codebuddy/projects/. In the context of delegated coding tasks, these logs can capture prompts, code, paths, model/session identifiers, environment-derived details, and potentially sensitive operational context, creating a durable audit trail that may leak secrets or proprietary data if stored insecurely or committed accidentally.

Content

Scanner excerpt · SKILL.md (reported line 580)May include surrounding context.

  • 统一日志落盘,便于事后审计:
sh
mkdir -p .cli-runs
LOG=".cli-runs/$(date +%H%M%S)-<任务名>.log"
timeout 1500 node .workbuddy/bin/codebuddy-cli.js \
  -p "$(cat .cli-runs/<任务名>.prompt.md)" \

Rp1

Medium
Category
MCP Rug Pull
Confidence
70% confidence
Finding

npx commands without a version suffix (e.g. @1.0.0) create a rug-pull risk if the upstream server is compromised and publishes a malicious update.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The header comments, usage text, and runtime error/help strings are written entirely in Chinese, and the file provides no indication that language choice is configurable or that Chinese is required for a region-specific purpose. This creates a natural-language locale policy issue because the skill effectively imposes a specific language on users without opt-in.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The template is entirely written in Chinese and prescribes mandatory reporting/output conventions without offering a language-selection mechanism or user opt-in. In a delegation/bootstrap prompt, this can cause the agent to ignore or override the user's preferred language, increasing the chance of misunderstanding task requirements, incorrect execution, and inaccessible audit output for operators who do not read Chinese.

Content

No source excerpt is available for this finding.

Scope Creep

Low
Category
Excessive Agency
Confidence
70% confidence
Finding

Skill's behavior or capabilities extend beyond its stated purpose. Scope creep allows an agent to perform actions unrelated to its documented functionality, increasing the attack surface.

Content

Scanner excerpt · LICENSE (reported line 12)May include surrounding context.

text
permit persons to whom the Software is furnished to do so.

THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED,
INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A
PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT
HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION
OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
92% confidence
Finding

The file explicitly says the skill body and the rest of the README are written in Chinese, identifying Chinese as the primary audience, but it does not offer the user a language or locale choice. This is a natural-language locale constraint that is not presented as optional or justified as a region-specific compliance requirement.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
95% confidence
Finding

The natural-language content of the skill is overwhelmingly in Chinese, with only a small English keyword fragment, and it does not state that the user can choose another language. This can violate a language/locale policy when a skill effectively forces one language by default without documented user opt-in or a justified locale restriction.

Content

No source excerpt is available for this finding.

Static analysis

Detected: suspicious.dangerous_exec

Shell command execution detected (child_process).

Critical
Code
suspicious.dangerous_exec
Location
scripts/codebuddy-cli.js:168