Back to skill

Security audit

crew-contract

Security checks for vulnerabilities and agentic risk

Overview

This skill is a disclosed workflow guide for using bounded subagents and reviewers, with practical cost and supply-chain cautions but no hidden malicious behavior found.

Install only from a source you trust, and consider pinning or reviewing the resolved package before using the npx command. Expect higher token/API cost when orchestration is used, and use it for bounded, verifiable work rather than simple questions or sensitive actions without explicit approval.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Supply ChainUnpinned Dependencies, External Script Fetching, Obfuscated Code
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
Findings (21)

Rp1

Medium
Category
MCP Rug Pull
Confidence
90% confidence
Finding

The README instructs users to run npx skills add mfang0126/crew-contract without pinning a specific package version. This creates a supply-chain risk: users may fetch whatever version is current at execution time, and if the package or dependency chain is compromised later, the same installation command could execute unreviewed code.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
87% confidence
Finding

The README instructs users to run npx skills add mfang0126/crew-contract without pinning a specific package version or immutable reference. That can cause consumers to fetch whatever version is current at install time, increasing supply-chain risk if the package, tag, or upstream dependency is later changed or compromised.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
90% confidence
Finding

The 'When to use' section lists trigger phrases such as 'delegate this' and also says the skill applies to 'any work that materially benefits from decomposition,' which is a broad condition that overlaps with common conversational requests. Although some negative examples are provided later, the activation language remains open-ended enough to risk unintentional invocation in ordinary task discussions.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
92% confidence
Finding

The README instructs users to run npx skills add mfang0126/crew-contract without pinning a specific package version or immutable source reference. That means installation behavior can change over time or be affected by a compromised upstream package/publisher, creating a supply-chain risk for anyone following the documented command.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
89% confidence
Finding

The README instructs users to run npx skills add mfang0126/crew-contract without any version pinning or integrity control. Because npx resolves and executes the latest published package by default, a compromised upstream package, typo-squatted dependency, or malicious update could lead to arbitrary code execution on the user's machine during installation.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The activation criteria are broad enough to apply to most non-trivial tasks, which can cause the agent to enter an expensive multi-agent delegation mode by default. In an autonomous-agent skill, overbroad triggering increases unnecessary autonomy, token/cost exposure, and the chance that risky or sensitive tasks are delegated when a simpler, more controllable main-thread flow would be safer.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
80% confidence
Finding

This section endorses opening subagents based on the main agent's judgment and explicitly frames child work as proceeding without asking questions, which can enable autonomous decisions under incomplete requirements. In this skill's context, that raises the risk of incorrect or unauthorized actions being taken through delegated agents, especially because the skill is designed as a general-purpose orchestration method across engineering, ops, and other potentially sensitive domains.

Content

Scanner excerpt · skills/crew-contract/SKILL.md (reported line 113)May include surrounding context.

md
## Orchestration signals (when it pays)

Orchestration pays when the work is **parallelizable, bounded, and verifiable**: streams can run independently; each stream can be pinned down without asking questions (children can't ask); each has a check (test, ground truth, source, artifact) someone else can apply. It degrades on strictly sequential work (controlled study across 180 agent configurations: -39% to -70%), shared-file/state coupling, ambiguous requirements, or tiny tasks; coordination overhead grows with dependencies and tool count. Cost is ~15x chat tokens — the task's value must cover it. Signals, not gates: run ONE bounded child first, extend only if it clearly pays.

**Admission gate:** open subagents only when ≥2 of these hold — (a) genuinely parallelizable/decomposable; (b) does not fit one context window; (c) task value covers orchestration cost. Record a one-line written justification on the route card (the format is checkable; the judgment stays with the main agent).

Rp1

Medium
Category
MCP Rug Pull
Confidence
92% confidence
Finding

The README instructs users to run npx skills add mfang0126/crew-contract without pinning a specific version or immutable source. That means installation behavior can change over time or be influenced by a compromised upstream package, exposing users to supply-chain risk during install. The skill context makes this somewhat more dangerous because the document is explicitly an installation guide, so readers are likely to copy-paste the command directly.

Content

No source excerpt is available for this finding.

Rp1

Medium
Category
MCP Rug Pull
Confidence
70% confidence
Finding

npx commands without a version suffix (e.g. @1.0.0) create a rug-pull risk if the upstream server is compromised and publishes a malicious update.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The description says the skill applies to engineering, research, writing, data, ops — 'anything' and further says to use it for 'any work' benefiting from decomposition. Although some examples are provided, the trigger scope remains very broad and overlaps with many ordinary task requests, making it unclear when the skill should activate versus remain inactive.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · SKILL.md (reported line 113)May include surrounding context.

md
## Orchestration signals (when it pays)

Orchestration pays when the work is **parallelizable, bounded, and verifiable**: streams can run independently; each stream can be pinned down without asking questions (children can't ask); each has a check (test, ground truth, source, artifact) someone else can apply. It degrades on strictly sequential work (controlled study across 180 agent configurations: -39% to -70%), shared-file/state coupling, ambiguous requirements, or tiny tasks; coordination overhead grows with dependencies and tool count. Cost is ~15x chat tokens — the task's value must cover it. Signals, not gates: run ONE bounded child first, extend only if it clearly pays.

**Admission gate:** open subagents only when ≥2 of these hold — (a) genuinely parallelizable/decomposable; (b) does not fit one context window; (c) task value covers orchestration cost. Record a one-line written justification on the route card (the format is checkable; the judgment stays with the main agent).

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · skills/crew-contract/skills/crew-contract/SKILL.md (reported line 112)May include surrounding context.

md
## Orchestration signals (when it pays)

Orchestration pays when the work is **parallelizable, bounded, and verifiable**: streams can run independently; each stream can be pinned down without asking questions (children can't ask); each has a check (test, ground truth, source, artifact) someone else can apply. It degrades on strictly sequential work (controlled study across 180 agent configurations: -39% to -70%), shared-file/state coupling, ambiguous requirements, or tiny tasks; coordination overhead grows with dependencies and tool count. Cost is ~15x chat tokens — the task's value must cover it. Signals, not gates: run ONE bounded child first, extend only if it clearly pays.

**Admission gate:** open subagents only when ≥2 of these hold — (a) genuinely parallelizable/decomposable; (b) does not fit one context window; (c) task value covers orchestration cost. Record a one-line written justification on the route card (the format is checkable; the judgment stays with the main agent).

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
76% confidence
Finding

The file includes a dedicated Chinese-language section and points readers to a Chinese README, while the policy requires avoiding forced language or locale constraints unless there is user choice or clear justification. Although the top link bar lists both English and Chinese, the inline Chinese section itself does not clearly frame language as an optional user choice.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
90% confidence
Finding

The file presents itself as the Chinese version and labels the language switch as '[English] | 中文', making this document Chinese-only by default. Under the policy, forcing a specific language without user opt-in is a natural-language locale constraint unless clearly documented as region-specific or offering an explicit choice within the skill behavior.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
78% confidence
Finding

The description states the skill should be used when a task should use main-sub delegation and includes the Chinese phrase '主-子分工' as part of the canonical invocation description. This can create a locale-specific expectation without explicitly offering user language choice or clarifying that the phrase is optional rather than required.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
77% confidence
Finding

The file includes a dedicated Chinese summary section and an English/Chinese link selector, but it does not explicitly frame the bilingual presentation as user-selectable policy or opt-in behavior. Because SQP-3 covers language/locale policy concerns, this could be read as a mild locale assumption in the skill's natural-language presentation.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

This markdown file is written as a Chinese variant README and line L05 explicitly presents '中文' for the current document, without any in-file statement that the skill adapts to the user's preferred language. Under the policy rule, forcing a specific language without user opt-in is a natural-language locale violation unless the constraint is clearly documented and justified.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
95% confidence
Finding

Line L05 presents the Chinese README as the active content ('中文') with only a link out to English, so this file itself is language-specific and does not offer an in-flow language choice or opt-in. Under the stated policy, forcing a specific language without user opt-in is a natural-language policy concern.

Content

No source excerpt is available for this finding.

Scope Creep

Low
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill's behavior or capabilities extend beyond its stated purpose. Scope creep allows an agent to perform actions unrelated to its documented functionality, increasing the attack surface.

Content

Scanner excerpt · SKILL.md (reported line 106)May include surrounding context.

md
## Failure and scope handling

- A subagent hitting an architectural decision, schema/API change, new dependency, security-sensitive choice, unclear requirement, or overlapping ownership must **stop and report** — not expand scope.
- On failure: inspect why; narrow, retry, reassign, or handle it in the main thread. Never silently ignore.
- Never claim delegated work finished unless it actually ran and its result was checked.
- Before finishing: every required agent has completed or explicitly failed; nothing required is still running.

Scope Creep

Low
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill's behavior or capabilities extend beyond its stated purpose. Scope creep allows an agent to perform actions unrelated to its documented functionality, increasing the attack surface.

Content

Scanner excerpt · skills/crew-contract/SKILL.md (reported line 106)May include surrounding context.

md
## Failure and scope handling

- A subagent hitting an architectural decision, schema/API change, new dependency, security-sensitive choice, unclear requirement, or overlapping ownership must **stop and report** — not expand scope.
- On failure: inspect why; narrow, retry, reassign, or handle it in the main thread. Never silently ignore.
- Never claim delegated work finished unless it actually ran and its result was checked.
- Before finishing: every required agent has completed or explicitly failed; nothing required is still running.

Scope Creep

Low
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill's behavior or capabilities extend beyond its stated purpose. Scope creep allows an agent to perform actions unrelated to its documented functionality, increasing the attack surface.

Content

Scanner excerpt · skills/crew-contract/skills/crew-contract/SKILL.md (reported line 105)May include surrounding context.

md
## Failure and scope handling

- A subagent hitting an architectural decision, schema/API change, new dependency, security-sensitive choice, unclear requirement, or overlapping ownership must **stop and report** — not expand scope.
- On failure: inspect why; narrow, retry, reassign, or handle it in the main thread. Never silently ignore.
- Never claim delegated work finished unless it actually ran and its result was checked.
- Before finishing: every required agent has completed or explicitly failed; nothing required is still running.

Static analysis

No suspicious patterns detected.