Back to skill

Security audit

skill-authoring

Security checks for vulnerabilities and agentic risk

Overview

This is a disclosed skill-authoring guide with strong behavior-shaping advice, but it is markdown-only and does not show hidden execution, exfiltration, or persistence.

Install this if you want an opinionated guide for building and testing Claude Code skills. Review the broad activation triggers and the persuasion-focused modules first, and only run the example validation, deployment, git, or authentication commands when they match your repo and you intend those local changes.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • Memory PoisoningPersistent Context Injection, Context Window Stuffing, Memory Manipulation
  • Rogue AgentSelf-Modification, Session Persistence
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • YARA SignaturesMalware Match, Webshell Match, Cryptominer Match
Findings (17)

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 82)May include surrounding context.

md
te the smallest intervention that addresses the documented failures. Write the `SKILL.md` with required frontmatter and content that directly counters the basel

Memory Manipulation

High
Category
Memory Poisoning
Confidence
90% confidence
Finding

Skill manipulates agent memory, state, or stored context. Memory corruption can alter personality, override safety rules, or cause unpredictable behavior.

Content

Scanner excerpt · modules/error-handling.md (reported line 201)May include surrounding context.

md
For idempotent reads (a `gh api` GET), one retry is reasonable.
For state-changing operations (commits, pushes, file writes),
silent retries cause duplicate work or corrupt state. Default
to no retry. Add retries only for specific commands with
documented idempotence.

YARA rule 'agent_skill_prompt_injection_hidden_instructions': Prompt injection or hidden instructions embedded in AI agent skill text [agent_skills]

High
Category
YARA Match
Confidence
84% confidence
Finding

This module explicitly teaches authors how to increase an LLM's compliance using authority, commitment, scarcity, social proof, and identity framing. In an agent skill context, those same techniques can be used to override user intent, suppress model safety checks, or make downstream prompts more resistant to correction, so the finding is a real prompt-injection risk even if presented as 'best practices' guidance.

Content

Scanner excerpt · modules/persuasion-principles.md (reported line 11)May include surrounding context.

md
n Principles for Skill Design

## Overview

Skills are behavioral interventions. Research shows that incorporating compliance psychology principles can dramatically improve adherence rates. This module covers evidence-based persuasion techniques for skill design.

## Research Foundation

### Meincke et al. (2025): Persuasion Doubles Compliance

**Study**: Persuasive Paraphrasing in Large Language Models

**Key Finding**: Incorporating persuasion principles into instructions doubled compliance rates:
- **Baseline**: 33% compliance with standard instructions
- **Persuasive**: 72% compliance with persuasion-enhanced instructions
- **Improvement**: 118% increase (more than doubled)

**Implication**: How you phrase requirements dramatically affects whether Claude follows them.

### Six Principles of Persuasion (Cialdini)

The study validated Cialdini's classic persuasion principles in LLM contexts:

1. **Authority**: Directive language from credible sources
2. **Commitment & Consistency**:

Self-Modification

High
Category
Rogue Agent
Confidence
85% confidence
Finding

Skill modifies its own code, configuration, or behavior at runtime. Self-modification enables an agent to escalate privileges, disable safety constraints, or install persistent backdoors.

Content

Scanner excerpt · modules/progressive-disclosure.md (reported line 170)May include surrounding context.

md
## Example: Secure API Skill

1. Document baseline (3 scenarios)
2. Write skill addressing failures
3. Add anti-rationalization

For additional examples, see `modules/examples.md`

Self-Modification

High
Category
Rogue Agent
Confidence
85% confidence
Finding

Skill modifies its own code, configuration, or behavior at runtime. Self-modification enables an agent to escalate privileges, disable safety constraints, or install persistent backdoors.

Content

Scanner excerpt · modules/progressive-disclosure.md (reported line 277)May include surrounding context.

md
## Example: Secure API Skill

1. Document baseline (3 scenarios)
2. Write skill addressing failures
3. Add anti-rationalization

For additional examples, see `modules/examples.md`

Self-Modification

High
Category
Rogue Agent
Confidence
85% confidence
Finding

Skill modifies its own code, configuration, or behavior at runtime. Self-modification enables an agent to escalate privileges, disable safety constraints, or install persistent backdoors.

Content

Scanner excerpt · modules/tdd-methodology.md (reported line 461)May include surrounding context.

md
## Example: Secure API Skill

1. Document baseline (3 scenarios)
2. Write skill addressing failures
3. Add anti-rationalization

For additional examples, see `modules/examples.md`

Self-Modification

High
Category
Rogue Agent
Confidence
85% confidence
Finding

Skill modifies its own code, configuration, or behavior at runtime. Self-modification enables an agent to escalate privileges, disable safety constraints, or install persistent backdoors.

Content

Scanner excerpt · modules/tdd-methodology.md (reported line 285)May include surrounding context.

3. Add Explicit Counters

Update skill with direct counters to observed rationalizations:

markdown
## Common Rationalizations (DO NOT USE)

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The trigger list contains broad terms such as "writing" and "validation" that are common in many unrelated prompts, which can cause this skill to activate unintentionally. In a skill that teaches persuasion and behavior-shaping patterns, accidental invocation can bias outputs, increase prompt surface area, and interfere with user intent even if no direct code execution or data exfiltration occurs.

Content

No source excerpt is available for this finding.

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
94% confidence
Finding

This markdown file discusses a skill that loads when the user says "implement X" or "fix bug Y." Those phrases are broad, common software-assistance requests and the text does not provide narrowing constraints or negative examples, so they could cause unintended activation collisions.

Content

No source excerpt is available for this finding.

Skill Enumeration

Medium
Category
Agent Snooping
Confidence
80% confidence
Finding

Skill enumerates or reads other installed skills. Access to other skills' SKILL.md files or the skills directory reveals prompt instructions, capabilities, and secrets that should be invisible to peer skills.

Content

Scanner excerpt · modules/examples.md (reported line 29)May include surrounding context.

md
**Skill**: `Skill(pensive:tiered-audit)`
**Path**:
`/home/alext/claude-night-market/plugins/pensive/skills/tiered-audit/SKILL.md`

What to copy from this skill:

Skill Enumeration

Medium
Category
Agent Snooping
Confidence
80% confidence
Finding

Skill enumerates or reads other installed skills. Access to other skills' SKILL.md files or the skills directory reveals prompt instructions, capabilities, and secrets that should be invisible to peer skills.

Content

Scanner excerpt · modules/examples.md (reported line 55)May include surrounding context.

md
**Skill**: `Skill(abstract:subagent-testing)`
**Path**:
`/home/alext/claude-night-market/plugins/abstract/skills/subagent-testing/SKILL.md`

This skill exists because `skill-authoring` requires baseline
tests in fresh subagents. Rather than restating the testing

Skill Enumeration

Medium
Category
Agent Snooping
Confidence
80% confidence
Finding

Skill enumerates or reads other installed skills. Access to other skills' SKILL.md files or the skills directory reveals prompt instructions, capabilities, and secrets that should be invisible to peer skills.

Content

Scanner excerpt · modules/examples.md (reported line 84)May include surrounding context.

md
**Skill**: `Skill(imbue:proof-of-work)`
**Path**:
`/home/alext/claude-night-market/plugins/imbue/skills/proof-of-work/SKILL.md`

This skill enforces evidence before "done." It is referenced
by `skill-authoring/SKILL.md` as the source of detailed

Skill Enumeration

Medium
Category
Agent Snooping
Confidence
80% confidence
Finding

Skill enumerates or reads other installed skills. Access to other skills' SKILL.md files or the skills directory reveals prompt instructions, capabilities, and secrets that should be invisible to peer skills.

Content

Scanner excerpt · modules/examples.md (reported line 117)May include surrounding context.

md
**Skill**: `Skill(scribe:slop-detector)`
**Path**:
`/home/alext/claude-night-market/plugins/scribe/skills/slop-detector/SKILL.md`

What to copy:

Skill Enumeration

Medium
Category
Agent Snooping
Confidence
80% confidence
Finding

Skill enumerates or reads other installed skills. Access to other skills' SKILL.md files or the skills directory reveals prompt instructions, capabilities, and secrets that should be invisible to peer skills.

Content

Scanner excerpt · modules/examples.md (reported line 138)May include surrounding context.

md
**Skill**: `Skill(abstract:skill-graph-audit)`
**Path**:
`/home/alext/claude-night-market/plugins/abstract/skills/skill-graph-audit/SKILL.md`

What to copy:

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · modules/troubleshooting.md (reported line 217)May include surrounding context.

md
2. **Skill instructs Claude to run a command without
   capturing output**. Claude infers what the output would
   have been.
3. **No verification step**. The skill produces an artifact
   without checking it.

### Diagnosis

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · modules/troubleshooting.md (reported line 218)May include surrounding context.

md
capturing output**. Claude infers what the output would
   have been.
3. **No verification step**. The skill produces an artifact
   without checking it.

### Diagnosis

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
88% confidence
Finding

This markdown file includes executable verification commands such as wc -l and python plugins/abstract/scripts/skills_auditor.py --skill ... but does not provide any warning that readers will be running local commands in their environment. Under the markdown-specific SQP-2 criteria, operational steps that affect system integrity or execute code should include a brief disclosure or caution.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.