Back to skill

Security audit

skills-eval

Security checks for vulnerabilities and agentic risk

Overview

The skill is a coherent skill-evaluation guide, but some instructions could run tools from untrusted skills on the user's machine without adequate containment.

Install only if you intend to audit skills and are comfortable with the skill reading skill files you point it at. Do not run its integration or benchmarking workflows against untrusted skills on your normal machine; use a disposable sandbox with no secrets and no network by default. Review any referenced external plugin scripts separately before using auto-fix, scan-all, or dynamic test modes.

Vulnerability Patterns
  • Insecure Skill Coding PracticesFinds exploitable flaws such as hardcoded secrets or command injection
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
Findings (2)

T09 · Insecure Skill Coding Practices

Error
Location
modules/integration-testing.md:60
Finding

Integration testing executes untrusted target-skill tools without isolation

Content
View full analysis

Vulnerability Details

File Location: modules/integration-testing.md, lines 60-87
Vulnerability Type: Unsafe execution of an untrusted executable during skill evaluation
Risk Level: High

Vulnerable Code

python
def test_tool_integration(skill_path: str) -> ToolIntegrationResults:
    """Test tool compatibility and integration"""
    results = ToolIntegrationResults()

    # Parse skill frontmatter
    frontmatter = parse_frontmatter(skill_path)
    declared_tools = frontmatter.get('tools', [])

    # Test each declared tool
    for tool in declared_tools:
        tool_path = find_tool(tool, skill_path)
        if tool_path and tool_path.exists():
            results.tools_found.append(tool)

            # Test tool is executable
            if os.access(tool_path, os.X_OK):
                results.tools_executable.append(tool)

            # Test tool runs with --help
            try:
                subprocess.run([tool_path, '--help'],
                              capture_output=True,
                              timeout=5)
                results.tools_functional.append(tool)
            except Exception as e:
                results.tool_errors.append(f"{tool}: {e}")
        else:
            results.tools_missing.append(tool)

    return results

Technical Analysis

The documented integration-testing procedure resolves executable paths from the skill being evaluated and launches each executable directly with subprocess.run. A --help argument is not a security boundary: an executable controls its own argument handling and may run arbitrary initialization or payload logic before displaying help, or ignore the argument entirely.

The procedure does not establish the provenance or integrity of the executable and does not use a sandbox, reduced-privilege account, read-only filesystem, network isolation, environment sanitization, or syscall restrictions. The ...[truncated 1555 chars]

Remediation
View remediation

Remediation Suggestions

  1. Make static inspection the default and do not execute tools contained in an untrusted target skill.
  2. Require explicit user authorization before entering a clearly labeled dynamic-analysis mode.
  3. Execute dynamic tests only in a disposable container or virtual machine configured with:
    • An unprivileged user and no additional Linux capabilities.
    • No host credential, SSH-agent, cloud-token, or Docker-socket mounts.
    • No network access by default.
    • Read-only target artifacts and a disposable writable directory.
    • A minimal, sanitized environment and controlled PATH.
    • CPU, memory, process-count, output-size, and wall-clock limits.
    • Seccomp, AppArmor, SELinux, or equivalent syscall restrictions where available.
  4. Resolve and validate the final canonical executable path, rejecting symlinks or paths outside the copied sandbox input.
  5. Verify tool hashes or signatures against trusted metadata when provenance is available.
  6. Treat exit status explicitly: use check=True, record nonzero results as failures, and avoid marking a tool functional merely because process creation succeeded.
  7. Destroy the complete execution environment after every target and preserve only bounded, sanitized diagnostic output.

T09 · Insecure Skill Coding Practices

Error
Location
modules/performance-benchmarking.md:30
Finding

Performance benchmark repeatedly executes untrusted tools without containment or timeout

Content
View full analysis

Vulnerability Details

File Location: modules/performance-benchmarking.md, lines 30-49
Vulnerability Type: Repeated unsafe execution of a target-supplied executable
Risk Level: High

Vulnerable Code

python
def measure_tool_execution(self, tool_path: str) -> Dict[str, float]:
    """Measure tool execution time"""
    results = {}

    # Warm-up run
    subprocess.run([tool_path, '--help'], capture_output=True)

    # Benchmark runs
    times = []
    for _ in range(10):
        start = time.perf_counter()
        subprocess.run([tool_path, '--help'], capture_output=True)
        end = time.perf_counter()
        times.append((end - start) * 1000)

    results['mean'] = statistics.mean(times)
    results['median'] = statistics.median(times)
    results['std_dev'] = statistics.stdev(times)
    results['min'] = min(times)
    results['max'] = max(times)

    return results

Technical Analysis

The benchmarking example directly invokes the supplied tool_path once for warm-up and ten more times for measurement. Any executable referenced by that path can perform arbitrary actions, regardless of the --help argument.

No trust validation or execution containment is shown. Unlike the integration-testing example, these calls also have no timeout, allowing a target executable to hang indefinitely. Repeated invocation can amplify destructive or resource-exhaustion behavior. Capturing standard output and standard error does not restrict filesystem, network, process, or credential access, and unbounded captured output may itself consume excessive memory.

Attack Path

  1. An attacker provides or causes the benchmark to select a malicious executable as tool_path.
  2. The benchmark starts the executable during its warm-up phase.
  3. The malicious program executes with the benchmark process's permissions and can ignore --help.
  4. If it returns normally, the benchmark invokes i ...[truncated 948 chars]
Remediation
View remediation

Remediation Suggestions

  1. Do not benchmark executables from untrusted skills on the host system.
  2. Separate parsing and static analysis from explicitly authorized dynamic benchmarking.
  3. Run each benchmark target in a fresh, disposable sandbox with no secrets, no network by default, read-only inputs, an unprivileged identity, and strict filesystem and syscall policies.
  4. Apply a short timeout to every invocation and terminate the complete process group on expiration.
  5. Bound captured output and redirect excess output to a size-limited disposable file to prevent memory exhaustion.
  6. Enforce CPU, memory, process-count, and file-size limits independently of the Python timeout.
  7. Validate that tool_path resolves inside the sandboxed target directory and is not a symlink, wrapper, or path traversal to a host executable.
  8. Avoid repeated execution until one isolated validation run has completed successfully; recreate the sandbox between benchmark iterations when testing untrusted artifacts.
  9. Check and report exit codes rather than treating any completed invocation as a valid benchmark sample.
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Excessive AgencyUnrestricted Tool Access, Autonomous Decision Making, Scope Creep
  • System Prompt LeakageDirect Leakage, Indirect Extraction, Tool-Based Exfiltration
  • Memory PoisoningPersistent Context Injection, Context Window Stuffing, Memory Manipulation
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
Findings (10)

Memory Manipulation

High
Category
Memory Poisoning
Confidence
80% confidence
Finding

Skill manipulates agent memory, state, or stored context. Memory corruption can alter personality, override safety rules, or cause unpredictable behavior.

Content

Scanner excerpt · modules/evaluation-criteria.md (reported line 441)May include surrounding context.

md
### Activation (20 points)
- [ ] **5+ trigger phrases in description**
- [ ] Clear context indicators
- [ ] Differentiates from alternatives
- [ ] Easy to discover

Anti-Refusal Statement

High
Category
Anti-Refusal
Confidence
80% confidence
Finding

Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.

Content

Scanner excerpt · modules/troubleshooting.md (reported line 20)May include surrounding context.

md
**Technical Limits**:
- **Default budget**: Skill description budget scales at 2% of context window (~20,000 chars for 1M context)
- **No warning system**: There's currently no notification when you exceed this threshold
- **Silent failure**: Skills beyond the budget are simply not included in Claude's system prompt

**Solutions**:

Agent Config Directory Access

High
Category
Agent Snooping
Confidence
85% confidence
Finding

Skill reads from agent configuration directories (.claude/, .codex/, .gemini/). These directories may contain API keys, personal settings, and other credentials that the skill has no legitimate need to access.

Content

Scanner excerpt · modules/troubleshooting.md (reported line 44)May include surrounding context.

  1. Audit Skill Count and Size:
bash
# Count total skills
find ~/.claude/skills -name "SKILL.md" | wc -l

# Measure description field sizes
grep -A 5 "^description:" ~/.claude/skills/*/SKILL.md

Agent Config Directory Access

High
Category
Agent Snooping
Confidence
85% confidence
Finding

Skill reads from agent configuration directories (.claude/, .codex/, .gemini/). These directories may contain API keys, personal settings, and other credentials that the skill has no legitimate need to access.

Content

Scanner excerpt · modules/troubleshooting.md (reported line 47)May include surrounding context.

md
find ~/.claude/skills -name "SKILL.md" | wc -l

# Measure description field sizes
grep -A 5 "^description:" ~/.claude/skills/*/SKILL.md

# Estimate total description budget usage
# (requires custom script - see skills-eval tools)

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The trigger list includes broad, generic phrases such as 'evaluation', 'improvement', 'skills', and 'optimization' that are likely to match many ordinary user requests. This can cause unintended activation of the skill in unrelated contexts, increasing prompt-surface exposure and creating opportunities for misrouting, overreach, or interference with other skills and instructions.

Content

No source excerpt is available for this finding.

Autonomous Decision Making

Medium
Category
Excessive Agency
Confidence
75% confidence
Finding

Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.

Content

Scanner excerpt · modules/skill-authoring-best-practices.md (reported line 271)May include surrounding context.

md
| Abstract Quick Start | Reader cannot copy-paste | Use literal commands |
| Relative cross-refs | Breaks across installs | Use `Skill()` form |
| Second-person voice | Treated as user docs | Convert to third person |
| No verification section | Claude declares done early | Add explicit checks |
| Stale cited paths | Hallucinated content | Re-verify on each release |

## How to use this module

Skill Enumeration

Medium
Category
Agent Snooping
Confidence
85% confidence
Finding

Skill enumerates or reads other installed skills. Access to other skills' SKILL.md files or the skills directory reveals prompt instructions, capabilities, and secrets that should be invisible to peer skills.

Content

Scanner excerpt · modules/troubleshooting.md (reported line 43)May include surrounding context.

  1. Audit Skill Count and Size:
bash
# Count total skills
find ~/.claude/skills -name "SKILL.md" | wc -l

# Measure description field sizes

Skill Enumeration

Medium
Category
Agent Snooping
Confidence
85% confidence
Finding

Skill enumerates or reads other installed skills. Access to other skills' SKILL.md files or the skills directory reveals prompt instructions, capabilities, and secrets that should be invisible to peer skills.

Content

Scanner excerpt · modules/troubleshooting.md (reported line 100)May include surrounding context.

Issue: No skills found during discovery

bash
# Solution: Verify Claude configuration and skill locations
ls ~/.claude/skills/
skills/skills-eval/scripts/skills-auditor --discover

Missing User Warnings

Low
Category
Not specified by scanner
Confidence
92% confidence
Finding

This markdown file provides a command that writes benchmark output to benchmark-report.md, which affects local user data. The surrounding documentation does not warn that running the command will create or overwrite a file, so the data-modifying behavior is not explicitly disclosed.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
85% confidence
Finding

The example elicitation dialogue hard-codes an English conversation between the agent and user, including suggested prompts and responses, with no indication that other languages are acceptable. Because this file is instructional guidance for creating tests, the examples can function as normative templates and implicitly force English usage without user opt-in.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.