Back to skill

Security audit

Skill Compass

Security checks for vulnerabilities and agentic risk

Overview

The skill is mostly a coherent skill-auditing utility, but its optional auto-fix mode can persistently edit installed skill files with weak scoping and misleading side-effect documentation.

Review before installing if you let agents run remediation automatically. Use the audit mode read-only first, pass an explicit --skills-dir, and avoid --fix unless you have backups or have reviewed the exact SKILL.md files it may change.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Rogue AgentSelf-Modification, Session Persistence
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • MCP Least PrivilegeUnderdeclared Capability, Wildcard Permission, Missing Permission Declaration
  • MCP Tool PoisoningHidden Instructions, Unicode Deception, Parameter Description Injection
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (8)

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
93% confidence
Finding

The README advertises an auto-fix flag but does not clearly warn that it will modify skill files. In an agent or automation context, omission of write-side effects can cause users to run the command expecting a read-only audit, leading to silent changes across a skills directory and possible breakage, data loss, or unauthorized edits.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Medium
Category
Not specified by scanner
Confidence
96% confidence
Finding

The README states 'Side effects: None' while elsewhere documenting an auto-fix mode that edits skill files. This is a security-relevant misrepresentation because users or agents may trust the documentation and invoke the tool in write mode without realizing it can modify local files, increasing the chance of unintended configuration changes or damage to installed skills.

Content

No source excerpt is available for this finding.

Undeclared Tool Scope

Medium
Category
MCP Least Privilege
Confidence
70% confidence
Finding

Without declared permissions the skill's intent is opaque and cannot be validated.

Content

No source excerpt is available for this finding.

Session Persistence

Medium
Category
Rogue Agent
Confidence
60% confidence
Finding

Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.

Content

Scanner excerpt · SKILL.md (reported line 23)May include surrounding context.

md
**Bad:** `"Helps with documents."` — no trigger, just capability
**Good:** `"Use when the user asks to extract form fields, fill, redact, or parse tables from a PDF file."`

**Fix:** Rewrite using the Description Design Pattern below.

### 2. YAML Frontmatter Format Error

Description-Behavior Mismatch

Medium
Category
Not specified by scanner
Confidence
95% confidence
Finding

The script is presented as an audit/diagnostic tool, but when run with --fix it modifies SKILL.md files in place. This mismatch can lead users or higher-level agents to invoke it assuming read-only behavior, causing unintended persistent changes to skill definitions and potentially degrading or altering downstream agent behavior.

Content

No source excerpt is available for this finding.

Missing User Warnings

Medium
Category
Not specified by scanner
Confidence
89% confidence
Finding

In auto-fix mode, the script writes directly to SKILL.md files with no interactive confirmation, backup, or explicit warning at the point of modification. In an agent skill ecosystem, silent writes to configuration/trigger files can unexpectedly change what skills activate, making this more dangerous than a normal text-processing utility because it can alter agent behavior across future runs.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
78% confidence
Finding

SQP-3 applies to all file types and covers language or locale policy violations in natural-language content. Line L086 directs adding multilingual keywords, and the nearby example at L071 hard-codes Chinese trigger terms; this nudges skills toward forcing locale-specific behavior without explicit user opt-in or a documented regional justification.

Content

No source excerpt is available for this finding.

Intent-Code Divergence

Low
Category
Not specified by scanner
Confidence
84% confidence
Finding

The file docstring and CLI parser describe this as 'Skill Compass' for auditing trigger issues, but the report banner prints 'SKILL SENTINEL — AUDIT REPORT'. This is not just missing detail; it presents a different tool identity than the surrounding documentation, which can mislead operators about what component is running.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.