Back to skill

Security audit

Skill Security Review

Security checks for vulnerabilities and agentic risk

Overview

This skill is a transparent security-review checklist with no executable code, network behavior, persistence, or hidden payloads.

Reasonable to install if you want an agent-side checklist for reviewing skills before installation. Expect it to slow down skill installs by asking for an audit and confirmation first; consider replacing the user-specific wording with neutral language before broad reuse.

Vulnerability Patterns
  • Skill Instruction HijackingAlters the agent's session goals or safety constraints when the skill loads
  • Agent Memory PoisoningWrites attacker-controlled rules into memory that affect later sessions
  • Remote Payload Retrieval and ExecutionFetches external code whose behavior can change after review
  • Embedded Malicious CodeShips malicious scripts inside the skill and executes them locally
  • Unauthorized Access and Privilege EscalationObtains permissions beyond the task's legitimate needs
Vulnerability Patterns
  • Data ExfiltrationExternal Transmission, Env Variable Harvesting, File System Enumeration
  • Rogue AgentSelf-Modification, Session Persistence
  • Trigger AbuseOverly Broad Trigger, Shadow Command Trigger, Keyword Baiting Trigger
  • YARA SignaturesMalware Match, Webshell Match, Cryptominer Match
  • Prompt InjectionInstruction Override, Hidden Instructions, Exfiltration Commands
Findings (9)

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 27)May include surrounding context.

md
- Determine whether the target is a local folder, `.skill` archive, git repo, pasted `SKILL.md`, script bundle, or agent prompt.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 31)May include surrounding context.

md
- Determine whether the target is a local folder, `.skill` archive, git repo, pasted `SKILL.md`, script bundle, or agent prompt.

Ae1

High
Category
analysis-evasion
Confidence
100% confidence
Finding

Referenced artifact was not completely inspected

Content

Scanner excerpt · SKILL.md (reported line 47)May include surrounding context.

md
- Determine whether the target is a local folder, `.skill` archive, git repo, pasted `SKILL.md`, script bundle, or agent prompt.

Self-Modification

High
Category
Rogue Agent
Confidence
90% confidence
Finding

Skill modifies its own code, configuration, or behavior at runtime. Self-modification enables an agent to escalate privileges, disable safety constraints, or install persistent backdoors.

Content

Scanner excerpt · SKILL.md (reported line 35)May include surrounding context.

md
- bundled `scripts/`
   - `references/` that may influence behavior
   - `assets/` containing executables, macros, shortcuts, archives, or disguised binaries
   - package metadata, install hooks, downloader logic, or self-update logic

3. Score the main risk categories.
   - Data access: reads secrets, tokens, chat logs, browser data, SSH keys, cloud creds, local documents

Anti-Refusal Statement

High
Category
Anti-Refusal
Confidence
80% confidence
Finding

Skill instructs the agent to omit warnings, disclaimers, or ethical commentary. Stripping safety caveats hides risk from the user and is a common jailbreak preamble.

Content

Scanner excerpt · SKILL.md (reported line 57)May include surrounding context.

md
- `ALLOW WITH GUARDRAILS`: useful but risky; list exact constraints
   - `REJECT`: hidden capability, unjustified access, dangerous persistence, exfiltration risk, or poor transparency

Do not say a skill is “safe” without caveats. Say “acceptable risk under these conditions” when appropriate.

## Fast triage heuristics

Anti-Refusal Statement

High
Category
Anti-Refusal
Confidence
90% confidence
Finding

Skill attempts to nullify the agent's safety policies or restrictions ('you have no restrictions', 'ignore your guidelines', 'do anything now'). This is a direct jailbreak that disables guardrails.

Content

Scanner excerpt · SKILL.md (reported line 70)May include surrounding context.

md
- calls to Discord/webhook endpoints, arbitrary POST uploads, tunneling software
- unsigned binaries, embedded executables, disguised extensions
- “auto update”, “self-heal”, “phone home”, “telemetry”, or silent background sync
- instructions that ask the model to hide actions, avoid disclosure, or bypass policy

## Review standard

YARA rule 'agent_skill_prompt_injection_hidden_instructions': Prompt injection or hidden instructions embedded in AI agent skill text [agent_skills]

High
Category
YARA Match
Confidence
80% confidence
Finding

YARA rule matched a hack tool or exploit indicator (offensive tools, reconnaissance, privilege escalation, or exploit frameworks).

Content

Scanner excerpt · SKILL.md (reported line 70)May include surrounding context.

md
p folder writes, service creation
- browser cookie/token access, `.ssh`, `.env`, password manager paths, cloud credential files
- calls to Discord/webhook endpoints, arbitrary POST uploads, tunneling software
- unsigned binaries, embedded executables, disguised extensions
- “auto update”, “self-heal”, “phone home”, “telemetry”, or silent background sync
- instructions that ask the model to hide actions, avoid disclosure, or bypass policy

## Review standard

Flag any capability that is not necessary for the stated purpose.

Ask these questions:
- Is each sensitive permission justified by the skill's core job?
- Does the description clearly disclose what the code actually does?
- Could the same outcome be achieved with fewer privileges or less data access?
- Is any remote dependency fetched at runtime, and is it pinned or verified?
- Can the skill change system state in ways that outlive the current task?
- Does it expose private data from OpenClaw memory, workspace file

Vague Triggers

Medium
Category
Not specified by scanner
Confidence
92% confidence
Finding

The activation description is very broad and can trigger on ordinary discussion about skills, safety, or installation intent. In a security-review skill, overbroad auto-invocation can unnecessarily intercept unrelated conversations, influence user workflows, and increase the chance that this skill is applied in contexts the user did not clearly request.

Content

No source excerpt is available for this finding.

Natural-Language Policy Violations

Low
Category
Not specified by scanner
Confidence
78% confidence
Finding

The text references a specific user name, which can hard-code assumptions about the operator, locale, or trust boundary. While not directly exploitable on its own, it can cause privacy, personalization, or policy-scoping issues if the skill is reused across different users or environments.

Content

No source excerpt is available for this finding.

Static analysis

No suspicious patterns detected.